Image processing device, image processing method, and program

JP7927507B2Active Publication Date: 2026-10-01CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022129247
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2026-10-01
Estimated Expiration
2042-08-15

AI Technical Summary

Benefits of technology

【0007】 本開示によれば、オブジェクトの概略形状を表す三次元形状データから、高精度の三次元形状データを得ることができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927507000002
    Figure 0007927507000002
  • Figure 0007927507000003
    Figure 0007927507000003
  • Figure 0007927507000004
    Figure 0007927507000004
Patent Text Reader

Abstract

To allow for obtaining high-precision three-dimensional shape data from three-dimensional shape data representing a rough shape of an object.SOLUTION: An image processing device disclosed herein is configured to: acquire three-dimensional shape data of an object existing in an imaging space; acquire multiple distance images representing distances to the object and corresponding to multiple viewpoints; evaluate the three-dimensional shape data on the basis of the multiple distance images; and make a correction involving eliminating unit elements estimated to be not representing the shape of the object among unit elements constituting the three-dimensional shape data on the basis of a result of the evaluation.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing technique for generating three-dimensional shape data of an object.

Background Art

[0002] As a method for generating three-dimensional shape data of an object (generally also referred to as a "3D model") based on a plurality of captured images obtained by imaging a subject (object) from different viewpoints, the visual hull intersection method is known. While the visual hull intersection method can stably obtain three-dimensional shape data of an object at high speed, it has the disadvantage of being prone to errors. Specifically, when the surface of an object has a curved or concave shape, the shape is approximated by a plane, which causes a fundamental problem that an error increases. To address this problem, Patent Document 1 discloses a technique for restoring an accurate three-dimensional shape of an object through the following procedures 1) to 4). 1) Generate a 3D model (approximate shape model) circumscribed to the subject by the visual hull intersection method based on the silhouette image of the subject generated from the captured image. 2) Generate approximate distance information from the imaging camera to the surface of the approximate shape model and local shape information of the subject based on the approximate shape model. 3) Perform a search using a function that stores local shape information with the approximate distance information as an initial value, to generate distance information from the imaging camera to the subject. 4) Restore the three-dimensional shape of the subject based on the distance information generated in 3) above and the silhouette image.

Prior Art Literature

Patent Literature

[0003]

Patent Document 1

Summary of the Invention

Problem to be Solved by the Invention

[0004] Even with the technology described in Patent Document 1, for example, in recessed areas, the difference between the local shape obtained based on the general shape model and the original local shape of the object could not be filled, resulting in errors in distance information and making it impossible to obtain three-dimensional shape data with sufficient accuracy.

[0005] This disclosure aims to obtain high-precision three-dimensional shape data from three-dimensional shape data representing the general shape of an object. [Means for solving the problem]

[0006] The image processing apparatus according to this disclosure includes: a first acquisition means for acquiring three-dimensional shape data of an object present in the imaging space; a second acquisition means for acquiring a plurality of distance images representing the distance to the object, each corresponding to a plurality of viewpoints; and a correction means for evaluating the three-dimensional shape data based on the plurality of distance images, and correcting the three-dimensional shape data by deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data based on the results of the evaluation. The correction means includes setting means for setting a threshold for each local space obtained by dividing the imaging space, the correction means sequentially determines a unit element of interest from among the unit elements constituting the three-dimensional shape data, and adds an evaluation value to the unit element of interest if the distance from the viewpoint corresponding to the distance image to the unit element of interest is smaller than the distance at the pixel position on the distance image corresponding to the unit element of interest, and deletes the unit element of interest if the cumulative value of the evaluation value exceeds the threshold, the three-dimensional shape data is data using voxels as the unit elements, the correction means determines whether the cumulative value of the evaluation value exceeds the threshold using a threshold from among the thresholds set for each local space that corresponds to the position of the voxel set shown in the three-dimensional shape data in the imaging space, and the setting means divides the imaging space according to the division conditions and sets the threshold for each local space obtained by the division according to the threshold pattern. It is characterized by the following: [Effects of the Invention]

[0007] According to this disclosure, high-precision three-dimensional shape data can be obtained from three-dimensional shape data representing the general shape of an object. [Brief explanation of the drawing]

[0008] [Figure 1] A diagram showing an example configuration of an image processing system. [Figure 2] A diagram showing an example of the hardware configuration of an image processing device. [Figure 3] A diagram showing an example of the functional configuration (software configuration) of an image processing device according to Embodiment 1. [Figure 4] A flowchart showing the processing flow executed by the image processing device according to Embodiment 1. [Figure 5] A figure showing an example of an captured image. [Figure 6]Figure illustrating the acquisition of an approximate shape by the visual volume intersection method. [Figure 7] Figure showing an example of generating a distance image. [Figure 8] (a) is a diagram showing an example of a distance image of a gourd-shaped object, and (b) and (c) are diagrams showing changes in depth values before and after correction of the distance image. [Figure 9] Flowchart showing details of processing for acquiring surface three-dimensional information according to the first embodiment. [Figure 10] Figure illustrating how spatial corresponding points are derived from feature point pairs. [Figure 11] Flowchart showing details of processing for sorting surface three-dimensional information according to the first embodiment. [Figure 12] Figure illustrating sorting of spatial corresponding points. [Figure 13] Flowchart showing details of threshold setting processing according to the first embodiment. [Figure 14] (a) and (b) are diagrams showing an example of a threshold pattern. (c) is a diagram showing an example of a user interface screen for designating the number of divisions and a threshold pattern. [Figure 15] Flowchart showing details of processing for correcting an approximate shape. [Figure 16] Figure showing an example of a voting result for a voxel set. [Figure 17] Flowchart showing details of threshold setting processing according to the first modification of the first embodiment. [Figure 18] (a) to (d) are diagrams showing an example of provisional thresholds for each group. [Figure 19] Figure showing the distribution of thresholds set in an imaging space. [Figure 20] Flowchart showing details of shape correction processing according to the second modification of the first embodiment. [Figure 21] Figure showing how approximate shape data is divided. [Figure 22] Figure illustrating a problem that is a premise of the second embodiment. [Figure 23] Figure showing an example of the functional configuration (software configuration) of an image processing apparatus according to the second embodiment. [Figure 24] A flowchart showing the flow of processing executed by the image processing apparatus according to the second embodiment. [Figure 25] A flowchart showing details of processing for acquiring surface three-dimensional information according to the second embodiment. [Figure 26] Figure showing an example of a face landmark. [Figure 27] Figure showing an example of face detection in a captured image. [Figure 28] Figure showing a specific example where spatial corresponding points of a face landmark occur at a position where no human face exists. [Figure 29] A flowchart showing details of processing for sorting surface three-dimensional information according to the second embodiment. [Figure 30] (a) to (c) are diagrams showing an example of a UI screen for setting control parameters. [Figure 31] (a) and (b) are diagrams showing a specific example of face landmark sorting. [Figure 32] A flowchart showing details of processing for integrating surface three-dimensional information according to the second embodiment. [Figure 33] (a) and (b) are diagrams showing an example of a UI screen for setting control parameters. MODES FOR CARRYING OUT THE INVENTION

[0009] Hereinafter, the present embodiment will be described with reference to the drawings. Note that the following embodiments do not necessarily limit the present invention. In addition, not all combinations of features described in the present embodiment are necessarily essential to the solving means of the present invention.

[0010] [Embodiment 1] In this embodiment, surface 3D information of an object (subject) is obtained from each captured image used to generate 3D shape data representing the approximate shape of the object (subject), and distance images representing the distance from each camera to the object are corrected based on this surface 3D information. Then, based on the corrected distance images, the 3D shape data representing the approximate shape is corrected to obtain high-precision 3D shape data.

[0011] <System Configuration> Figure 1 is a diagram showing an example configuration of an image processing system according to this embodiment. The image processing system in this embodiment includes 12 cameras 101a to 101l, an image processing device 102, a user interface (UI) 103, a storage device 104, and a display device 105. Note that the 12 cameras 101a to 101l may be collectively referred to simply as "camera 101". Each camera 101a to 101l, which is an imaging device, acquires images by synchronously imaging an object 107 in the imaging space 106 from different viewpoints according to the imaging conditions. Multiple images acquired from different viewpoints in this way may be collectively referred to as "multiple viewpoint images". In this embodiment, the multiple viewpoint images are assumed to be a video consisting of multiple frames, but they may also be still images. The image processing device 102 controls the cameras 101 and generates three-dimensional shape data (hereinafter referred to as "approximate shape data") representing the approximate shape of the object 107 based on the multiple images acquired from the cameras 101. The UI 103 is a user interface for the user to perform imaging conditions and various settings, and is composed of a display with touch panel functionality, etc. The UI103 may have separate hardware buttons, and may also have a mouse or keyboard as an input device. The storage device 104 is a large-capacity storage device that receives and stores the schematic shape data generated by the image processing device 102. The display device 105 is, for example, a liquid crystal display that receives and displays the schematic shape data or corrected high-precision three-dimensional shape data generated by the image processing device 102. The imaging space 106 is shown two-dimensionally for convenience in Figure 1, but it is a rectangular space surrounded by 12 cameras 101a to 101l, and the rectangular area shown by a solid line in Figure 1 represents the contour in the front-to-back and left-to-right directions on the floor surface.

[0012] <Hardware configuration of the image processing device> Figure 2 shows an example of the hardware configuration of the image processing device 102. The image processing device 102 has a CPU 201, RAM 202, ROM 203, HDD 204, control interface (I / F) 205, input interface (I / F) 206, output interface (I / F) 207, and main bus 208. The CPU 201 is a processor that comprehensively controls each part of the image processing device 102. The RAM 202 functions as the main memory and work area of ​​the CPU 201. The ROM 203 stores a group of programs executed by the CPU 201. The HDD 204 stores applications executed by the CPU 201, data used for image processing, etc. The control (I / F) 205 is connected to each camera 101a to 101l and is an interface for setting imaging conditions, starting and stopping imaging, etc. The input I / F 206 is a serial bus interface such as SDI or HDMI (registered trademark). Through this input I / F 206, multiple viewpoint images obtained by each camera 101a to 101l performing synchronized imaging are acquired. The output I / F 207 is a serial bus interface such as USB or IEEE1394. The generated three-dimensional shape data is output to the storage device 104 and display device 105 via this output I / F 207. The main bus 208 is a transmission path that connects each module within the image processing device 102.

[0013] In this embodiment, twelve cameras 101 of the same specifications are used to image a single object from four directions: front, back, left, and right, with three cameras from each direction (101a-101c, 101d-101f, 101g-101i, 101j-101l). Three cameras 101 imaging from the same direction are positioned on a straight line perpendicular to the optical axis so that their optical axes are parallel to each other. Camera parameters (internal parameters, external parameters, distortion parameters, etc.) for each camera 101 are stored in the HDD 204. Here, internal parameters represent the coordinates of the image center and the lens focal length, while external parameters represent the camera's position and orientation. In this embodiment, twelve cameras of the same specifications are used, but the camera configuration is not limited to this. For example, the number of cameras may be increased or decreased, or the distance to the imaging space and the lens focal length may be changed depending on the imaging direction.

[0014] <Image Processing Device Functional Configuration> Figure 3 is a diagram showing an example of the functional configuration (software configuration) of the image processing device 102, and Figure 4 is a flowchart showing the processing flow by each functional unit. The series of processes shown in the flowchart of Figure 4 are realized by the CPU 201 reading the program stored in the ROM 203 or HDD 204, loading it into the RAM 202, and executing it. Below, the process by which high-precision three-dimensional shape data is created in the image processing device 102 will be explained according to the flowchart of Figure 4. In the following explanation, the symbol "S" means step.

[0015] In S401, the image acquisition unit 301 acquires multiple images from different viewpoints (multiple viewpoint images) obtained by synchronous imaging from 12 cameras 101a to 101l via the input I / F 206. Alternatively, it may acquire multiple viewpoint images stored in the HDD 204. The data of the acquired multiple viewpoint images is held in the RAM 202. Figure 5 shows images obtained by three cameras 101a to 101c, which are facing the same direction among the 12 cameras 101. Currently, the three cameras 101a to 101c are imaging the object 107 from the front. Image 501 corresponds to camera 101a, image 502 corresponds to camera 101b, and image 503 corresponds to camera 101c. When a video consisting of multiple frames is input, each frame at the same time corresponds to one of these multiple viewpoint images. That is, when a video with multiple viewpoint images is input, each process from S401 onwards will be executed on a frame-by-frame basis.

[0016] In S402, the schematic shape generation unit 302 generates three-dimensional shape data ( schematic shape data) representing the schematic shape of the object 107 as seen in the multiple viewpoint images acquired in S401. There are various formats for this three-dimensional shape data, but in this embodiment, we will explain using the case where the schematic shape data in voxel format, which represents the three-dimensional shape as a collection of tiny cubes called "voxels," is generated using the viewing volume cross-eyed method as an example. First, for each of the multiple synchronized captured images, the schematic shape generation unit 302 acquires an image representing the silhouette of the object 107 shown in the captured image (called a "silhouette image" or "foreground image") based on the difference with the background image. The background image for acquiring the silhouette image can be, for example, one obtained by capturing an image beforehand when the object 107 is not in the imaging space 106 and saved on the HDD 204 or the like. Then, based on the camera parameters of each camera 101, each voxel included in the voxel set corresponding to the imaging space 106 is projected onto the respective silhouette image. Then, in all silhouette images, only the voxels projected into the silhouette of object 107 are retained. The voxel set consisting of these remaining voxels is used as the approximate shape data of object 107. Figure 6 is a diagram illustrating how the voxel set as the approximate shape data of an object is obtained using the viewing volume cross-section method. In the case of a gourd-shaped object as shown in Figure 6, the approximate shape data obtained by the viewing volume cross-section method will have a shape like an ellipsoid that encompasses the entire object. That is, even if the actual object has a concave part, that concave part will not appear in the voxel set. Note that the method for obtaining approximate shape data is not limited to the viewing volume cross-section method. For example, approximate shape data may be obtained by applying alignment and deformation processing according to multiple viewpoint images to a basic model prepared in advance for each object. In this case, the three-dimensional posture of a person is estimated using a depth sensor, and the basic model of the person is obtained using the joint information obtained from that. Alternatively, for example, the three-dimensional coordinates of the region corresponding to the object may be obtained based on a distance image obtained separately using a low-resolution distance camera, and the solid (such as a rectangular prism) circumscribing that region may be used as the approximate shape data.

[0017] In S403, the surface three-dimensional information acquisition unit 303 acquires three-dimensional information of the surface corresponding to the contour of the object (hereinafter referred to as "surface three-dimensional information"). Specifically, it first extracts points (feature points) that characterize the object being captured from each of the captured images that make up the multiple viewpoint images. Then, it obtains the three-dimensional coordinates of the position obtained by projecting two corresponding feature points (feature point pairs) extracted from different captured images onto the imaging space 106. Hereinafter, the point representing the three-dimensional position in the imaging space corresponding to a feature point pair will be called a "spatial correspondence point". Details of this surface three-dimensional information acquisition process will be described later.

[0018] In S404, the surface 3D information selection unit 304 selects only the highly reliable surface 3D information from the surface 3D information acquired in S403, based on the approximate shape data generated in S402. In this embodiment, spatial correspondence points of numerous feature point pairs are acquired as surface 3D information, and the most reliable spatial correspondence points are selected from among them. Details of this surface 3D information selection process will be described later.

[0019] In S405, the distance image generation unit 305 generates a distance image representing the distance from each camera 101 to an object, based on the multiple viewpoint images acquired in S401. This distance image is also commonly called a "depth map". In this embodiment, the distance image is generated by stereo matching using two captured images corresponding to two adjacent cameras. Figure 7 shows how the distance image is generated based on captured images 501 to 503, corresponding to the three cameras 101a to 101c shown in Figure 5. Stereo matching methods include block matching and semi-global matching. In the block matching method, the captured image corresponding to one of two adjacent cameras is used as the reference image, and the captured image corresponding to the other camera is used as the target image. The pixel corresponding to the pixel of interest in the reference image is searched for in the target image, and the disparity at the pixel of interest is determined. Then, based on the camera parameters, the disparity at the pixel of interest is converted to depth and used as the pixel value in the distance image. By performing the above processing for each pixel of the reference image, a distance image corresponding to the reference image is obtained. Through this processing, a depth image 701 corresponding to the captured image 501 is obtained from the captured image 501 and the captured image 502, and a depth image 702 corresponding to the captured image 502 is obtained from the captured image 502 and the captured image 503. Furthermore, the depth image generation unit 305 modifies the depth images obtained as described above based on the surface three-dimensional information selected in S404. Specifically, it processes the depth values ​​of pixels on the depth image corresponding to each feature point to approach the depth values ​​calculated based on their spatial correspondence points. Figure 8(a) shows the depth image when the gourd-shaped object shown in Figure 6 exists in the imaging space, and Figures 8(b) and (c) show the change in depth values ​​before and after the modification of the depth image. As shown in Figure 8(b), the depth value calculated based on the spatial correspondence point of feature point A at the center of the recessed part of the object is far from the line profile of the pre-modification depth map. This is modified to the line profile of the modified depth map shown in Figure 8(c). In this example, the depth value of the pixel on the depth image corresponding to feature point A is replaced with a depth value calculated based on its spatially corresponding point, and then the depth values ​​of the surrounding pixels are modified to ensure a smooth transition.Such corrections allow for the correction of inaccurate depth values ​​around pixels in the depth image corresponding to a particular feature point, even if they are incorrect, by changing them to more appropriate depth values ​​based on selected three-dimensional surface information, thereby obtaining a more accurate depth image. Note that the method for generating the depth image is not limited to the stereo matching method described above. For example, a camera capable of acquiring distance information using Time of Flight (ToF) or pattern projection may be used.

[0020] In S406, the threshold setting unit 306 sets a threshold (a threshold for determining which unit elements to be deleted from the unit elements constituting the approximate shape data) for each small space (hereinafter referred to as "local space") obtained by dividing the imaging space into predetermined sizes. The larger the threshold set here, the more likely it is that unit elements constituting the approximate shape data will remain, and the stronger the tolerance to errors in the depth image. Details of this threshold setting process will be described later.

[0021] In S407, the shape correction unit 307 corrects the approximate shape data generated in S402 based on the distance image generated in S405 and the threshold set in S406. Specifically, it removes extra unit elements that are presumed not to represent the shape of the object from among the unit elements that make up the approximate shape data. Details of this shape correction process will be described later.

[0022] In S408, the output unit 308 outputs the approximate shape data corrected in S407, that is, three-dimensional shape data that more accurately represents the three-dimensional shape of the object, to the storage device 104 and the display device 105 via the output I / F 207.

[0023] In S409, a decision is made, for example, based on user instructions entered via UI103, to continue or terminate the process of generating the object's three-dimensional shape data. If the generation process is to continue, the system returns to S401 and continues the series of processes for the new multi-viewpoint images.

[0024] The above describes the process by which high-precision three-dimensional shape data is produced in the image processing system shown in Figure 1. Although the above example describes the case where there is one object in the captured image, it can also be applied when there are multiple objects. In this case, first, in S402, the voxel set obtained by the viewing volume cross-eyed method is separated into connected components, and the separated voxel set is obtained as approximate shape data for each object. Then, in S407, shape correction processing can be applied to each of the obtained approximate shape data.

[0025] <Surface 3D Information Acquisition Processing> Figure 9 is a flowchart detailing the surface three-dimensional information acquisition process (S403) according to this embodiment, performed by the surface three-dimensional information derivation unit 303. In this process, feature points extracted from each captured image are associated with different captured images, and the spatial correspondences of these feature points are acquired as surface three-dimensional information. The process will be explained in detail below following the flow chart in Figure 9.

[0026] In S901, feature points are extracted from each captured image that makes up a multi-viewpoint image. For feature point extraction, known methods such as SIFT (Scale-Invariant Feature Transform) or SURF (Speeded-Up Robust Features) can be applied. In the case of SIFT, after detecting feature points using a DoG (Difference of Gaussian) filter or the like, a process is performed to describe the feature quantity based on the orientation calculated from the gradient direction and gradient intensity.

[0027] In S902, a process is performed to associate the feature points extracted from each captured image in S901 with two captured images from different viewpoints. In this embodiment, for each combination of captured images corresponding to two different cameras, a process is performed to associate each feature point extracted from one captured image with the feature point extracted from the other captured image that has the smallest distance between the feature points. In this way, combinations of corresponding feature points (hereinafter referred to as "feature point pairs") are determined. Here, the combinations of captured images for which feature point association is performed may be predetermined based on camera parameters. For example, pairs of cameras whose distance from each other is within a predetermined range and whose difference in optical axis (camera orientation) is within a predetermined range may be determined in advance, and feature points may be associated with the captured images obtained by those camera pairs.

[0028] In S903, the spatial correspondence point described above is derived for each feature point pair obtained in S902. Specifically, based on the camera parameters of both cameras that captured the two images from which the feature points related to the target feature point pair were extracted, two rays corresponding to the feature point are determined, and their intersection is determined as the spatial correspondence point. If the two rays do not intersect, the midpoint of the line segment with the shortest distance between the two rays may be determined as the spatial correspondence point. Furthermore, if the distance between the two rays is greater than a predetermined value, it may be determined that the correspondence for that feature point pair was incorrect, and the feature point may be excluded from the derivation of the spatial correspondence point. Figure 10 is a diagram illustrating how spatial correspondence points are derived from feature point pairs. Figure 10 shows three spatial correspondence points p1 to p2. Spatial correspondence point p1 is the spatial correspondence point derived from the feature point pair of feature point f0 in the image captured by camera 101b and feature point f1 in the image captured by camera 101c. The spatial correspondence point p2 is derived from the feature point pair of feature point f0 in the image captured by camera 101b and feature point f2 in the image captured by camera 101a. The spatial correspondence point p3 is derived from the feature point pair of feature point f0 in the image captured by camera 101b and feature point f3 in the image captured by camera 101d. In this way, in this step, the spatial correspondence point is derived for each feature point pair obtained in S902, and the spatial correspondence point information for each feature point pair is acquired as surface three-dimensional information.

[0029] The above describes the surface three-dimensional information derivation process according to this embodiment.

[0030] <Surface 3D Information Sorting Process> Figure 11 is a flowchart detailing the process (S404) performed by the surface 3D information selection unit 304 to select more reliable surface 3D information from the surface 3D information derivation unit 303. In this embodiment, the surface 3D information selection process removes all but the spatial correspondence points derived for each feature point pair, keeping only those spatial correspondence points that are close to the surface of the 3D shape indicated by the schematic shape data. The surface 3D information selection process will now be described in detail following the flow shown in Figure 11.

[0031] In S1101, the surface (contour) of the three-dimensional shape is extracted based on the schematic shape data generated in S402. Hereafter, the shape surface extracted from the schematic shape data will be referred to as the " schematic shape surface". If the schematic shape data is in voxel format, the voxels adjacent to the background are identified from the voxel set representing the three-dimensional shape of the object, and the set of these identified voxels is extracted as the schematic shape surface. If the schematic shape data is in point cloud format, the process is the same as for the voxel format, and the set of points adjacent to the background should be extracted as the schematic shape surface. If the data is in mesh format, each polygon face constituting the mesh should be extracted as the schematic shape surface.

[0032] In S1102, the distance to the approximate shape surface extracted in S1101 is calculated for each spatially corresponding point acquired for each feature point pair in the flow chart of Figure 9. In this embodiment, where the approximate shape surface is represented by a background and a set of adjacent voxels, the distance from the three-dimensional position of the spatially corresponding point to all voxels included in that set of voxels is determined, and the minimum distance among them is taken as the distance from the spatially corresponding point to the approximate shape surface. Here, "distance to voxel" is the distance to the three-dimensional coordinate of the voxel center, and is expressed in units such as mm. This process is performed for all spatially corresponding points from which three-dimensional coordinates have been derived.

[0033] In S1103, based on the distances calculated for each spatial corresponding point in S1102, only highly reliable spatial corresponding points are retained, and other spatial corresponding points are removed. Specifically, processing is performed in which only spatial corresponding points whose distance to the rough shape surface is equal to or less than a predetermined distance are retained, and spatial corresponding points whose distance to the rough shape surface is greater than the predetermined distance are deleted. Here, the predetermined distance is defined in units of voxel resolution, for example, as "n × voxel resolution (n is a constant)", and is set in advance based on how much thickness of correction is desired for the rough shape surface. FIG. 12 is a diagram illustrating a state in which the sorting according to the present embodiment is performed on the three spatial corresponding points p1 to p3 shown in FIG. 10 described above. Here, let d1, d2, and d3 be the distances calculated for p1, p2, and p3 respectively. If p1 is assumed to be on the rough shape surface, 0 = d1 < d3 < d2 holds. Here, when the predetermined distance is smaller than d3, among the three spatial corresponding points p1 to p3, only the spatial corresponding point p1 whose distance to the rough shape surface is zero is retained, and the spatial corresponding points p2 and p3 are deleted.

[0034] The above is the content of the surface three-dimensional information sorting process according to the present embodiment. By this process, spatial corresponding points derived from incorrectly correlated feature point pairs, that is, spatial corresponding points with low reliability can be removed, and only spatial corresponding points with high reliability can be retained.

[0035] <Threshold Setting Processing> FIG. 13 is a flowchart showing details of threshold setting processing (S406) executed by a threshold setting unit 306. In the threshold setting processing of the present embodiment, a threshold is set for each local space obtained by dividing an imaging space according to a predetermined division condition, based on a prepared threshold pattern. Hereinafter, the threshold setting processing will be described in detail along the flow of FIG. 13.

[0036] In S1301, the imaging space is divided into multiple local spaces according to predetermined division conditions. In this embodiment, the imaging space is divided into equal intervals in the front-to-back and left-to-right directions according to a predetermined number of divisions, and each is divided into a small rectangular spatial unit. Hereinafter, each small space obtained by the division will be referred to as a "local space". Note that the above division method is just an example and is not limited thereto. For example, the division may be made so that the intervals become smaller closer to the center of the imaging area, rather than at equal intervals. Also, the local spaces may be divided so that their shape is a tetrahedron or other shape.

[0037] In S1302, a threshold is set for each local space divided in S1301 based on a predetermined threshold pattern. In this embodiment, for example, a threshold pattern is used to set a threshold value for each local space, such as the one shown in Figures 14(a) and (b), which is designed so that a larger threshold value is set for local spaces closer to the center of the imaging space. The content of the threshold pattern is arbitrary; for example, a threshold pattern may be created in which the threshold is different for local spaces containing objects without recessed areas and local spaces containing objects with recessed areas. Alternatively, previously created threshold patterns may be saved so that the user can select and specify from them.

[0038] The above describes the threshold setting process. Alternatively, instead of using a predetermined number of divisions and threshold patterns, the user may specify the number of divisions and threshold patterns each time via a user interface screen (UI screen) as shown in Figure 14(c). Or, it may be possible to specify an arbitrary threshold for each local space. Note that the UI screen in Figure 14(c) shows the two-dimensional number of divisions and threshold patterns when the imaging space is viewed from directly above, and the height direction (Z-axis direction) is common.

[0039] <Shape correction processing> Figure 15 is a flowchart detailing the process (S407) performed by the shape correction unit 307 to correct the approximate shape data generated in S402. In this embodiment, the approximate shape data is evaluated based on a distance image, and based on the results of the evaluation, a process is performed to delete unit elements that are presumed not to represent the shape of an object among the unit elements constituting the approximate shape data. The following will be explained in detail following the flow in Figure 15.

[0040] In S1501, a threshold is set for the approximate shape data to determine whether a voxel is to be deleted. Specifically, first, the centroid coordinates of the voxel set representing the approximate shape are calculated. Then, a local space containing the calculated centroid coordinates is identified, and the threshold set in the aforementioned threshold setting process is set for the identified local space as the threshold to be applied to the approximate shape data. For example, if thresholds are set for each local space according to the threshold patterns shown in Figures 14(a) and (b) above, a threshold of "2" will be set for the approximate shape data of an object located in the center of the imaging space. Similarly, a threshold of "1" will be set for the approximate shape data of an object located at the edge of the imaging space.

[0041] In S1502, an evaluation is performed on each voxel that constitutes the voxel set representing the approximate shape, based on the depth image generated in S405. This evaluation is performed by voting for voxels that are deemed unnecessary. Since depth images are generated for each camera 101, all generated depth images are processed sequentially. Specifically, a voxel of interest is sequentially selected from the voxel set, and the depth value at the pixel position on the depth image corresponding to the voxel of interest is compared with the depth value from the camera corresponding to the depth image to the voxel of interest. If the latter depth value is smaller, one vote is cast for the voxel of interest. This is equivalent to adding "1" to the evaluation value. As a result, in each voxel that constitutes the voxel set representing the approximate shape, the more likely it is that a voxel does not represent the original object shape, the larger its number of votes (cumulative evaluation value). Here, the following equation (1) is used to compare the depth values.

[0042]

number

[0043] In the above equation (1), D * vi represents the depth value from the voxel center v to the camera corresponding to the depth image i. Di(x,y) represents the depth value of the pixel position in the depth image i specified by coordinates (x,y). (xvi,yvi) is the coordinate that indicates the pixel position when the voxel center v is projected onto the depth image i. In this case, the "depth value of the pixel position on the depth image corresponding to the voxel of interest" can be obtained by the following procedure. First, based on the camera parameters of the camera corresponding to the depth image i, the voxel center v of the voxel of interest is projected onto the depth image to obtain the coordinates (xvi,yvi) on the depth image i corresponding to the voxel of interest. Next, the depth value at coordinates (xvi,yvi) in the depth image i is obtained if there is a pixel at the corresponding position, or if there is no pixel at the corresponding position, it is obtained by interpolation (nearest neighbor interpolation, etc.) of the depth values ​​of surrounding pixels. The value obtained in this way is the depth value of the pixel position on the depth image corresponding to the voxel of interest. The "depth value from the camera corresponding to the depth image to the voxel of interest" can be obtained using the following procedure. First, based on the camera parameters of the camera corresponding to the depth image i, the voxel center v of the voxel of interest is transformed into a coordinate system based on the camera corresponding to the depth image i. Next, the depth (ignoring front, back, left, and right) to the transformed voxel center v is calculated. The value obtained in this way is the depth value from the camera corresponding to the depth image to the voxel of interest.

[0044] Then, if the voxel of interest satisfies the conditions of equation (1) above, one vote (evaluation value "1") is added to that voxel of interest. As a result of this process, if the depth values ​​in all distance images are accurate (i.e., no distance images contain incorrect depth values), the number of votes for the voxel representing the original object shape will be "0". If, for example, there is only one distance image among the distance images corresponding to each camera that contains an incorrect depth value, the number of votes for the voxel representing the original object shape will be "1". Figure 16 shows an example of the voting results for a set of voxels representing the approximate shape of the gourd-shaped object shown in Figure 6. As a result of a distance image containing an incorrect depth value for the voxel corresponding to the constricted part in the center of the object, one vote was also cast for each of the four voxels 1600 that should not be deleted. In this embodiment, 12 distance images are obtained, each corresponding to one of the 12 cameras 101, so the above process is repeated 12 times.

[0045] In S1503, based on the voting results (evaluation results), voxels with a number of votes (= cumulative evaluation value) exceeding the threshold set in S1501 are removed from the voxel set representing the general shape. Here, we will explain the removal of voxels based on the voting results by referring to Figure 16 above. In Figure 16, the thick line 1601 shows the outline of the voxel set representing the general shape before correction. If a threshold of "2" was set for the general shape data, voxels with a number of votes of "2" or more will be removed, and voxels with a number of votes of "1" or less will remain. As a result, the voxel set will be corrected to have an outline as shown by the dashed line 1602.

[0046] The above describes the shape correction process. Note that if the rough shape data is in point cloud format, the above explanation can be applied by replacing "voxel" with "point," but it cannot be applied directly if it is in mesh format. If the rough shape data is given in mesh format, the data format must be converted to replace the area enclosed by the mesh with a set of voxels, and then the flow shown in Figure 15 above must be applied. If you want to output the corrected shape data in the original mesh format, you should convert the data format back to mesh format and output it.

[0047] As described above, according to this embodiment, three-dimensional surface information of an object is obtained from each captured image used to generate the object's schematic shape data, and the depth image is corrected based on this three-dimensional surface information. Then, by correcting the schematic shape data based on the corrected depth image, the three-dimensional shape of an object with a complex shape including indentations can be accurately restored.

[0048] <Example 1> In the threshold setting process described above, the imaging space is divided into a predetermined number of divisions, and a threshold is set for the local space based on a pre-prepared threshold pattern. However, the method of setting the threshold is not limited to this. For example, the depth images may be grouped, and a threshold may be set for the local space based on the visibility of the depth images belonging to each group to each local space. Figure 17 is a flowchart detailing the threshold setting process related to this modified example. The threshold setting process of this modified example will be described below in accordance with the flow shown in Figure 17.

[0049] S1701 is the same as S1301 described above, and the imaging space is divided into multiple local spaces according to predetermined division conditions. In the following S1702, the distance images corresponding to each of the multiple cameras are divided based on the imaging direction specified by the camera parameters, so that distance images corresponding to cameras with a common imaging direction are grouped together. Here, they are divided into four groups, from the 1st group to the 4th group. Note that the above grouping is just one example, and for example, distance images with similar position and orientation indicated by the camera parameters may be grouped together.

[0050] In S1703, for each group, the number of distance images that are visible to the local space of interest out of all local spaces is counted. Here, a "visible distance image" refers to a distance image in which the local space of interest is contained within its field of view, and will be referred to as a "visible distance image" below.

[0051] In S1704, a provisional threshold for the local space of interest is determined for each group based on the number of visible-distance images obtained for each group. Here, the provisional threshold is set to a value smaller than the count of visible-distance images. Figures 18(a) to (d) show the provisional thresholds for the local space 1801 determined for each directional group (Group 1 to Group 4) when six cameras are arranged in four directions at 90-degree intervals surrounding the imaging space 1800. In this example, in each group, five distance images are obtained by performing stereo matching between images from adjacent cameras using six images captured by the six cameras. Next, the number of visible-distance images is counted for the five distance images. Then, in each group, if the local space of interest belongs to a region where the number of visible-distance images is "2", a provisional threshold of "1" is determined, and if it belongs to a region where the number of visible-distance images is "3" or more, a provisional threshold of "2" is determined.

[0052] In S1705, the minimum value among the provisional thresholds determined in S1704 on a group basis is set as the threshold for the local space of interest. Figure 19 shows the distribution of thresholds set based on the provisional thresholds on a group basis shown in Figure 18 above. The height direction is omitted, but it is assumed that the same threshold is set. In Figure 19, it can be seen that a threshold of "2" is set for local spaces belonging to the dark gray area within the imaging space 1800, and a threshold of "1" is set for local spaces belonging to the light gray area.

[0053] As described above, thresholds may be set for each local space based on the grouped distance images. In this modified example, the distance images were divided into four groups, but the number of groups to be divided is not limited to this. Also, in this modified example, the distance images were divided so that the groups are mutually exclusive, but they may also be divided so that the groups overlap.

[0054] <Modification 2> In the embodiment described above, the minimum value among the provisional thresholds determined at the group level was set as the threshold for the local space, and only one threshold was set for each local space. However, the provisional thresholds determined at the group level may be set directly as the thresholds for the local space. In this case, the approximate shape data can be divided according to the imaging direction, and shape correction can be performed by applying each of the multiple thresholds to the divided shape data. The shape correction process related to this modified example will be explained below in accordance with the flowchart shown in Figure 20.

[0055] In S2001, the schematic shape data is divided according to the aforementioned groups. In the example above, where the data is divided into four groups according to the imaging direction, it is sufficient to divide it into four parts based on the faces passing through the vertices of the bounding box that encompasses the voxel set as schematic shape data. Figure 21 shows how schematic shape data with a contour as shown in Figure 12 is divided according to the aforementioned four groups. The data representing a part of the schematic shape data obtained by the division will be called "partial shape data". In the next S2002, a threshold is set for each piece of partial shape data based on a provisional threshold corresponding to each group. Specifically, the centroid coordinates of the voxel set as partial shape data are found, and the provisional threshold for each group determined for the local space to which these centroid coordinates belong is set as the threshold. Now, as shown in 18 above, let's assume that provisional thresholds for the local space 1801 have been set for each group in the first to fourth directions. In this case, as shown in the example in Figure 21 above, a threshold of "2" is set for the sub-shape data corresponding to the first, third, and fourth directions, and a threshold of "1" is set for the sub-shape data corresponding to the second direction. The following S2003 corresponds to S1502 in the flow of Figure 15 above. That is, for each set of voxels as sub-shape data, a vote is made for each voxel based on the distance image generated in S405. Note that the distance image referenced during voting may be limited to those belonging to the same group as the target sub-shape data group. S2004 corresponds to S1503 in the flow of Figure 15 above. That is, for each set of voxels as sub-shape data, based on the voting results, voxels that have received more votes than the threshold set in S2002 are removed from the voxel set.

[0056] The above describes the shape correction process related to this modified example. This modified example also allows for the accurate restoration of the object's three-dimensional shape.

[0057] <Other variations> In the above-described embodiment, the three-dimensional surface information after sorting was used to correct the distance image, but the method of use is not limited to this. For example, it may be used to set the search range when identifying pixels in the target image that correspond to pixels of interest in the reference image. Specifically, the search range near the feature points in the reference image is set to be narrower based on the three-dimensional coordinates of the spatially corresponding points of the feature point pairs. This is because the search range corresponds to the range in which an object may exist, and by using the spatially corresponding points of the feature point pairs, which are the three-dimensional surface information of the object, an appropriate search range can be set.

[0058] Furthermore, although thresholds are set for each local space in the above embodiment, if there are multiple objects in the imaging space, different thresholds may be set for each object. For example, a larger threshold for a person (player) than for a ball, or a smaller threshold for objects with simpler shapes. Alternatively, if there are multiple people in the imaging space, different thresholds may be used for people who require correction and those who do not. When setting different thresholds for each object, objects can be identified in S1501 by template matching or the like, and predetermined thresholds can be set for each object.

[0059] Furthermore, in the above embodiment, a single threshold is set for the entire schematic shape data, but different thresholds may be set for each part of the three-dimensional shape represented by the schematic shape data (for example, for each part such as the head, arms, torso, and legs in the case of a human object). In this case, first, the voxel set representing the schematic shape is divided into multiple voxel sets corresponding to each part (part-level schematic shape data). Then, the centroid coordinates for each part-level schematic shape data are identified, and the threshold corresponding to the local space containing those centroid coordinates can be set as the threshold for the part-level schematic shape data.

[0060] Furthermore, in the above embodiment, the approximate shape data is corrected according to the number of votes for each voxel based on the distance image, but weighting may also be applied to each distance image. For example, if the distance resolution differs for each distance image, a larger weight may be set for distance images with higher distance resolution. In this way, distance images with higher distance resolution will be reflected more in the evaluation results. Alternatively, the weight may be reduced for regions of the approximate shape represented by the voxel set that should not be corrected, or the number of votes may be reduced. For example, by reducing the weight for voxels whose distance to the surface of the approximate shape is greater than a predetermined value, it becomes more difficult to delete those voxels. In this way, the contribution rate to the evaluation results may be controlled by weighting each distance image and approximate shape data.

[0061] Furthermore, in the above-described embodiment, the distance image is corrected based on the spatially corresponding points of the feature point pairs, and the schematic shape data is corrected based on the corrected distance image. However, the schematic shape data may also be corrected based on the distance image before correction. In this case, the processing by the surface three-dimensional information derivation unit 303 and the surface three-dimensional information sorting unit 304 is skipped.

[0062] Furthermore, in the above embodiment, a process is performed to compare the number of votes for each voxel constituting the voxel set representing the general shape with a set threshold to determine whether to delete a voxel on a voxel-by-voxel basis. However, if a common threshold of "1" is set for all local spaces, the determination process by threshold comparison becomes unnecessary. In other words, voxels that satisfy equation (1) above in any of the depth images can be deleted immediately. This makes it possible to perform shape correction processing more simply.

[0063] [Embodiment 2] As mentioned above, Embodiment 1 described above is also applicable when multiple objects are visible. Here, we consider a case where multiple objects of the same type exist in the input multi-viewpoint images (for example, multiple people are shown side by side). In such a case, in the surface 3D information derivation process, two or more points corresponding to the facial features that characterize each person's face, such as the eyes, nose, and mouth (generally called "face landmarks"), will be extracted as feature points. As a result of extracting face landmarks for each of the multiple people from each captured image, a large number of incorrect spatial correspondence points based on incorrect combinations of face landmarks are generated. Figure 22 illustrates how a huge number of "spatial correspondence points for the right eye" are generated, including an incorrect 3D position where the right eye does not actually exist. The same thing happens with other facial organs such as the nose and mouth, resulting in a large amount of incorrect surface 3D information that represents faces of a normal size based on a large number of incorrect combinations of face landmarks. Therefore, in this embodiment, we will describe a method for appropriately obtaining surface 3D information corresponding to each face when two or more people are visible in a multi-viewpoint image. The following explanation will focus on the differences from Embodiment 1.

[0064] <Image Processing Device Functional Configuration> Figure 23 is a diagram showing an example of the functional configuration (software configuration) of the image processing apparatus 102 according to this embodiment, and Figure 24 is a flowchart showing the processing flow by each functional unit. The main difference from Embodiment 1 is that the surface three-dimensional information integration unit 2301 is added in Figure 23, and the surface three-dimensional information integration process (S2401) is added in Figure 24. However, this is not the only difference from Embodiment 1; the contents of the surface three-dimensional information derivation process (S403) and the surface three-dimensional information selection process (S404) are also different. The processes of surface three-dimensional information derivation, selection, and integration in this embodiment will be described in detail below.

[0065] <Surface 3D Information Derivation Processing> Figure 25 is a flowchart detailing the surface three-dimensional information derivation process (S403) according to this embodiment. The following explanation will follow the flow shown in Figure 25.

[0066] In S2501, two or more feature points are extracted from each captured image for each object of the same type. In this embodiment, the face landmarks of each of the multiple people depicted in each captured image are detected and extracted as feature points. For detecting face landmarks, known face recognition techniques such as Dlib or OpenCV can be used. Here, as shown in Figure 26, a total of seven face landmarks are detected and extracted as feature points: the outer corner of the right eye 2601, the inner corner of the right eye 2602, the outer corner of the left eye 2603, the inner corner of the left eye 2604, the tip of the nose 2605, the right corner of the mouth 2606, and the left corner of the mouth 2607. Note that the face landmarks to be extracted are not limited to the above seven. For example, it is not necessary to include any of the above seven, or it may include more facial parts such as the space between the eyebrows, points on the cheeks, and points on the jawline.

[0067] In S2502, a process is performed to associate two or more feature points for each object of the same type extracted from each captured image with the two captured images from different viewpoints. This determines the combination of feature point groups that correspond to each of the above-mentioned objects between the captured images. The "combination of feature point groups" determined here corresponds to the "feature point pair" in Embodiment 1. However, if multiple objects of the same type are captured in each captured image, the "combination of feature point groups" determined may not be for the same object captured in the two captured images. This will be explained using a specific example. Figure 27 shows images 2700, 2710, and 2720 obtained by capturing two people 107a and 107b with three cameras 101a to 101c. Now, person 107a is captured in images 2700 and 2710, and person 107b is captured in all three captured images 2700, 2710, and 2720. Then, for each of the face portions of persons 107a and 107b detected by extracting seven face landmarks from each captured image, the bounding rectangle face frames 2701, 2702, 2711, 2712, and 2721 are shown. In this example, there are eight possible combinations of "faces (≒seven face landmarks)" as follows. • F1 (Face frame 2701 and Face frame 2711): Incorrect response • F2 (Face frame 2701 and Face frame 2712): Correctly compatible • F3 (Face frame 2701 and Face frame 2721): Correctly compatible • F4 (Face frame 2702 and Face frame 2711): Correctly compatible • F5 (Face frame 2702 and Face frame 2712): Incorrect response • F6 (Face frame 2702 and Face frame 2721): Incorrect response • F7 (Face frame 2711 and Face frame 2721): Incorrect response • F8 (Face frame 2712 and Face frame 2721): Correctly compatible

[0068] As described above, eight possible combinations of feature point clouds are obtained, but some of them contain incorrect correspondences between faces (mismatches). Therefore, the "combinations of feature point clouds" determined in this embodiment will be referred to as "candidate feature point pairs" below. In addition, "combinations of faces" that use face landmarks as feature point clouds will be referred to as "candidate faces."

[0069] In S2503, for each of the feature point pair candidates determined in S2502, the spatially corresponding points of the feature point group are derived. In this embodiment, one face candidate contains seven face landmarks. Therefore, based on the camera parameters of the cameras corresponding to the two captured images related to the face candidate of interest, the intersection of the two corresponding light rays is determined as the spatially corresponding point for each face landmark. In this way, seven spatially corresponding points for each face candidate are derived as surface three-dimensional information.

[0070] The above describes the surface three-dimensional information derivation process according to this embodiment.

[0071] <Surface 3D Information Sorting Process> In the process of determining candidate feature point pairs (face candidates in this embodiment) in the surface three-dimensional information derivation process described above, the process of matching which person in the captured image they belong to is not performed. Therefore, as mentioned above, the face candidates include mismatched faces that combine the faces of different people. As a result, among the spatial correspondence points of face landmarks for each face candidate derived as surface three-dimensional information, there are some that indicate three-dimensional locations where no human face actually exists. A specific example is shown in Figure 28. Figure 28 is a view of the imaging space of Figure 27 from directly above, showing the head 2800a of person 107a and the head 2800b of person 107b. In Figure 27, the face detected for person 107a was shown as face frame 2702 in captured image 2700 and as face frame 2711 in captured image 2710. Similarly, the face detected for person 107b was shown as face frame 2701 in captured image 2700 and as face frame 2712 in captured image 2710. Point 2801 in Figure 28 represents the spatial correspondence point of the right outer corner of the eye 2601 derived from face candidate F1 (face frame 2701 and face frame 2711), which is a combination of faces of different people. It can be seen that this indicates the three-dimensional position of the right outer corner of the eye, which does not actually exist. Thus, the spatial correspondence points of face landmarks derived from face candidates where people are mismatched are highly likely to indicate a location where the person does not actually exist. Therefore, from the spatial correspondence points of the derived face landmarks, we select only those whose positions are close to the approximate shape surface obtained from the approximate shape data.

[0072] Figure 29 is a flowchart detailing the surface three-dimensional information selection process (S404) according to this embodiment. When acquiring schematic shape data by the view volume cross-section method, the true shape does not exist outside the schematic shape surface, provided there are no errors in the silhouette image used. On the other hand, recessed parts of an object cannot be reproduced, and areas that are not the true shape may be included in the voxel set if part of the object is occluded or if there is an insufficient number of silhouette images (≒ number of viewpoints). Therefore, the spatial correspondence points indicating the three-dimensional position outside the schematic shape obtained by the view volume cross-section method are likely to be incorrect. In this embodiment, only the spatial correspondence points of the derived face landmarks that indicate the three-dimensional position inside the voxel set representing the schematic shape are retained. The following explanation follows the flow shown in Figure 29.

[0073] In S2901, a face candidate to be focused on for processing is selected from all face candidates. In the following S2902, one face landmark to be focused on for processing is selected from the face landmarks that are feature points. In this embodiment, one face landmark is selected at a time from among the seven face landmarks.

[0074] In S2903, the next process to be executed is determined by whether or not the spatially corresponding points of the face landmark of interest, set in S2902, are contained within the outline shape. That is, if the spatially corresponding points of the face landmark of interest are contained within the voxel set representing the outline shape, the process in S2904 is executed next; otherwise, the process in S2907 is executed next.

[0075] In S2904, the next process to be executed is determined based on whether all voxels within a radius N [mm] centered on the spatially corresponding point of the face landmark of interest are included inside the schematic shape. Here, N is a control parameter, and ideally, the maximum value of the difference between the "spatial shape surface obtained based on the viewing volume cross-section method" and the "true position of the face landmark" is used, and this is determined taking into account the number of viewpoints, etc. Figure 30(a) shows an example of a user interface screen (UI screen) for the user to set the control parameter N. In the image display area 3001 of the UI screen shown in Figure 30(a), the schematic shapes of two people are shown as gray silhouettes, and the derived face candidates are shown as dashed lines. The user can operate the seek bar 3002 in the UI screen to specify any value for the control parameter N, which represents the distance from the schematic shape, within the range of 0 to 40 [mm], and here N is set to 20 [mm]. If all voxels within a circle of radius N mm centered on the spatially corresponding point of the face landmark of interest are not contained within the schematic shape (i.e., the spatially corresponding point is within N mm of the schematic surface shape), then process S2905 is executed next. On the other hand, if all voxels within a circle of radius N mm centered on the spatially corresponding point of the face landmark of interest are contained within the schematic shape (i.e., the spatially corresponding point is not within N mm of the schematic surface shape), then process S2907 is executed next.

[0076] In S2905, the next process to be executed is determined by whether or not processing has been completed for all face landmarks included in the candidate faces of interest. If all face landmarks have been processed, the process in S2906 is executed next. On the other hand, if there are any face landmarks that have not been processed, the process returns to S2902, the next face landmark of interest is set, and processing continues.

[0077] In S2906, the spatial correspondence points of all face landmarks included in the candidate faces are added to the list. The resulting list will contain a set of spatial correspondence points for each candidate face that are presumed to correctly represent the surface shape of a person's face in the imaging space.

[0078] In S2907, the next process to be executed is determined based on whether or not processing for all face candidates has been completed. If there are any unprocessed face candidates, the process returns to S2901, the next face candidate of interest is set, and processing continues. On the other hand, if processing for all face candidates has been completed, this process is terminated. Through this process, only face candidates in which it has been determined that the spatially corresponding points of face landmarks are located inside the general shape and within a certain distance from the surface of the general shape can be selected. Figures 31(a) and (b) show specific examples of selection using the above flow. In Figures 31(a) and (b), the dashed-dot curve 3100 represents the general shape surface of a person's face, and the solid circle 3101 represents a circle with radius N mm centered on the spatially corresponding point of the face landmark. Figure 31(a) is an example of the three-dimensional surface information remaining after selection, i.e., the group of spatially corresponding points for each face candidate listed in the aforementioned list. In this example, it can be seen that the spatial correspondence points 2601' to 2607' of all face landmarks for the face candidates are located inside the schematic shape surface 3100 and within N mm of the schematic shape surface 3100. Figure 31(b) is an example of surface three-dimensional information that is not retained by the selection process, i.e., a group of spatial correspondence points for face candidates that are not included in the aforementioned list. In this example, it can be seen that among the spatial correspondence points 2601" to 2607" of all face landmarks for the face candidates, the spatial correspondence point 2601" of the outer corner of the right eye is located outside the schematic shape surface 3100, and the two spatial correspondence points 2603" and 2604" of the inner and outer corners of the left eye are not within 30 mm of the schematic shape surface 3100.

[0079] The above describes the surface three-dimensional information selection process according to this embodiment. This process excludes the spatial correspondence points of incorrect face landmarks derived from face candidates that mismatch between people, thereby obtaining highly accurate surface three-dimensional information corresponding to the faces of real people. In the image display area 3001 of the UI screen in Figure 30(a), the face candidates remaining after selection are shown with solid lines, and the selection results (number of face candidates before and after selection) are shown on the right side of the screen. In the flow in Figure 29, only face landmarks with spatial correspondence points inside the schematic shape are retained, but considering the errors that the silhouette image may contain, it is also possible to allow the retention of face landmarks with spatial correspondence points outside the schematic shape. In this case, for example, a threshold is set separately to determine how far outside the surface of the schematic shape is to be allowed. It is preferable that the external threshold in this case be a smaller value than the internal threshold, for example, 5 [mm]. Furthermore, this external threshold may also be specified by the user in the UI screen shown in Figure 30(a) above, similar to the control parameter N that defines the internal threshold. Furthermore, the detection accuracy of face landmarks varies depending on the area; for example, the outer corners of the eyes and corners of the mouth have strong image features and are easy to detect accurately, while accurately detecting the position of the tip of the nose is difficult. Therefore, for example, the above selection could be performed using only the spatially corresponding points of the six face landmarks excluding the tip of the nose: the inner corners of both eyes, the outer corners of both eyes, and the corners of the mouth.

[0080] <Surface 3D Information Integration Processing> Figure 32 is a flowchart detailing the surface three-dimensional information integration process (S2401) according to this embodiment, performed by the surface three-dimensional information integration unit 2301. This process integrates the selected feature point pairs so that there is one for each object (in this case, one for each face candidate). The following explanation follows the flow shown in Figure 32.

[0081] In S3201, the position and orientation of each face candidate remaining after selection are derived in the imaging space. First, the average value of the three-dimensional coordinates of the spatially corresponding points is calculated for each of the multiple (seven in this embodiment) face landmarks that make up the face candidate. Then, the three-dimensional position of the face candidate unit, which is identified by the average value of the three-dimensional coordinates calculated for all face landmarks, is determined as the position of the face in the imaging space. In this case, for example, a face landmark with low accuracy (e.g., the tip of the nose) among the seven face landmarks may be excluded from the calculation of the average value. Next, the forward direction of the face is determined by finding the normal of the triangle formed by the midpoints of both outer corners of the eyes and both corners of the mouth, and the rightward direction of the face is determined by finding the direction vector from the outer corner of the left eye to the outer corner of the right eye, thereby determining the orientation of the face. In this way, the position and orientation of the face are derived for each face candidate.

[0082] In S3202, based on the "face position and orientation" derived for each face candidate, those with similar "face position and orientation" are merged. Here, the criteria for "similar" are as follows: First, regarding face position, the condition is that the distance between them is M [mm] or less. Second, regarding face orientation, the condition is that the angle θf between the forward directions and the angle θr between the right directions are both θt or less. Here, M and θt are control parameters, which are set by the user via the UI screen shown in Figure 30(b), for example. The UI screen in Figure 30(b) is displayed by pressing the "Advanced Settings" button 3004 in the UI screen in Figure 30(a) mentioned above. Here, the width of a typical human face is about 160 [mm], and it is considered abnormal for the distance between two faces to be 100 [mm] or less under the constraint that the difference in face orientation is, for example, 30° or less. Therefore, in the UI screen of Figure 30(c), the maximum value of the distance M between face candidates is set to 200 [mm], and the maximum value of the angle θf formed by the face pose is set to 40 [°]. However, if the subject being imaged is a child with a small face, M should be set to a smaller value, or if the accuracy of face pose estimation is low, θt should be set to a larger value, and these settings should be adjusted according to the situation. Control parameters M and θt are set via such a UI screen. When two or more "face positions and poses" that satisfy the conditions defined by the above control parameters M and θt are identified, they are judged to represent the "face position and pose" of the same person and are merged. Specifically, for each face landmark of the two or more face candidates determined to be merged, the median of the three-dimensional coordinates of its spatially corresponding point is calculated. The three-dimensional coordinates indicated by the calculated median are then used as the three-dimensional coordinates of the spatially corresponding point of each face landmark in the merged face candidate. Note that the above method of adopting the median when merging multiple three-dimensional coordinates is just one example; the mean, mode, or the midpoint between the maximum and minimum values ​​may also be adopted. Furthermore, the closest or furthest three-dimensional coordinates from the approximate surface shape may be used. In this way, the three-dimensional surface information (in this case, the spatial correspondence points of seven face landmarks) that are thought to represent the faces of the same person are integrated, and three-dimensional surface information that corresponds one-to-one with each person's face present in the imaging space is obtained.Figure 30(c) shows the image display area 3001 of the UI screen after the integration process, and it can be seen that the two selected face candidates for the person on the left have been integrated into one. The integration result is also reflected in the "Number of remaining face candidates" on the right side of the screen, which has changed from "3" to "2".

[0083] In S3203, the position and orientation of one face per person are derived for each person present in the imaging space, based on the three-dimensional coordinates of the spatially corresponding points of each face landmark in the integrated face candidate unit. The same method used in S3101 above can be used for this derivation.

[0084] The above describes the surface three-dimensional information integration process. In this way, surface three-dimensional information of the face can be obtained that corresponds one-to-one with the person captured in each image of the multiple viewpoint images. In this embodiment, a person's face was used as an example, but the method is not limited to this, and may be applied to other parts of the body (for example, arms or legs), or even to objects other than people, such as the tires of a car or motorcycle.

[0085] <Variation> In the embodiments described above, an example was given in which the control parameters M and θt in the surface three-dimensional information integration process are set based on user operation, but the method of setting the control parameters is not limited to this. For example, the user may be allowed to specify the number of people present in the target scene via a UI screen as shown in Figures 33(a) and (b), and the control parameters M and θt may be automatically determined so that the number of faces after integration is less than or equal to the number of people specified. In this case, for example, the control parameter M can be determined by decreasing it from 40[mm] in increments of 1[mm] until the number of candidate faces after integration reaches 2. The control parameters may be set in this manner.

[0086] [Other embodiments] This disclosure can also be implemented by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions.

[0087] Furthermore, the disclosure of this embodiment includes the following configurations and methods.

[0088] (Composition 1) A first acquisition means for acquiring three-dimensional shape data of an object present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition means for acquiring multiple distance images corresponding to multiple viewpoints, Correction means for correcting the three-dimensional shape data by evaluating the three-dimensional shape data based on the plurality of distance images and, based on the results of the evaluation, deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data, An image processing apparatus characterized by having

[0089] (Configuration 2) The correction means is From among the unit elements constituting the three-dimensional shape data, the unit element of interest is determined in order. If the distance from the viewpoint corresponding to the distance image to the unit element of interest is smaller than the distance to the pixel position on the distance image corresponding to the unit element of interest, an evaluation value is added to the unit element of interest. If the cumulative value of the aforementioned evaluation value exceeds the threshold, the unit element of interest is deleted. The image processing apparatus according to configuration 1, characterized in that...

[0090] (Composition 3) The aforementioned three-dimensional shape data is data using voxels as the unit elements, The system further includes setting means for setting the threshold for each local space obtained by dividing the imaging space, The correction means determines whether the cumulative value of the evaluation value exceeds the threshold, using a threshold from among the thresholds set for each local space that corresponds to the position of the voxel set indicated by the three-dimensional shape data in the imaging space. The image processing apparatus according to configuration 2, characterized in that...

[0091] (Composition 4) The setting means is, The imaging space is divided according to the division conditions. The threshold is set for each local space obtained by the division according to the threshold pattern. The image processing apparatus according to configuration 3, characterized in that

[0092] (Composition 5) It further has a user interface that accepts user instructions, The division conditions and threshold patterns are specified via the user interface. The image processing apparatus according to configuration 4, characterized in that...

[0093] (Composition 6) The image processing apparatus according to configuration 4 or 5, characterized in that the division conditions include the number of divisions, the interval between divisions, and the shape of the local space.

[0094] (Composition 7) The image processing apparatus according to configuration 4 or 5, characterized in that the threshold pattern has larger thresholds for local spaces closer to the center of the imaging space.

[0095] (Composition 8) The image processing apparatus according to configuration 3, characterized in that the setting means groups the plurality of distance images and sets the threshold for each local space based on the visibility of the distance images belonging to each group to each local space.

[0096] (Composition 9) The setting means is, For each group, count the number of distance images that are visible to the local space of interest out of all local spaces. Based on the number of visible distance images obtained for each group, a provisional threshold for the local space of interest is determined for each group. The minimum value among the provisional thresholds determined on a group basis is set as the threshold for the local space of interest. The image processing apparatus according to configuration 8, characterized by the above.

[0097] (Composition 10) The setting means is, For each group, count the number of distance images that are visible to the local space of interest out of all local spaces. Based on the number of visible distance images obtained for each group, a threshold is set for the local space of interest for each group. The correction means is The three-dimensional shape data is divided according to the group, and for each of the divided three-dimensional shape data, it is determined whether the cumulative value of the evaluation value exceeds the threshold set for each group. The image processing apparatus according to configuration 8, characterized by the above.

[0098] (Composition 11) The image processing apparatus according to any one of configurations 8 to 10, characterized in that the setting means groups the plurality of distance images based on the camera parameters of the imaging device corresponding to the distance image, such that distance images corresponding to cameras having a common imaging direction belong to the same group, or distance images having similar positions and orientations indicated by the camera parameters belong to the same group.

[0099] (Composition 12) The image processing apparatus according to configuration 3, characterized in that the setting means sets a larger threshold for local spaces that are contained within the field of view of more of the multiple distance images.

[0100] (Composition 13) The setting means sets a plurality of thresholds according to the imaging direction for each local space into which the imaging space is divided, The correction means is The voxel set is divided according to the imaging direction, For each of the divided voxel sets, it is determined whether the cumulative value of the evaluation value exceeds the threshold, using a threshold corresponding to the divided imaging direction from among a plurality of thresholds set for the local space corresponding to the position in the imaging space. The image processing apparatus according to configuration 3, characterized in that

[0101] (Composition 14) The first acquisition means, when multiple objects exist in the imaging space, acquires multiple three-dimensional shape data corresponding to each of the multiple objects, The correction means determines whether the cumulative value of the evaluation value exceeds the threshold, using the thresholds set for each local space, which correspond to the respective positions in the imaging space of the voxel sets represented by each of the plurality of three-dimensional shape data. The image processing apparatus according to configuration 3, characterized in that

[0102] (Composition 15) The image processing apparatus according to configuration 1, characterized in that the correction means controls the contribution rate to the evaluation result by weighting the unit elements constituting the plurality of distance images or the three-dimensional shape data.

[0103] (Composition 16) The system further includes a third acquisition means for acquiring multiple captured images showing the object, obtained by multiple imaging devices corresponding to the multiple viewpoints, The first acquisition means acquires the three-dimensional shape data by generating it based on the plurality of captured images. An image processing apparatus according to any one of configurations 1 to 15, characterized by the above.

[0104] (Composition 17) The image processing apparatus according to configuration 16, characterized in that the first acquisition means generates the three-dimensional shape data by a viewing volume cross-section method using the plurality of captured images.

[0105] (Composition 18) The system further includes a third acquisition means for acquiring multiple captured images showing the object, obtained by multiple imaging devices corresponding to the multiple viewpoints, The second acquisition means acquires the plurality of distance images by generating them based on the plurality of captured images. An image processing apparatus according to any one of configurations 1 to 15, characterized by the above.

[0106] (Composition 19) The second acquisition means generates the plurality of distance images by stereo matching using the plurality of captured images. The image processing apparatus according to configuration 18, characterized by the above.

[0107] (Composition 20) A derivation means for deriving three-dimensional surface information of the object based on the plurality of captured images and the three-dimensional shape data, A sorting means for sorting the derived three-dimensional surface information based on the distance from the shape surface of the object represented by the three-dimensional shape data, It further possesses, The second acquisition means modifies the plurality of distance images based on the selected surface three-dimensional information, The correction means performs the evaluation using the multiple corrected distance images. The image processing apparatus according to configuration 1, characterized in that...

[0108] (Composition 21) The image processing apparatus according to configuration 20, characterized in that the correction is a process that brings the depth value of the pixel on the depth image corresponding to each feature point extracted from the plurality of captured images closer to the depth value calculated based on the selected surface three-dimensional information.

[0109] (Method 1) The first acquisition step involves obtaining three-dimensional shape data of objects present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition step of acquiring multiple distance images corresponding to multiple viewpoints, A correction step in which the three-dimensional shape data is corrected by evaluating the three-dimensional shape data based on the plurality of distance images and, based on the results of the evaluation, deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data, An image processing method characterized by including

[0110] (Composition 22) A program for causing a computer to function as an image processing device as described in any one of items 1 to 21.

Claims

1. A first acquisition means for acquiring three-dimensional shape data of an object present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition means for acquiring multiple distance images corresponding to multiple viewpoints, Correction means for correcting the three-dimensional shape data by evaluating the three-dimensional shape data based on the plurality of distance images and, based on the results of the evaluation, deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data, Setting means for setting a threshold for each local space obtained by dividing the imaging space, It has, The correction means is From among the unit elements constituting the three-dimensional shape data, the unit element of interest is determined in order. If the distance from the viewpoint corresponding to the distance image to the unit element of interest is smaller than the distance to the pixel position on the distance image corresponding to the unit element of interest, an evaluation value is added to the unit element of interest. If the cumulative value of the aforementioned evaluation exceeds the threshold, the unit element of interest is deleted. The aforementioned three-dimensional shape data is data using voxels as the unit elements, The correction means determines whether the cumulative value of the evaluation value exceeds the threshold, using a threshold from among the thresholds set for each local space that corresponds to the position of the voxel set indicated by the three-dimensional shape data in the imaging space. The setting means is, The imaging space is divided according to the division conditions. The threshold is set for each local space obtained by the division according to the threshold pattern. An image processing apparatus characterized by the following:

2. It further has a user interface that accepts user instructions, The division conditions and threshold patterns are specified via the user interface. The image processing apparatus according to feature 1.

3. The image processing apparatus according to claim 1 or 2, characterized in that the division conditions include the number of divisions, the interval between divisions, and the shape of the local space.

4. The image processing apparatus according to claim 1 or 2, characterized in that the threshold pattern has larger thresholds for local spaces closer to the center of the imaging space.

5. A first acquisition means for acquiring three-dimensional shape data of an object present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition means for acquiring multiple distance images corresponding to multiple viewpoints, Correction means for correcting the three-dimensional shape data by evaluating the three-dimensional shape data based on the plurality of distance images and, based on the results of the evaluation, deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data, Setting means for setting a threshold for each local space obtained by dividing the imaging space, It has, The correction means is From among the unit elements constituting the three-dimensional shape data, the unit element of interest is determined in order. If the distance from the viewpoint corresponding to the distance image to the unit element of interest is smaller than the distance to the pixel position on the distance image corresponding to the unit element of interest, an evaluation value is added to the unit element of interest. If the cumulative value of the aforementioned evaluation exceeds the threshold, the unit element of interest is deleted. The aforementioned three-dimensional shape data is data using voxels as the unit elements, The correction means determines whether the cumulative value of the evaluation value exceeds the threshold, using a threshold from among the thresholds set for each local space that corresponds to the position of the voxel set indicated by the three-dimensional shape data in the imaging space. The setting means groups the plurality of distance images and sets the threshold for each local space based on the visibility of the distance images belonging to each group to each local space. An image processing apparatus characterized by the following:

6. The setting means is, For each group, count the number of distance images that are visible to the local space of interest out of all local spaces. Based on the number of visible distance images obtained for each group, a provisional threshold for the local space of interest is determined for each group. The minimum value among the provisional thresholds determined on a group basis is set as the threshold for the local space of interest. The image processing apparatus according to feature 5.

7. The setting means is, For each group, count the number of distance images that are visible to the local space of interest out of all local spaces. Based on the number of visible distance images obtained for each group, a threshold is set for the local space of interest for each group. The correction means is The three-dimensional shape data is divided according to the group, and for each of the divided three-dimensional shape data, it is determined whether the cumulative value of the evaluation value exceeds the threshold set for each group. The image processing apparatus according to feature 5.

8. The image processing apparatus according to any one of claims 5 to 7, characterized in that the setting means groups the plurality of distance images based on the camera parameters of the imaging device corresponding to the distance image, such that distance images corresponding to cameras having a common imaging direction belong to the same group, or distance images having similar positions and orientations indicated by the camera parameters belong to the same group.

9. A first acquisition means for acquiring three-dimensional shape data of an object present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition means for acquiring multiple distance images corresponding to multiple viewpoints, Correction means for correcting the three-dimensional shape data by evaluating the three-dimensional shape data based on the plurality of distance images and, based on the results of the evaluation, deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data, The system includes setting means for setting a threshold value for each local space obtained by dividing the imaging space, The correction means is From among the unit elements constituting the three-dimensional shape data, the unit element of interest is determined in order. If the distance from the viewpoint corresponding to the distance image to the unit element of interest is smaller than the distance to the pixel position on the distance image corresponding to the unit element of interest, an evaluation value is added to the unit element of interest. If the cumulative value of the aforementioned evaluation exceeds the threshold, the unit element of interest is deleted. The aforementioned three-dimensional shape data is data using voxels as the unit elements, The correction means determines whether the cumulative value of the evaluation value exceeds the threshold, using a threshold from among the thresholds set for each local space that corresponds to the position of the voxel set indicated by the three-dimensional shape data in the imaging space. The setting means sets a larger threshold for local spaces that are contained within the field of view of more of the multiple distance images. An image processing apparatus characterized by the following:

10. A first acquisition means for acquiring three-dimensional shape data of an object present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition means for acquiring multiple distance images corresponding to multiple viewpoints, Correction means for correcting the three-dimensional shape data by evaluating the three-dimensional shape data based on the plurality of distance images and, based on the results of the evaluation, deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data, The system includes setting means for setting a threshold value for each local space obtained by dividing the imaging space, The correction means is From among the unit elements constituting the three-dimensional shape data, the unit element of interest is determined in order. If the distance from the viewpoint corresponding to the distance image to the unit element of interest is smaller than the distance to the pixel position on the distance image corresponding to the unit element of interest, an evaluation value is added to the unit element of interest. If the cumulative value of the aforementioned evaluation exceeds the threshold, the unit element of interest is deleted. The aforementioned three-dimensional shape data is data using voxels as the unit elements, The correction means determines whether the cumulative value of the evaluation value exceeds the threshold, using a threshold from among the thresholds set for each local space that corresponds to the position of the voxel set indicated by the three-dimensional shape data in the imaging space. The setting means sets a plurality of thresholds according to the imaging direction for each local space into which the imaging space is divided, The correction means is The voxel set is divided according to the imaging direction, For each of the divided voxel sets, it is determined whether the cumulative value of the evaluation value exceeds the threshold, using a threshold corresponding to the divided imaging direction from among a plurality of thresholds set for the local space corresponding to the position in the imaging space. An image processing apparatus characterized by the following:

11. A first acquisition means for acquiring three-dimensional shape data of an object present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition means for acquiring multiple distance images corresponding to multiple viewpoints, Correction means for correcting the three-dimensional shape data by evaluating the three-dimensional shape data based on the plurality of distance images and, based on the results of the evaluation, deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data, The system includes setting means for setting a threshold value for each local space obtained by dividing the imaging space, The correction means is From among the unit elements constituting the three-dimensional shape data, the unit element of interest is determined in order. If the distance from the viewpoint corresponding to the distance image to the unit element of interest is smaller than the distance to the pixel position on the distance image corresponding to the unit element of interest, an evaluation value is added to the unit element of interest. If the cumulative value of the aforementioned evaluation exceeds the threshold, the unit element of interest is deleted. The aforementioned three-dimensional shape data is data using voxels as the unit elements, The correction means determines whether the cumulative value of the evaluation value exceeds the threshold, using a threshold from among the thresholds set for each local space that corresponds to the position of the voxel set indicated by the three-dimensional shape data in the imaging space. The first acquisition means, when multiple objects exist in the imaging space, acquires multiple three-dimensional shape data corresponding to each of the multiple objects, The correction means determines whether the cumulative value of the evaluation value exceeds the threshold, using the thresholds set for each local space, which correspond to the respective positions in the imaging space of the voxel sets represented by each of the plurality of three-dimensional shape data. An image processing apparatus characterized by the following:

12. A first acquisition means for acquiring three-dimensional shape data of an object present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition means for acquiring multiple distance images corresponding to multiple viewpoints, The device includes a correction means for evaluating the three-dimensional shape data based on the plurality of distance images, and correcting the three-dimensional shape data by deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data based on the results of the evaluation. The correction means controls the rate of contribution to the evaluation result by weighting the unit elements constituting the plurality of distance images or the three-dimensional shape data. An image processing apparatus characterized by the following:

13. The system further includes a third acquisition means for acquiring multiple captured images showing the object, obtained by multiple imaging devices corresponding to the multiple viewpoints, The first acquisition means acquires the three-dimensional shape data by generating it based on the plurality of captured images. The image processing apparatus according to feature 1.

14. The image processing apparatus according to claim 13, characterized in that the first acquisition means generates the three-dimensional shape data by a viewing volume cross-section method using the plurality of captured images.

15. The system further includes a third acquisition means for acquiring multiple captured images showing the object, obtained by multiple imaging devices corresponding to the multiple viewpoints, The second acquisition means acquires the plurality of distance images by generating them based on the plurality of captured images. The image processing apparatus according to feature 1.

16. The second acquisition means generates the plurality of distance images by stereo matching using the plurality of captured images. The image processing apparatus according to feature 15.

17. A first acquisition means for acquiring three-dimensional shape data of an object present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition means for acquiring multiple distance images corresponding to multiple viewpoints, Correction means for correcting the three-dimensional shape data by evaluating the three-dimensional shape data based on the plurality of distance images and, based on the results of the evaluation, deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data, A derivation means for deriving three-dimensional surface information of the object based on the plurality of distance images and the three-dimensional shape data, A sorting means for sorting the derived three-dimensional surface information based on the distance from the shape surface of the object represented by the three-dimensional shape data, It has, The second acquisition means modifies the plurality of distance images based on the selected surface three-dimensional information, The correction means performs the evaluation using the multiple corrected distance images. The image processing apparatus according to feature 1.

18. The image processing apparatus according to claim 17, characterized in that the correction is a process of bringing the depth value of the pixels on the depth image corresponding to each feature point extracted from the plurality of depth images closer to the depth value calculated based on the selected three-dimensional surface information.

19. The first acquisition step involves acquiring three-dimensional shape data of an object present in the imaging space, A distance image representing the distance to the aforementioned object, a second acquisition step of acquiring multiple distance images corresponding to multiple viewpoints, A correction step in which the three-dimensional shape data is corrected by evaluating the three-dimensional shape data based on the plurality of distance images and, based on the results of the evaluation, deleting unit elements that are presumed not to represent the shape of the object among the unit elements constituting the three-dimensional shape data, A setting step of setting a threshold for each local space obtained by dividing the imaging space, Includes, In the correction step, From among the unit elements constituting the three-dimensional shape data, the unit element of interest is determined in order. If the distance from the viewpoint corresponding to the distance image to the unit element of interest is smaller than the distance to the pixel position on the distance image corresponding to the unit element of interest, an evaluation value is added to the unit element of interest. If the cumulative value of the aforementioned evaluation exceeds the threshold, the unit element of interest is deleted. The aforementioned three-dimensional shape data is data using voxels as the unit elements, In the correction step, a threshold value corresponding to the position of the voxel set shown by the three-dimensional shape data in the imaging space is used from among the threshold values ​​set for each local space to determine whether the cumulative value of the evaluation value exceeds the threshold value. In the aforementioned setup step, The imaging space is divided according to the division conditions. The threshold is set for each local space obtained by the division according to the threshold pattern. An image processing method characterized by the following:

20. A program for causing a computer to perform the image processing method described in claim 19.

Citation Information

Patent Citations

  • Distance information output device and three-dimensional shape restoring device

    JP2008015863A