Image processing device, image processing method and program
The image processing device enhances 3D model generation by selecting optimal viewpoints and minimizing errors in feature point calculations, ensuring accurate three-dimensional coordinate determination for objects in diverse poses.
Patent Information
- Application Number
- JP2022011602
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2042-01-28
AI Technical Summary
Existing methods for generating 3D models from multiple viewpoints face challenges in accurately obtaining three-dimensional coordinates of feature points due to varying detection accuracy across different angles, particularly when capturing human faces, leading to reduced accuracy in world coordinates.
An image processing device that includes detection, assignment, and correction means to determine three-dimensional coordinates by selecting candidate viewpoints based on attribute information and minimizing errors in feature point calculations, using a combination of two-dimensional coordinates and direction vectors.
Enables high-accuracy acquisition of three-dimensional coordinates of feature points from multiple viewpoints, even when capturing objects in varying poses, by optimizing the selection of viewpoints and reducing calculation errors.
Smart Images

Figure 0007746176000009 
Figure 0007746176000010 
Figure 0007746176000011
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technique for reconstructing three-dimensional coordinates of feature points of an object. [Background technology]
[0002] A technology for generating a 3D model (three-dimensional shape data) of an object based on a plurality of captured images obtained by capturing images of the subject (object) from different viewpoints is widely used in fields such as computer graphics. Patent Document 1 discloses a method for selecting an optimal viewpoint when reconstructing the three-dimensional shape of a human head using image data obtained by capturing images of the head surrounded in three dimensions. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-317000 [Patent Document 2] Japanese Patent Application Laid-Open No. 2007-102601 Summary of the Invention [Problem to be solved by the invention]
[0004] When generating a 3D model of an object from multiple captured images corresponding to multiple viewpoints, it is necessary to accurately obtain the three-dimensional coordinates (world coordinates) of the object's feature points. Patent Document 2 discloses a method for correcting feature points of a standard face model to match the shape of the captured human face using image coordinates of feature points, such as the corners of the eyes and the corners of the mouth, in each image captured from multiple viewpoints. Image coordinates are two-dimensional coordinate information representing a point on an image. To accurately obtain the world coordinates of feature points of an object that can assume any pose, it is important to select an appropriate viewpoint from among the multiple viewpoints corresponding to each captured image so that the image coordinates of the feature points can be obtained with high accuracy. For example, when detecting feature points from a captured image of a human face, while feature points on the right half of the face can be detected with high accuracy in an image captured diagonally to the right, the detection accuracy of feature points on the left half (opposite side) often decreases. This is due to the fact that a human face has a three-dimensional structure that is symmetrical and oblique with respect to the nose. However, if the image captures the human face from the front, all feature points can be detected with high accuracy. However, a certain amount of parallax is required to accurately obtain world coordinates from image coordinates, and it is not possible to accurately obtain three-dimensional coordinates of facial feature points from only a captured image that captures a person's face from the front.
[0005] As described above, there are advantages to using images captured from an oblique angle to reconstruct the world coordinates of feature points. On the other hand, there is a disadvantage in that the accuracy of the image coordinates of feature points farther from the imaging viewpoint decreases, resulting in a problem of reduced accuracy in the acquired 3D coordinates. [Means for solving the problem]
[0006] The image processing device according to the present disclosure includes a detection means for detecting feature points of an object from a plurality of images acquired by capturing images from a plurality of viewpoints, an assignment means for assigning attribute information indicating a region of the object to the detected feature points, a determination means for determining three-dimensional coordinates of the feature points to which the same attribute information has been assigned by the assignment means based on two-dimensional coordinates of the feature points in images corresponding to two or more viewpoints that are less than the plurality of viewpoints, and a correction means for correcting three-dimensional shape data of the object using information on the three-dimensional coordinates. the determining means extracts candidate viewpoints from among the plurality of viewpoints for each of the same attribute information, and determines three-dimensional coordinates of the feature points to which the same attribute information has been assigned based on two-dimensional coordinates of the feature points on the image corresponding to a viewpoint selected from the candidate viewpoints; the assigning means acquires three-dimensional shape data having a basic structure of the object and position information of its feature points, and determines the content of the attribute information to be assigned by clustering normals of the feature points specified by the position information, assigns the attribute information corresponding to the detected feature points from the determined attribute information, acquires normals of the feature points specified by the position information, defines one or more direction vectors with respect to a local coordinate system of the three-dimensional shape data, and assigns the attribute information corresponding to the direction vectors based on an angle between the direction vectors and the normals. It is characterized by: [Effects of the Invention]
[0007] According to the technology of the present disclosure, it is possible to acquire with high accuracy three-dimensional coordinates of feature points of an object from a plurality of captured images taken from different viewpoints. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram showing an example of the hardware configuration of an image processing apparatus. [Figure 2] FIG. 2 is a block diagram showing the software configuration of the image processing apparatus according to the first embodiment. [Figure 3] 6 is a flowchart showing the flow of processing for deriving the world coordinates of feature points according to the first embodiment. [Figure 4] Schematic diagram showing how a person's face is captured from different viewpoints. [Figure 5] FIG. 4A is a diagram showing an example of a captured image of the left face, and FIG. 4B is a diagram showing an example of facial feature points. [Figure 6] FIG. 1A is a diagram showing feature points of a frontal face, and FIGS. 1B and 1C are diagrams showing examples of attribute labels assigned to the facial feature points. [Figure 7] 1 is a diagram illustrating roll, pitch, and yaw in a right-handed system. [Figure 8] Shows the definition of the camera coordinate system. [Figure 9] FIG. 10 is a diagram illustrating extraction of candidate viewpoints. [Figure 10]FIG. 1A is a diagram showing how two faces are captured from two viewpoints, and FIGS. 1B and 1C are diagrams showing captured images corresponding to the two viewpoints. [Figure 11] (a) is a diagram explaining the calculation error of the world coordinates of feature points, and (b) and (c) are diagrams explaining the person identification results. [Figure 12] FIG. 10 is a block diagram showing the software configuration of an image processing device according to a second embodiment. [Figure 13] 10 is a flowchart showing the flow of processing for deriving the world coordinates of feature points according to the second embodiment. [Figure 14] FIG. 10 is a diagram showing an example of feature points of an automobile. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the configurations shown in the following embodiments are merely examples, and the scope of the present disclosure is not limited to only those configurations.
[0010] [Embodiment 1] <Hardware configuration> FIG. 1 is a diagram illustrating an example of the hardware configuration of an image processing device 100 according to this embodiment. In FIG. 1, a CPU 101 uses a RAM 102 as a work memory, executes programs stored in a ROM 103 and a hard disk drive (HDD) 105, and controls the operation of each block (described later) via a system bus 112. An HDD interface (hereinafter, interface will be referred to as "I / F") 104 connects a secondary storage device such as the HDD 105 or an optical disk drive. The HDD I / F 104 is, for example, an I / F such as a serial ATA (SATA). The CPU 101 can read data from and write data to the HDD 105 via the HDD I / F 104. Furthermore, the CPU 101 can load data stored in the HDD 105 into the RAM 102, and conversely, can save data loaded in the RAM 102 to the HDD 105. The CPU 101 can then execute the data loaded in the RAM 102 as a program. The input I / F 106 connects to an input device 107 such as a keyboard, mouse, digital camera, or scanner. The input I / F 106 is, for example, a serial bus I / F such as USB or IEEE1394. The CPU 101 can read various data such as captured images from the input device 107 via the input I / F 106. The output I / F 108 connects the image processing device 100 to a display, which is an output device 109. The output I / F 108 is, for example, a video output I / F such as DVI or HDMI (registered trademark). The CPU 101 can send data to the display via the output I / F 108 and display a predetermined video on the display. The network I / F 1108 connects the image processing device 100 to an external server 111.
[0011] <Software configuration> Fig. 2 is a block diagram showing the software configuration of the image processing device 100 according to this embodiment. Below, each function of the image processing device 100 according to this embodiment will be explained with reference to the flowchart shown in Fig. 3. In the following explanation, the symbol "S" means step.
[0012] In S301, the data acquisition unit 201 reads and acquires data of a plurality of images captured from different viewpoints (hereinafter referred to as "multi-viewpoint images") and their camera parameters from the HDD 105 or the like. FIG. 4 is a schematic diagram showing how a person's head 400, as an object to be imaged, is imaged from six different viewpoints 401 to 406 from which the face can be seen. Here, the explanation will be given assuming that the multi-viewpoint images obtained by capturing images from six different directions as shown in the figure and the camera parameters at that time have been acquired. The camera parameters include the position, orientation, focal length, and principal point of the viewpoint, and are information that can convert two-dimensional coordinates on the image into a ray passing through the position of the viewpoint.
[0013] In S302, the feature point detection unit 202 detects feature points of an object from each captured image constituting the acquired multi-view image. Facial feature points can be detected from captured images showing a human face using known face recognition technologies, such as Dlib and OpenCV. Here, seven points—the outer and inner corners of the eyes, the corners of the mouth, and the tip of the nose—are detected as facial feature points. Note that the seven facial feature points are merely an example, and the facial feature points may not include any of the seven points, or may include other points such as the space between the eyebrows, points on the cheeks, or points on the jaw line. FIG. 5(a) shows a captured image obtained by capturing an image of the head 400 from the viewpoint 402 in FIG. 4. FIG. 5(b) shows the positions (image coordinates) of the seven facial feature points detected from the captured image of FIG. 5(a): the outer corner 501 of the right eye, the inner corner 502 of the right eye, the outer corner 503 of the left eye, the inner corner 504 of the left eye, the tip of the nose 505, the right corner 506 of the mouth, and the left corner 507 of the mouth. In this way, in a captured image in which the face is photographed from the left side, the detection accuracy of feature points on the right side of the face, i.e., the right eye outer corner 501, the right eye inner corner 502, and the right mouth corner 506, is relatively lower. Furthermore, in this embodiment, the feature point detection unit 202 also estimates the pose of the object. For example, the above-mentioned Dlib has a function to detect facial feature points as well as a function to estimate the pose of the face, and by using this, it is possible to also obtain pose information of the face. Here, the pose of the object is relative to the imaging viewpoint and is represented by roll, pitch, and yaw.
[0014] In S303, the labeling unit 203 assigns, to each of the feature points detected in S302, a label (hereinafter referred to as an "attribute label") as attribute information indicating to which region of the object the feature point belongs. FIG. 6(a) shows the seven facial feature points (right eye outer corner 601, right eye inner corner 602, left eye outer corner 603, left eye inner corner 604, nose tip 605, right mouth corner 606, and left mouth corner 607) detected from the captured image corresponding to the viewpoint 404. FIG. 6(b) shows the attribute labels assigned to the facial feature points 601 to 607 in FIG. 6(a). As shown in FIG. 6(b), right labels are assigned to the feature points belonging to the right side of the face, including the center (right eye outer corner 601, right eye inner corner 602, nose tip 605, and right mouth corner 606). Furthermore, left labels are assigned to the feature points belonging to the left side of the face, including the center (left eye outer corner 603, left eye inner corner 604, nose tip 605, and left mouth corner 607). It should be noted that attribute labels are assigned for each feature point. For example, if a right label is assigned to the "outer corner of the right eye," the right label will be assigned to the "outer corner of the right eye" in all captured images constituting the multi-viewpoint image. Here, attribute labels are assigned to classify the face into either the left or right region. However, the types of attribute labels are not limited to this. For example, as shown in FIG. 6(c), attribute labels may be assigned to indicate whether the face belongs to the upper or lower region. Furthermore, combining left and right and top and bottom labels may result in four classifications, such as an upper right label, a lower right label, an upper left label, and a lower left label. The classification of attribute labels may be appropriately determined according to the shape characteristics of the target object. In this embodiment, it is assumed that attribute labels are assigned automatically based on the results of feature point detection, but attribute labels may also be assigned manually by an operator.
[0015] In S304, the world coordinate determination unit 204 calculates the world coordinates of the feature points detected in S302 for each attribute label assigned in S303. In this calculation, first, from among the viewpoints in the multi-view image, viewpoints (candidate viewpoints) that are candidates for the viewpoints used to calculate the world coordinates of the feature points are extracted for each attribute label based on the posture information of the object identified in S302. Then, the image coordinates (two-dimensional coordinates) of the feature points on the captured image corresponding to the extracted candidate viewpoints are used to calculate the world coordinates (three-dimensional coordinates) of the feature points. Here, a specific flow of processing for calculating the world coordinates of the feature points for each attribute label when the object is a human face and two types of attribute labels (left and right) are assigned will be described in detail with reference to the drawings.
[0016] <Extraction of candidate viewpoints> As described above, face pose information is expressed by roll, pitch, and yaw. FIG. 7 is a diagram illustrating roll, pitch, and yaw rotation in a right-handed system. In this embodiment, the right-handed system is used, but a left-handed system may also be used. Yaw represents left and right turning relative to the viewpoint. When yaw is 0 degrees, the face faces forward. Roll represents rotation relative to the viewpoint. When roll is 0 degrees, the face is upright (when roll is 180 degrees, the face is upside down). Pitch represents the elevation and depression angles relative to the viewpoint. When pitch is 0 degrees, the face faces forward, and as the pitch increases, the face faces downward. For example, when roll and pitch are 0 degrees, if yaw is positive, it can be determined that the right side of the face is captured, and if yaw is negative, it can be determined that the left side of the face is captured. Therefore, roll, pitch, and yaw are converted into direction vectors (unit vectors) in a three-dimensional camera coordinate system. If the x-component of the viewpoint is equal to or less than a threshold R, it is considered a candidate viewpoint labeled as left, and if it is equal to or greater than a threshold L, it is considered a candidate viewpoint labeled as right. FIG. 8 shows the definition of the camera coordinate system. Here, the camera coordinate system is a coordinate system in which the position of the imaging device (camera) is defined as the origin, the direction of the camera's optical axis as z, the right direction as x, and the downward direction as y. When the face is facing the viewpoint, the z-axis value of the directional vector is negative. The left-right direction of the face is represented by the x-axis value, which corresponds to the sine of the left-right angle of the face. Therefore, for example, when assigning attribute labels to a range of up to 25 degrees in the opposite direction from the front, the threshold R can be set to sin(25°) and the threshold L can be set to -sin(25°). Figure 9 shows the range of ±25 degrees from the front direction of the head 400 in the specific example of Figure 4. In this example, four viewpoints 401 to 404 on the left side of a line segment 901 indicating +25° are extracted as candidate viewpoints labeled as left, and four viewpoints 403 to 406 on the right side of a line segment 902 indicating -25° are extracted as candidate viewpoints labeled as right. In this way, two or more viewpoints that are less than the plurality of viewpoints corresponding to the input multi-viewpoint image are extracted as candidate viewpoints.
[0017] <Calculating world coordinates> Next, a pair of two viewpoints is selected from the candidate viewpoints extracted for each attribute label, and the image coordinates of the feature points on the captured image corresponding to the two viewpoints are used to calculate the world coordinates of the feature points assigned the same attribute label. Once the calculation is complete for all pairs of two viewpoints, the world coordinates of the pair with the smallest error are set as the world coordinates for the same attribute label. Here, the error is treated as the distance between two rays in a twisted three-dimensional space. The error calculation method will be explained in detail with reference to a specific example in Figure 9. First, the image coordinates of the feature points detected at viewpoints 401 to 406 are calculated as q ij where i represents the viewpoint number and j represents the feature point number. Next, the pose information of each viewpoint in the world coordinate system is defined as R i , location information d i Let R i and d i are generally called the extrinsic parameters of the camera. Next, the focal length and principal point of each camera are calculated as a 3x3 matrix. The intrinsic parameters are calculated as A i Using these parameters, the ray r corresponding to the feature point j at the viewpoint i is ij is calculated using the following formula (1).
[0018]
number
[0019] In the above formula (1), t is a coefficient. ij q ij It is a homogeneous coordinate (3-dimensional) of the ray, which is generated by adding 1 to the last element of the 2-dimensional image coordinate. It is rare for two rays consisting of feature points obtained independently to intersect, and in most cases they are in a twisted relationship. Therefore, when finding the intersection point, the midpoint of the line segment consisting of two points on the two rays is approximately obtained when it is shortest. Here, the ray r ij Two rays, r1(t1) and r2(t2), are rewritten as in the following equations (2) and (3), respectively.
[0020]
number
[0021]
number
[0022] At this time, the coefficients t1 and t2 corresponding to the points on each ray of the above-mentioned shortest line segment are expressed by the following equations (2) and (3), respectively.
[0023]
number
[0024]
number
[0025] Therefore, the intersection point h obtained is the midpoint of the two points obtained from these coefficients t1 and t2, and is expressed by the following equation (6).
[0026]
number
[0027] The error e is half the length of the line segment and can be calculated using the following equation (7).
[0028]
number
[0029] In this way, the error e is calculated for a pair of two viewpoints selected from the candidate viewpoints, and the world coordinates obtained from the pair with the smallest error e are set as the world coordinates of the feature points for that attribute label. For example, since the deviation of the feature points on the right side of the face is usually large in a captured image corresponding to viewpoint 401 capturing the left side of the face, the error e becomes large in any combination with viewpoints 402 to 404. Therefore, the world coordinates obtained from a pair of two viewpoints including viewpoint 401 are not adopted as the world coordinates for the left label. The same is true for viewpoint 406 capturing the right side of the face. In other words, since the error e becomes large in any combination with viewpoints 403 to 405, the world coordinates obtained from a pair of two viewpoints including viewpoint 406 are not adopted as the world coordinates for the right label.
[0030] To summarize the above, viewpoints 401 and 406 capture images from positions that are significantly tilted relative to the front direction of the face, resulting in a large deviation in the detected positions of feature points, and as a result, the above-mentioned error e becomes large. Also, viewpoints 402 to 405 capture the face from the front direction, so feature points can be detected more accurately than viewpoints 401 and 406. However, on the other hand, the detection accuracy of feature points on the opposite side of the imaging direction (the right half of the face when viewed from viewpoints 402 and 403, and the left half of the face when viewed from viewpoints 404 and 405) tends to decrease, resulting in a large error. As a result, for the left label, world coordinates calculated from the pair of viewpoints 402 and 403 are used, and for the right label, world coordinates calculated from the pair of viewpoints 404 and 405 are used.
[0031] In S305, the world coordinate determination unit 204 determines the world coordinates of the feature points in the entire object based on the world coordinates of the feature points calculated for each attribute label. In the example of FIG. 6(b) described above, a right label is assigned to each feature point on the right side of the face (the outer corner 601 of the right eye, the inner corner 602 of the right eye, the tip of the nose 605, and the right corner of the mouth 606), and world coordinates estimated from the selected viewpoints 404 and 405 are obtained for each of them. A left label is assigned to each feature point on the left side of the face (the outer corner 603 of the left eye, the inner corner 604 of the left eye, the tip of the nose 605, and the left corner of the mouth 607), and world coordinates estimated from the selected viewpoints 402 and 403 are obtained for each of them. In this case, only one attribute label is assigned to each feature point other than the tip of the nose 605, so the world coordinates calculated for each attribute label are used as they are. For the tip of the nose 605, world coordinates are obtained for both the right and left labels, so the midpoint of these labels is used as the world coordinate for the tip of the nose 605. If there are three or more attribute labels, the world coordinates of the target feature point can be determined by averaging them. Alternatively, the median or most frequent value of the world coordinates obtained for each attribute label can be used, or the world coordinate with the smallest reprojection error can be used.
[0032] In S306, the output unit 205 outputs the world coordinates derived by the world coordinate determination unit 204. The output world coordinate information can be used to correct the three-dimensional model. For example, the world coordinate information can be used to identify recessed portions in a pre-generated three-dimensional model, and the corresponding data can be removed. Alternatively, the data does not need to be removed, and the positions of the elements constituting the three-dimensional model can be changed. In this way, the world coordinate information can be used to accurately reproduce the unevenness of the three-dimensional model. The pre-generated three-dimensional model can be generated based on captured images of the subject, or can be generated using computer graphics (CG) technology, or can be created by combining these. The world coordinate information can also be used to estimate the posture of the subject (face or head), for example. The subject can also be an object other than a face.
[0033] The above is the flow of processing in the image processing device 100 according to this embodiment for obtaining the world coordinates of feature points of an object from a multi-viewpoint image. In this embodiment, for a pair of two viewpoints selected from the candidate viewpoints, the errors of feature points with the same attribute label are calculated, and the two viewpoints with the smallest maximum value are selected. As a result, even if the error for a feature point with a large deviation in a captured image corresponding to a certain viewpoint is accidentally estimated to be small, the errors for other feature points will be large, making it difficult to select that certain viewpoint. This ultimately makes it possible to select the most appropriate viewpoint.
[0034] <Variation 1> In the method of the above-described embodiment, when the distance b between the viewpoints is short (the rays are nearly parallel), the error e tends to become large in the depth direction relative to the viewpoint. Taking this into consideration, the distance from the line connecting the viewpoints to the estimated point may be defined as c, and the error e' expressed by the following equation (8) may be estimated as an error that becomes larger as the distance between the viewpoints becomes smaller.
[0035]
number
[0036] <Variation 2> In the above-described embodiment, pairs of two viewpoints are sequentially selected from candidate viewpoints for each attribute label, and the world coordinates of the feature points calculated from each pair are selected from the pair of two viewpoints with the smallest error, and are adopted as the world coordinates for that attribute label. In addition to this method, for example, the world coordinates of the feature points may be calculated using all candidate viewpoints, and the median or average of the calculated world coordinates may be adopted as the world coordinates for that attribute label. Alternatively, the viewpoint with the smallest sum of the distances to the rays of all candidate viewpoints may be selected to calculate the world coordinates of the feature points. Furthermore, these methods may be combined to exclude viewpoints with large reprojection errors (estimated to have large calculation errors) to determine the world coordinates of the feature points for the attribute label. Furthermore, viewpoints may be selected from the candidate viewpoints so that the angular density of the viewpoints relative to the object is constant.
[0037] According to this embodiment, it is possible to acquire with high accuracy the world coordinates of the feature points of an object that can assume any pose in the imaging environment.
[0038] [Embodiment 2] In the first embodiment, a specific example was described in which the world coordinates of facial feature points for one person's head are acquired with high accuracy. As shown in FIG. 10(a), when multiple people's heads 1003 and 1004 are simultaneously captured from different viewpoints 1001 and 1002, a captured image corresponding to viewpoint 1001 (FIG. 10(b)) and a captured image corresponding to viewpoint 1002 (FIG. 10(c)) are obtained. Each of the captured images contains multiple people's faces, and facial feature points for each person are detected from each captured image. However, as is, it is unclear to which person the facial feature points detected from each captured image correspond to across different viewpoints (different captured images). FIG. 11(a) shows an example in which world coordinates are calculated using image coordinates of facial feature points belonging to different people, resulting in facial feature points appearing in positions where no human faces actually exist. In FIG. 11(a), in addition to the left and right eyes 1101 of an actual head 1003 and the left and right eyes 1102 of an actual head 1004, left and right eyes 1103 appear in a position where no human head exists. To prevent such errors, it is necessary to associate feature points detected from each captured image with objects, i.e., to which object each belongs. If the object is a human head (face), it is possible to solve this problem by using face recognition technology, which can identify the same person appearing in different captured images. However, this would result in other problems, such as the need to acquire the facial features of each person in advance and the time required for processing.
[0039] Therefore, an aspect in which objects are identified between different viewpoints (between different captured images) using intermediate information obtained in the process of calculating the world coordinates of facial feature points will be described as embodiment 2. Note that a description of the content common to embodiment 1 will be omitted, and the following description will focus on the differences.
[0040] <Software configuration> Fig. 12 is a block diagram showing the software configuration of the image processing device 100 according to this embodiment. Below, each function of the image processing device 100 according to this embodiment will be described with reference to the flowchart shown in Fig. 13. In the following description, the symbol "S" means step.
[0041] S1301 to S1303 are the same as S301 to S303 in the flow of FIG. 3 of the first embodiment, and therefore description thereof will be omitted. In S1304, the identification unit 1201 identifies multiple objects appearing in multiple captured images between the captured images. From the captured images shown in FIGS. 10(b) and 10(c) described above, facial feature points for the head 1003 and facial feature points for the head 1004 are obtained, respectively. For simplicity of explanation, an example will be described in which the left and right eyes are detected as feature points. The solid arrows in FIG. 10(a) represent rays corresponding to both eyes of the face 1003 and the face 1004 when viewed from the viewpoint 1001, and the dashed arrows represent rays corresponding to both eyes of the head 1003 and the head 1004 when viewed from the viewpoint 1002. Furthermore, solid line ray 1011 is a ray directed toward the right eye of head 1003 imaged from viewpoint 1001, and dashed line ray 1012 is a ray directed toward the right eye of head 1003 imaged from viewpoint 1002. Solid line ray 1013 is a ray directed toward the right eye of head 1004 imaged from viewpoint 1001, and dashed line ray 1014 is a ray directed toward the right eye of head 1004 imaged from viewpoint 1002. Now, there are two rays heading toward the right eye at each imaging viewpoint, and four different world coordinates for the right eye are calculated based on the intersections of the rays resulting from these combinations. However, for example, in the combination of solid ray 1011 and dashed ray 1014, the intersection point of the rays (not shown) is far behind the imaging viewpoint. In other words, the world coordinates for the right eye obtained from this combination are located behind the imaging viewpoint and cannot be valid, so it is easy to see that they are incorrect. A similar result is obtained for the left eye. However, in the case of the left and right eyes 1103 appearing in front of the imaging viewpoints, as shown in 11(a) above, the world coordinates are valid, so it is not immediately possible to determine whether they are incorrect.
[0042] First, the world coordinates of all detected right and left eyes are calculated. Then, for each right eye, we check whether the world coordinate of the left eye exists at a likely position. For example, the distance between the left and right eyes of an adult Japanese woman is approximately 10 cm. Therefore, we add a margin to account for children and men, and check whether the world coordinate of the left eye is 8 to 15 cm away from the calculated world coordinate of the right eye. If the left eye is located at a likely position relative to the right eye, we determine that the left and right eyes associated with that combination actually exist and that their world coordinates are approximately accurate. This makes it possible to eliminate facial feature points that cannot actually exist. For ease of understanding, we have used a pair of left and right eyes as an example. However, in reality, we target a set of feature points with the same attribute label and check whether the distance between feature points (for example, the tip of the nose and the right corner of the mouth) is within a normal distance range. Combinations of feature points that fall outside this range are excluded. Then, we further check the distance between feature points based on the calculated 3D coordinates between different attribute labels to find combinations with consistent positional relationships between feature points. This allows a combination of feature points relating to the same person to be specified, and each of the faces of multiple people appearing in multiple captured images from different viewpoints to be identified. (b) and (c) in Figure 11 show groups of faces in captured images from each viewpoint that have been identified as the same face (person) as obtained by the above-mentioned combination search, with (b) showing the group of head 1004 and (c) showing the group of head 1003. In the steps from S1305 onwards, facial feature points are processed for each group of identified people, making it possible to obtain highly accurate world coordinates.
[0043] As described above, according to this embodiment, even in a situation where multiple objects are imaged simultaneously, it is possible to accurately acquire the world coordinates of the feature points of each object.
[0044] [Embodiment 3] In the first and second embodiments, a case where the world coordinates of facial feature points are derived using a human head as an example of an object has been described, but the object to be imaged is not limited to a human face. As an example, a case where the world coordinates of feature points of an automobile as an imaged object will be described as the third embodiment. Note that the hardware and software configurations of the image processing device are the same as those of the first embodiment, and therefore a description thereof will be omitted, and the following description will focus on the differences.
[0045] <Assigning attribute labels> In this embodiment, a basic model of the object is used when assigning attribute labels to the detected feature points (S303). Here, the basic model is three-dimensional shape data that includes the object's rough three-dimensional structure (basic structure) and the position information of its feature points. The feature points detected from each captured image correspond to all or some of the feature points represented by the basic model. Therefore, attribute labels such as left / right, up / down, and front / rear can be assigned to each detected feature point according to the normal direction of the surface of the basic model. Specifically, the normals are clustered and attribute labels are assigned to each cluster. FIG. 14 illustrates six feature points (front wheels 1401a and 1401b, rear wheels 1402a and 1402b, and front lights 1403a and 1403b) detected from a captured image of a car captured from an oblique front view. Here, by using feature point detection using deep learning, it is possible to estimate feature points of occluded parts of the car based on learning. In this case, the two right wheels 1401a and 1402a can be assigned a right label, the two left wheels 1401b and 1402b can be assigned a left label, and the two front lights 1403a and 1403b can be assigned a front label. Note that the normal direction is defined with respect to the local coordinate system of the basic model, and rotates according to the attitude of the car determined by a method described below, and the direction vector is calculated in a world coordinate system common to the imaging device.
[0046] <Posture estimation> In this embodiment, the world coordinate determination unit 204, rather than the feature point detection unit 202, estimates the pose of the object based on a basic model before extracting candidate viewpoints. The object in this embodiment is an automobile, and in general, the centers of the wheels of an automobile are located on a plane parallel to the ground, and further, the front lights are located parallel to the front wheels. These structural characteristics of an automobile are used to estimate the pose of the automobile captured in the captured image. The specific procedure is as follows.
[0047] First, the world coordinates of each feature point (here, the six feature points described above) are calculated using the method described in the first embodiment. Next, the world coordinates of the feature points of the four wheels among the calculated world coordinates are referenced to determine the up / down, left / right, and front / rear directions of the automobile shown in the image. Specifically, the lateral direction (left / right direction) is defined as the direction in which the angle formed by the line connecting the right front wheel 1401a and the left front wheel 1401b and the line connecting the right rear wheel 1402a and the left rear wheel 1402b is the smallest. The front / rear direction is defined as the direction that is perpendicular to the lateral direction and forms the smallest angle with the line connecting the right front wheel 1401a and the right rear wheel 1402a and the line connecting the left front wheel 1401b and the left rear wheel 1402b. The up / down direction is also determined from the cross product of the lateral direction and the front / rear direction. This makes it possible to determine the attitude of the automobile shown in the captured image. Furthermore, the three-dimensional position in the captured space can be identified by averaging the world coordinates of the feature points of the four wheels. By converting the posture of the object in the world coordinate system thus obtained into the camera coordinate system of each viewpoint, it becomes possible to extract candidate viewpoints for each attribute label (S304), as in embodiment 1. Note that although an automobile has been used as an example in the description here, it goes without saying that the object to which this embodiment is applicable is not limited to an automobile.
[0048] As described above, the configuration of this embodiment also makes it possible to acquire with high accuracy the three-dimensional coordinates of the feature points of an object that can assume any pose in the imaging environment.
[0049] (Other Examples) The present disclosure can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]
[0050] 100 Image processing device 202 Feature point detection unit 203 Label Assignment Unit 204 World Coordinate Determination Unit
Claims
1. a detection means for detecting feature points of an object from a plurality of images acquired by capturing images from a plurality of viewpoints; an attribute information assigning unit that assigns attribute information indicating a region of the object to the detected feature points; a determining means for determining three-dimensional coordinates of the feature points to which the same attribute information is assigned by the assigning means, based on two-dimensional coordinates of the feature points in images corresponding to two or more viewpoints not exceeding the plurality of viewpoints; a correcting means for correcting the three-dimensional shape data of the object using the three-dimensional coordinate information; and The determining means extracting candidate viewpoints from the plurality of viewpoints for each of the same attribute information; determining three-dimensional coordinates of feature points to which the same attribute information is assigned based on two-dimensional coordinates of the feature points on the image corresponding to a viewpoint selected from the candidate viewpoints; The applying means is Acquire three-dimensional shape data having the basic structure of the object and position information of its feature points; determining the content of the attribute information to be assigned by clustering normals of the feature points identified by the position information; assigning the attribute information corresponding to the detected feature points from among the determined attribute information; obtaining a normal to the feature point identified by the position information; defining one or more direction vectors for a local coordinate system of the three-dimensional shape data, and assigning the attribute information corresponding to the direction vectors based on an angle between the direction vectors and a normal line; 1. An image processing device comprising:
2. 2. The image processing apparatus according to claim 1, wherein said determining means selects a viewpoint from among said candidate viewpoints so that the angular density of viewpoints with respect to said object is constant.
3. an estimation means for estimating the pose of the object; the determining means extracts the candidate viewpoints based on the estimated posture of the object.
3. The image processing device according to claim 1, wherein the image processing device is a computer.
4. 4. The image processing device according to claim 1, wherein the assigning means assigns attribute information indicating that the feature points belong to a right region to feature points that belong to a right side including a center of the object, and assigns attribute information indicating that the feature points belong to a left region to feature points that belong to a left side including the center of the object.
5. 4. The image processing device according to claim 1, wherein the assigning means assigns attribute information indicating that feature points belonging to an upper region including a center of the object belong to an upper area, and assigns attribute information indicating that feature points belonging to a lower region including the center of the object belong to a lower area.
6. 4. The image processing device according to claim 1, wherein the assigning means assigns attribute information indicating that feature points belonging to a front region including a center of the object belong to a front area, and assigns attribute information indicating that feature points belonging to a rear region including the center of the object belong to a rear area.
7. 6. The image processing device according to claim 1, wherein the determining means determines one of an average value, a median value, and a mode value of the three-dimensional coordinates of the feature points to which the attribute information is assigned as the three-dimensional coordinates of the feature points.
8. 6. The image processing device according to claim 1, wherein the determining means determines, as the three-dimensional coordinates of the feature point, the three-dimensional coordinates having the smallest reprojection error among the three-dimensional coordinates of the feature point to which the attribute information is assigned.
9. The determining means If a plurality of the objects appear in the plurality of images, the objects are identified between the different images; determining, for each identified object, the three-dimensional coordinates of the detected feature points; 9. The image processing device according to claim 1, wherein the image processing device is a computer.
10. 10. The image processing device according to claim 9, wherein the determining means calculates distances between feature points for a set of feature points to which the same attribute information is assigned, and identifies a combination of feature points relating to the same object from a positional relationship between the feature points based on the calculated distances, thereby identifying the object.
11. a detection step of detecting feature points of an object from a plurality of images acquired by capturing images from a plurality of viewpoints; an assigning step of assigning attribute information to the detected feature points, the attribute information indicating the area of the object to which the feature points belong; a determining step of determining three-dimensional coordinates of the feature points to which the same attribute information has been assigned in the assigning step, based on two-dimensional coordinates of the feature points in images corresponding to two or more viewpoints not greater than the plurality of viewpoints; a correcting step of correcting three-dimensional shape data of the object using the three-dimensional coordinate information; Including, In the determining step, extracting candidate viewpoints from the plurality of viewpoints for each of the same attribute information; determining three-dimensional coordinates of feature points to which the same attribute information is assigned based on two-dimensional coordinates of the feature points on the image corresponding to a viewpoint selected from the candidate viewpoints; In the imparting step, Acquire three-dimensional shape data having the basic structure of the object and position information of its feature points; determining the content of the attribute information to be assigned by clustering normals of the feature points identified by the position information; assigning the attribute information corresponding to the detected feature points from among the determined attribute information; obtaining a normal to the feature point identified by the position information; defining one or more direction vectors for a local coordinate system of the three-dimensional shape data, and assigning the attribute information corresponding to the direction vectors based on an angle between the direction vectors and a normal line; An image processing method comprising:
12. A program for causing a computer to function as the image processing device according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method for determining set of optimal viewpoint to construct 3D shape of face from 2d image acquired from set of optimal viewpoint
JP2005317000A
Apparatus and method for generating solid model
JP2007102601A
Face image recognition apparatus, face image recognition method, face image recognition program, and recording medium recording this program
JP2009157767A
Image processing device, image processing method and program
JP2019128641A
JPP6908312B