Three-dimensional data generation device, generation method, program, and computer-readable non-transitory storage medium

The 3D data generation device improves camera pose estimation in 3D reconstruction by identifying subject areas and using peripheral feature points, markers, and conversion tables to correct errors, achieving accurate 3D reconstruction.

WO2026038486A1PCT designated stage Publication Date: 2026-02-19SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/027491
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-15
Filing Date
2025-08-04
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing 3D reconstruction methods face challenges in accurately estimating camera pose due to optical characteristics, occlusion, and lens distortion, which hinder the matching of feature points between images.

Method used

A 3D data generation device and method that identifies a subject area and extracts peripheral feature points to estimate camera posture, using markers and a conversion table to correct errors, thereby improving estimation accuracy.

Benefits of technology

Enhances the accuracy of camera posture estimation by overcoming obstacles caused by optical characteristics and occlusion, ensuring precise 3D reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025027491_19022026_PF_FP_ABST
    Figure JP2025027491_19022026_PF_FP_ABST
Patent Text Reader

Abstract

A three-dimensional data generation device according to the present disclosure comprises a subject region recognition unit, a peripheral feature point extraction unit, and a camera posture estimation unit. The subject region recognition unit detects a subject region from an image captured by a camera. The peripheral feature point extraction unit extracts, as a peripheral feature point group, a feature point group of a peripheral part of the subject region from a feature point group included in the image. The camera posture estimation unit estimates the camera posture by selectively using the peripheral feature point group.
Need to check novelty before this filing date? Find Prior Art

Description

Three-dimensional data generation device, generation method, program, and computer-readable non-transitory storage medium

[0001] The present invention relates to a three-dimensional data generation device, a generation method, a program, and a computer-readable non-transitory storage medium.

[0002] There is a known technology for reconstructing three-dimensional shapes from camera images. This technology is called 3D reconstruction or photogrammetry. Applications include generating free-viewpoint images, surveying using drones, 3D modeling for creating CG assets, and digital twins, which create environments identical to the real world in a virtual world.

[0003] Japanese Patent Application Laid-Open No. 2020-035388

[0004] When reconstructing a three-dimensional shape, it is necessary to estimate the camera pose from two-dimensional images. The camera pose includes information about the relative position (viewpoint) and pose (line of sight) of the subject and the camera. The camera pose is calculated by applying a group of feature points extracted from the camera image to the PnP problem. However, optical characteristics of the subject and occlusion can sometimes hinder accurate camera pose acquisition.

[0005] Therefore, the present disclosure proposes a three-dimensional data generation device, generation method, program, and computer-readable non-transitory storage medium that can improve the estimation accuracy of the camera posture.

[0006] According to the present disclosure, there is provided a 3D data generation device having a subject area recognition unit that detects a subject area from an image captured by a camera, a peripheral feature point extraction unit that extracts a group of feature points peripheral to the subject area from a group of feature points included in the image as a group of peripheral feature points, and a camera posture estimation unit that estimates the camera posture by selectively using the group of peripheral feature points.The present disclosure also provides a 3D data generation method in which information processing of the generation device is executed by a computer, a program that causes a computer to realize the information processing of the generation device, and a computer-readable non-transitory storage medium that stores the program.

[0007] 1 is a diagram illustrating a method for estimating a camera posture. FIG. 1 is a diagram illustrating an example of background occlusion by a subject. FIG. 2 is a diagram illustrating an example of reflected light becoming an obstacle. FIG. 3 is a diagram illustrating an example of light refraction becoming an obstacle. FIG. 4 is a diagram illustrating an example of lens distortion becoming an obstacle. FIG. 5 is a diagram illustrating an image in which the subject region is masked. FIG. 6 is a diagram illustrating an example of the configuration of an information processing device. FIG. 7 is a diagram illustrating an example of a processing flow related to a camera posture estimation process. FIG. 8 is a diagram illustrating an example of shooting using markers. FIG. 9 is a diagram illustrating an example of shooting using markers. FIG. 10 is a diagram illustrating an example of shooting using markers. FIG. 11 is a diagram illustrating an example of setting ring-shaped markers so as to border the periphery of a shooting stand. FIG. 12 is a diagram illustrating another example of shooting using ring-shaped markers. FIG. 13 is a diagram illustrating an example of the arrangement of a subject region and a camera posture detection region. FIG. 14 is a diagram illustrating an example of multiple markers being discretely placed around the periphery of a subject region. FIG. 15 is a diagram illustrating an example of multiple markers being discretely placed around the periphery of a subject region. FIG. 16 is a diagram illustrating an example of a calibration jig. FIG. 17 is a diagram illustrating another example of a calibration jig. FIG. 18 is a diagram illustrating an example of a conversion table creation procedure. FIG. 19 is a diagram illustrating an example of a conversion table creation procedure. FIG. 19 is a diagram illustrating a camera posture estimation process using a conversion table. FIG. 19 is a diagram illustrating an example of a method for estimating a camera posture for a camera located in a position where it is difficult to detect a group of peripheral feature points. FIG. 2 illustrates an example of a hardware configuration of an information processing device.

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0009] The description will be given in the following order: [1. Camera pose estimation method] [2. Causes of image analysis failure] [3. Estimation method of the present disclosure] [3-1. Example configuration of information processing device] [3-2. Processing flow] [4. Photographing a subject using markers] [5. Creation of a conversion table using a calibration jig] [6. Estimation of camera pose for a camera located in a position where it is difficult to detect a group of peripheral feature points] [7. Example hardware configuration] [8. Effects]

[0010] 1. Camera Pose Estimation Method FIG. 1 is a diagram illustrating a camera posture estimation method.

[0011] To estimate the camera pose, multiple images IM with different viewpoints are required. The camera pose includes information about the relative position (viewpoint) and attitude (line of sight) of the subject SB and the camera CM. In the example of FIG. 1, the subject SB placed on a plane is photographed from three viewpoints. The subject SB is, for example, a gift box. In the example of FIG. 1, there are three viewpoints, but the number of viewpoints is not limited to three. The more viewpoints there are, the higher the reproducibility of the three-dimensional shape becomes.

[0012] The camera pose is calculated by applying a group of feature points extracted from the image IM of the camera CM to a PnP (Perspective-n-Point) problem. PnP refers to a technique for determining the external parameters of the camera CM based on the positional relationship of the point groups between the images IM. Feature points are extracted from the subject SB or the entire image including the subject SB. The camera pose of each image IM is estimated by applying the images IM from each viewpoint (image IM1, image IM2, image IM3) to the PnP problem.

[0013] For example, a three-dimensional point P(X w , Y w , Z w ) is a point P in the camera coordinate system using external parameters R and T. c (X c , Y c , Z c ) where the external parameter R represents a rotation matrix. The external parameter T represents a translation vector. The camera pose can be formulated using the external parameters R and T. The image from the camera CM may contain lens distortion, and the camera pose can also be formulated taking the lens distortion into account.

[0014] [2. Obstacles to Image Analysis] As mentioned above, feature points are extracted from the subject SB or the entire image including the subject SB. However, due to obstacles such as the optical characteristics of the subject SB and occlusion, an accurate camera posture may not be obtained. For example, in the example of FIG. 1, feature points on the foreground of the subject SB are hidden from the viewpoint behind it by the ribbon. Occlusion OC becomes an obstacle when matching feature points between images IM. Occlusion OC also occurs in the background. FIG. 2 is a diagram showing an example of background occlusion OC caused by the subject SB.

[0015] Occlusion OC is not the only obstructive factor. The optical characteristics of the subject's body SB may also be obstructive, making it difficult to match feature points. Figure 3 shows an example in which reflected light is an obstructive factor. In an object with high reflectivity, a bright spot BS caused by reflected light may be recognized as a feature point. The appearance of the bright spot BS changes depending on the viewpoint. Depending on the viewing angle, the bright spot BS may become invisible or its position may change. Therefore, when a bright spot BS is detected as a feature point, feature points that cannot be matched between images IM may occur.

[0016] The same thing happens with refraction of light. Figure 4 shows an example where refraction of light becomes an obstruction factor. Refraction materials such as glass or plastic distort the shape of background objects seen through them. The positions of feature points within the refraction region RF change depending on the viewing angle, making it difficult to match feature points between images IM.

[0017] Lens distortion also causes distortion of an object. Figure 5 shows an example in which lens distortion DR becomes an obstacle. Lens distortion DR is largely eliminated by correction, but some distortion remains as a residue. The cause of the residue is not fully understood, but the residue often remains in a concentric circle around the center of the image IM. In Figure 5, there is little residue near the dashed line, and much residue near the solid line. When lens distortion DR is present, an error occurs in the conversion between feature points in a two-dimensional image and feature points in three dimensions. This reduces the accuracy of camera pose estimation.

[0018] It is also possible to improve the accuracy of posture estimation by attaching markers to the subject SB. However, the appearance of the subject SB changes depending on the marker, so a correction process is required. A method called structured light, which projects a pattern that serves as a landmark using laser light or the like, is also known. However, if the subject SB is a transparent object, the laser light will pass through, and if the subject SB is a reflective object, the light will be specularly reflected and will not be captured by the camera CM. If the subject SB is black, the laser light will also be absorbed by the subject SB and will not be captured by the camera CM.

[0019] [3. Estimation Method of the Present Disclosure] The present disclosure has been made in consideration of the above-mentioned problems. As described above, a decrease in the estimation accuracy of the camera pose may be due to the optical characteristics of the subject SB, occlusion OC, or the like. Therefore, in the present disclosure, an image area in which the subject SB appears is identified as a subject area SA, and the camera pose is estimated from a group of feature points in an image area other than the subject area SA. Figure 6 is a diagram showing an image IM in which the subject area SA is masked. The estimation method of the present disclosure will be described in detail below.

[0020] [3-1. Example configuration of information processing device] The information processing of the present disclosure is realized, for example, by an information processing device 1 shown in FIG. 7. FIG. 7 is a diagram showing an example configuration of the information processing device 1. The information processing device 1 is a 3D data generation device that generates data necessary for 3D reconstruction from images IM from multiple viewpoints. The information processing device 1 has an imaging unit 10, a feature point detection unit 20, a subject area recognition unit 30, a peripheral feature point extraction unit 40, and a camera posture estimation unit 50.

[0021] The imaging unit 10 controls the camera CM to capture images of the subject SB from multiple viewpoints. The feature point detection unit 20 acquires images IM of the subject SB from each viewpoint from the imaging unit 10. The feature point detection unit 20 acquires a group of feature points FT in the image IM for each viewpoint. For example, feature points include points (edges) where brightness or color changes significantly. Feature point extraction can be performed using known image analysis techniques.

[0022] The subject area recognition unit 30 detects a subject area SA from the image IM of the camera CM. The subject area SA refers to the area in the image IM in which the subject SB actually appears, or the area in the image IM in which the subject SB is expected to appear. For example, in three-dimensional reconstruction, the position and orientation of the camera CM that captures the subject SB are predetermined. If the background of the subject SB is captured in advance, the subject area SA can be identified by detecting the difference between an image of only the background and the image IM that includes the subject SB.

[0023] The subject SB is usually photographed with the subject SB positioned at the center of the angle of view. The size of the subject SB in the image IM is determined according to the distance between the camera CM and the subject SB. Therefore, an area in the center of the image IM large enough to encompass the subject SB can be set as the subject area SA. Alternatively, by using an object recognition technique such as segmentation, a segment corresponding to the subject SB can be acquired as the subject area SA.

[0024] The peripheral feature point extraction unit 40 extracts a group of feature points in the periphery of the subject area SA from the group of feature points FT included in the image IM as a peripheral feature point group SF. The peripheral feature point group SF is used as a calculation target when estimating the camera posture PS. The feature point group included in the subject area SA is excluded from the calculation target. The camera posture estimation unit 50 selectively uses the peripheral feature point group SF to estimate the camera posture PS.

[0025] The surrounding feature point group SF may be acquired from the pattern on the floor surface or from markers MK (see FIG. 10) for pose estimation placed on the floor surface or the like. In the latter case, the area where the markers MK are placed is used as the camera posture detection area PD (see FIG. 14). The camera posture detection area PD refers to an area where a feature point group used to estimate the camera posture PS exists. The surrounding feature point extraction unit 40 can extract the feature point group of the markers MK that indicate the camera posture detection area PD as the surrounding feature point group SF.

[0026] 3-2. Processing Flow FIG. 8 is a diagram showing an example of a processing flow relating to the processing for estimating the camera attitude PS.

[0027] The imaging unit 10 captures images of a subject SB from multiple viewpoints (step S1). The subject area recognition unit 30 detects a subject area SA from each image IM (step S2). The feature point detection unit 20 detects a group of feature points FT from the image IM from each viewpoint (step S3).

[0028] The peripheral feature point extraction unit 40 extracts a group of feature points (peripheral feature point group SF) in the periphery of the subject area SA from the detected feature point group FT (step S4). The camera posture estimation unit 50 selectively uses the peripheral feature point group SF to estimate the camera posture PS for each image IM (step S5).

[0029] When extracting the peripheral feature point group SF, the feature points of markers MK placed in the periphery of the subject area SA may be extracted, as described below. Alternatively, the feature points included in the subject area SA may be detected and then excluded, and all or part of the remaining part may be extracted as the peripheral feature point group SF. The latter method is employed when extracting a feature point group from a background image (an image of a floor, wall, etc.) without markers MK. In this example, the peripheral feature point extraction unit 40 extracts the peripheral feature point group SF by excluding the feature points of the subject area SA from the feature points included in the image IM of the camera CM.

[0030] The information processing device 1 reconstructs three-dimensional data of the subject SB using well-known photogrammetry techniques or the like from the data of the camera posture PS of each image IM estimated in this manner and the image data of the subject area SA of each image IM.

[0031] The three-dimensional data reconstructed in this manner is stored in a non-transitory storage medium such as a semiconductor memory element such as a flash memory, a hard disk, or an optical disk (not shown) connected to the information processing device 1. The three-dimensional data may be copied to a non-transitory storage medium in a remote location, for example, via another storage medium or an online communication line. In this way, a three-dimensional data storage medium is manufactured.

[0032] A non-transitory storage medium (three-dimensional data storage medium) storing three-dimensional data may be connected to a known playback device, such as a head-mounted display, or a computer or media player equipped with a display device such as a display or large screen, to form an image playback device. By writing the three-dimensional data generated by the present technology to the non-transitory storage medium, an image playback device (three-dimensional data playback device) capable of playing images based on the three-dimensional data can be manufactured. This makes it possible to provide content such as games to viewers, or to provide the image as a background image when producing other video content such as movies.

[0033] 4. Photographing a Subject Using Markers FIGS. 9 to 11 are diagrams illustrating an example of photographing a subject using markers MK.

[0034] As shown in Figure 9, the subject SB can be photographed on a photographing stand ST. A rotary table such as a turntable can be used as the photographing stand ST. By attaching multiple cameras CM to supports PL installed near the photographing stand ST and photographing while rotating the photographing stand ST, photographing can be performed from all directions (360°).

[0035] In this case, the background other than the imaging stand ST does not move. Therefore, the camera posture PS cannot be determined using only the feature points of the background other than the imaging stand ST. On the other hand, because the imaging stand ST rotates together with the subject SB, the camera posture PS can be estimated by using the feature point group on the imaging stand ST. Therefore, it is conceivable to place markers MK on the periphery of the imaging stand ST and use the feature point group of the markers MK as peripheral feature points SF (see FIG. 10 ).

[0036] When using markers MK, it is preferable that the markers MK are arranged in an area of ​​the image IM of the camera CM that does not overlap with the subject area SA. For example, the markers MK are arranged on the periphery of the imaging table ST so that the subject SB does not overlap part or all of the markers MK during imaging even when the imaging table ST and the subject SB rotate.

[0037] The markers MK only need to be placed on the periphery of the subject area SA. Even without rotating the subject SB, if the cameras CM are positioned in all 360° directions around the subject SB and photographing is performed, photographing similar to the example in Fig. 10 can be performed (see Fig. 11). In this case, if the markers MK are placed on the periphery of the subject area SA, the camera orientation PS can be estimated with high accuracy using the feature point group of the markers MK.

[0038] The marker MK may include a geometric pattern (designated marker) based on a specified rule, such as April Tag or Aruco. The peripheral feature point extraction unit 40 extracts a group of feature points included in the marker MK as a group of peripheral feature points. The camera posture estimation unit 50 can accurately estimate the identifier and posture of the marker MK based on the arrangement of the group of peripheral feature points.

[0039] The designated marker MK is provided as a square pattern surrounded by a black frame, but a marker MK with a two-dimensional spread may be generated by combining multiple designated markers. Fig. 12 is a diagram showing an example in which a ring-shaped marker MK is placed around the outer periphery of the imaging table ST. In the example of Fig. 12, a marker MK called ChArUco, which is a combination of a chessboard and ArUco, is used.

[0040] The ring-shaped marker MK can be used even without being combined with a turntable. Even without rotating the subject SB, if the camera CM is positioned in all 360° directions around the subject SB and photographing is performed, photographing similar to the example in FIG. 12 can be performed. FIG. 13 is a diagram showing another photographing example using a ring-shaped marker MK. In this case, if the ring-shaped marker MK is placed on the periphery of the subject area SA, the camera posture PS can be estimated with high accuracy using the feature point group of the marker MK.

[0041] The peripheral feature point extraction unit 40 identifies an installation area of ​​the marker MK as a camera posture detection area PD. The peripheral feature point extraction unit 40 acquires a group of feature points of the marker MK included in the camera posture detection area PD as a group of peripheral feature points. FIG. 14 is a diagram showing an example of the arrangement of the subject area SA and the camera posture detection area PD. In FIG. 14, the subject area SA is set as a cylindrical area that can include the subject SB. The camera posture detection area PD is set as a ring-shaped area surrounding the subject area SA.

[0042] A continuous ring-shaped marker MK may be placed in the camera posture detection area PD (see Figures 12 and 13), or multiple discretely arranged designated markers may be placed (see Figures 10 and 11).

[0043] 15 and 16 are diagrams showing an example in which multiple markers MK are discretely arranged around the periphery of the subject area SA. In this example, multiple isolated markers MK are arranged concentrically around the subject area SA. Each area in which the markers MK are arranged serves as a camera attitude detection area PD. By discretely arranging the markers MK, it is possible to prevent the markers MK from obscuring the subject SB when photographing from a low position (see FIG. 16).

[0044] 5. Creating a Conversion Table Using a Calibration Jig As described above, in the method disclosed herein, feature points in the subject area SA are excluded from the calculation. If feature points are not located near the subject SB, the accuracy of estimating the camera posture PS near the subject may decrease. Therefore, it may be possible to correct the camera posture PS estimated based on the surrounding feature point group SF using a conversion table.

[0045] The conversion table is a table for calibrating the error between the camera posture PS estimated from the feature point group of the subject area SA and the camera posture PS estimated from the peripheral feature point group SF. The conversion table converts the external parameters R and T of the camera CM so as to reduce the reprojection error of the feature point group of the subject area SA. The camera posture estimation unit 50 uses the conversion table to correct the camera posture PS estimated from the peripheral feature point group SF.

[0046] The conversion table can be created using a calibration jig CJ (see FIG. 17 ). FIG. 17 is a diagram showing an example of the calibration jig CJ. The calibration jig CJ is an object having a size similar to that of the subject SB and having a marker MK attached to its surface. In the example of FIG. 17 , the calibration jig CJ is configured such that a marker MK called ChArUco is attached to the surface of a cube, but the configuration of the calibration jig CJ is not limited to this. FIG. 18 is a diagram showing another example of the calibration jig CJ. In the example of FIG. 18 , a calibration jig CJ having a marker MK attached to the surface of a thin plate is illustrated.

[0047] 19 and 20 are diagrams showing an example of a procedure for creating a conversion table.

[0048] The imaging unit 10 photographs the calibration jig CJ from the same viewpoint as when photographing the subject SB. The photographing can be performed using the imaging stand ST and camera CM shown in FIG. 12. The imaging unit 10 photographs the calibration jig CJ by placing it in the center of the imaging stand ST. The imaging unit 10 photographs the calibration jig CJ with the camera CM installed near the imaging stand ST while rotating the imaging stand ST.

[0049] The feature point detection unit 20 detects feature points from the captured calibration images IM for each viewpoint. The extracted feature point group FT includes a feature point group of the markers MK of the calibration jig CJ (a jig feature point group CF) and a feature point group of the markers MK in the camera attitude detection area PD (a peripheral feature point group SF). The peripheral feature point extraction unit 40 acquires identifiers of the markers MK based on the arrangement of the feature points. The peripheral feature point extraction unit 40 distinguishes between the jig feature point group CF and the peripheral feature point group SF based on the identifiers.

[0050] The camera posture estimation unit 50 estimates the camera posture PS separately based on each of the feature point groups. The camera posture estimation unit 50 acquires, as a conversion table, a correspondence relationship between the camera posture PS estimated from the feature point group (jig feature point group CF) of the calibration jig CJ and the camera posture PS estimated from the feature point group (peripheral feature point group SF) of the peripheral part of the imaging table ST.

[0051] For example, the peripheral feature point extraction unit 40 extracts a fixture feature point group CF from the feature point group FT in the image IM (see the upper side of FIG. 19 ). The camera posture estimation unit 50 selectively uses the fixture feature point group CF to detect the camera posture PS (external parameters Rx, Tx). Next, the peripheral feature point extraction unit 40 extracts a peripheral feature point group SF from the feature point group FT in the image IM (see the lower side of FIG. 19 ). The camera posture estimation unit 50 selectively uses the peripheral feature point group SF to detect the camera posture PS (external parameters Rc, Tc).

[0052] In Fig. 20, external parameters Rx and Tx represent the camera posture PS estimated based on the jig feature point group CF. External parameters Rc and Tc represent the camera posture PS estimated based on the peripheral feature point group SF. External parameters Rx and Rc represent the rotation matrix of the camera CM. External parameters Tx and Tc represent the translation vector of the camera CM. The relationship between the external parameters Rx and Tx and the external parameters Rc and Tc is expressed by the following formula (1).

[0053]

[0054] By converting equation (1) into equation (2) below, conversion parameters Ra and Ta for converting external parameters Rc and Tc into external parameters Rx and Tx are obtained.

[0055]

[0056] The above calculation is performed for all images IM (photographing viewpoints). The camera posture estimation unit 50 acquires the correspondence between the viewpoint and the conversion parameters Ra and Ta as a conversion table. The conversion table defines the correspondence between the camera posture PS estimated from the feature point group (jig feature point group CF) of the calibration jig CJ installed in the center of the photographing table ST and the camera posture PS estimated from the feature point group (peripheral feature point group SF) of the peripheral part of the photographing table ST.

[0057] FIG. 21 is a diagram illustrating the process of estimating the camera attitude PS using a conversion table.

[0058] The user places the subject SB on the imaging stand ST instead of the calibration jig CJ. The imaging unit 10 places the subject SB in the center of the imaging stand ST and captures it. The feature point detection unit 20 detects a group of feature points FT from the captured images IM from each viewpoint. The peripheral feature point extraction unit 40 extracts a group of feature points in the periphery of the subject area SA from the detected group of feature points FT as a group of peripheral feature points SF. The camera posture estimation unit 50 selectively uses the group of peripheral feature points SF to estimate the camera posture PS.

[0059] The camera posture estimation unit 50 uses a conversion table to correct the camera posture PS estimated from the peripheral feature point group SF of the imaging stand ST. The camera posture estimation unit 50 associates the corrected camera posture PS with an image of the subject SB and stores them in a database. The user uses the data stored in the database to perform various processes (such as generating free-viewpoint images, surveying, 3D modeling for creating CG assets, and digital twins).

[0060] In the above example, the transformation parameters Ra and Ta were calculated for all viewpoints (calibration images IM) of the camera CM. However, if the transformation parameters Ra and Ta do not vary significantly depending on the viewpoint, the number of viewpoints for which calculations are performed can be reduced. This reduces the data size of the transformation table. The transformation parameters Ra and Ta for omitted viewpoints can be calculated by interpolation.

[0061] 6. Estimation of Camera Posture for Camera Located in Position Where Peripheral Feature Point Group is Difficult to Detect FIG. 22 is a diagram illustrating an example of a method for estimating the camera posture PS of a camera CM located in a position where it is difficult to detect the peripheral feature point group SF.

[0062] The camera CM installed at the bottom of the support PL is located close to the surface on which the surrounding feature point group SF is arranged. The camera CM installed at such a position (inferior camera CM) A ) makes it difficult to accurately detect the surrounding feature points SF.

[0063] In such a case, the camera CM located at a position where the group of peripheral feature points SF can be easily detected is designated as the reference camera CM. B and the reference camera CM B and inferior camera commercial ABased on the positional relationship with the reference camera CM B It is possible to convert the camera posture PS of the inferior camera CM. A The accuracy of estimating the camera posture PS at the reference camera CM is improved. The placement of each camera CM is set in advance by the user, and the reference camera CM is B and inferior camera commercial A The positional relationship with

[0064] For example, the camera posture estimation unit 50 may set a specific camera CM as a reference camera CM. B The camera posture estimation unit 50 sets the reference camera CM B The camera CM located at a position where it is more difficult to detect the surrounding feature point group SF than the camera CM is called the inferior camera CM. A The camera posture estimation unit 50 sets the reference camera CM B Based on the surrounding feature point group SF extracted from the image IM of the reference camera CM B The camera posture estimation unit 50 estimates the camera posture PS (reference posture) of the reference camera CM. B and inferior camera commercial A By converting the reference posture based on the positional relationship with the inferior camera CM A The camera pose PS is estimated.

[0065] The method of selecting the reference camera is arbitrary. A If the camera CM is located in a position where it is easier to detect the surrounding feature point group SF than the reference camera CM, any camera CM can be used. B In the example of FIG. 22, the camera CM located directly above the subject area SA is the reference camera CM. B The camera CM at this position can capture all of the surrounding feature points SF in the image IM without being blocked by the subject SB.

[0066] 7. Example of Hardware Configuration FIG. 23 is a diagram illustrating an example of the hardware configuration of the information processing device 1. As shown in FIG.

[0067] The information processing of the information processing device 1 is realized by, for example, a computer 1000. The computer 1000 has a CPU (Central Processing Unit) 1100, a RAM (Random Access Memory) 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0068] The CPU 1100 operates and controls each component based on a program (program data 1450) stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the program stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0069] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 starts up, as well as programs that depend on the hardware of the computer 1000 .

[0070] The HDD 1400 is a non-transitory computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the HDD 1400 is a recording medium that records an information processing program according to an embodiment as an example of program data 1450.

[0071] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0072] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display device, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.

[0073] For example, when the computer 1000 functions as the information processing device 1 according to the embodiment, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200 to realize the functions of the above-mentioned units. That is, each unit, such as the image capture unit 10 and the feature point detection unit 20, may be realized by the common CPU 1100 executing the information processing program, or may be realized by multiple CPUs sharing and executing the information processing program. When the computer 1000 is configured with multiple CPUs, each unit does not necessarily have to correspond one-to-one to each CPU, and the multiple CPUs may cooperate to realize the functions of the multiple units.

[0074] The information processing program, various models, and various data according to the present disclosure are stored in HDD 1400. CPU 1100 reads and executes program data 1450 from HDD 1400, but as another example, these programs may be obtained from other devices via external network 1550.

[0075] [8. Effects] The information processing device 1 has a subject area recognition unit 30, a peripheral feature point extraction unit 40, and a camera posture estimation unit 50. The subject area recognition unit 30 detects a subject area SA from an image IM of a camera CM. The peripheral feature point extraction unit 40 extracts a group of feature points in the periphery of the subject area SA from a group of feature points FT included in the image IM as a peripheral feature point group SF. The camera posture estimation unit 50 estimates the camera posture PS by selectively using the peripheral feature point group SF. In the three-dimensional data generation method of the present disclosure, the processing of the information processing device 1 is executed by a computer 1000. A program of the present disclosure causes the computer 1000 to realize the processing of the information processing device 1. A computer-readable non-transitory storage medium of the present disclosure stores a program that causes the computer 1000 to realize the processing of the information processing device 1.

[0076] This configuration eliminates obstacles caused by the optical characteristics of the subject SB, occlusion OC, etc., and therefore improves the accuracy of estimating the camera attitude PS.

[0077] The camera posture estimation unit 50 corrects the camera posture PS estimated from the peripheral feature point group SF using a conversion table for calibrating the error between the camera posture PS estimated from the feature point group of the subject area SA and the camera posture PS estimated from the peripheral feature point group SF.

[0078] This configuration makes it possible to suppress a decrease in estimation accuracy when using a group of feature points (surrounding feature points SF) located away from the subject area SA.

[0079] The conversion table defines the correspondence between the camera posture PS estimated from the feature point group (jig feature point group CF) of the calibration jig CJ installed in the center of the shooting table ST and the camera posture PS estimated from the peripheral feature point group SF of the periphery of the shooting table ST.

[0080] According to this configuration, the camera attitude PS is estimated with high accuracy based on a prior measurement using the calibration jig CJ.

[0081] The peripheral feature point extraction unit 40 extracts a group of feature points of the markers MK indicating the camera attitude detection area PD as a group of peripheral feature points SF.

[0082] According to this configuration, the camera attitude PS is estimated with high accuracy based on the marker MK.

[0083] The markers MK are arranged in an area of ​​the image IM that does not overlap with the subject area SA.

[0084] According to this configuration, the feature points of the marker MK are reliably acquired.

[0085] The peripheral feature point extraction unit 40 extracts a peripheral feature point group SF by excluding the feature point group of the subject area SA from the feature point group contained in the image IM.

[0086] According to this configuration, the surrounding feature points SF can be reliably acquired.

[0087] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0088] [Additional Notes] The present technology may also be configured as follows. (1) A three-dimensional data generation device comprising: a subject area recognition unit that detects a subject area from an image captured by a camera; a peripheral feature point extraction unit that extracts a group of feature points peripheral to the subject area from a group of feature points included in the image as a group of peripheral feature points; and a camera attitude estimation unit that estimates a camera attitude by selectively using the group of peripheral feature points. (2) The three-dimensional data generation device described in (1), wherein the camera attitude estimation unit corrects the camera attitude estimated from the group of peripheral feature points using a conversion table that calibrates an error between the camera attitude estimated from the group of feature points of the subject area and the camera attitude estimated from the group of peripheral feature points. (3) The three-dimensional data generation device described in (2), wherein the conversion table defines a correspondence between the camera attitude estimated from the feature point group of a calibration jig installed in the center of a shooting stand and the camera attitude estimated from the group of peripheral feature points on the periphery of the shooting stand. (4) The three-dimensional data generation device according to any one of (1) to (3), wherein the peripheral feature point extraction unit extracts a feature point group of a marker indicating a camera posture detection area as the peripheral feature point group. (5) The three-dimensional data generation device according to (4), wherein the marker is arranged in an area in the image that does not overlap with the subject area. (6) The three-dimensional data generation device according to any one of (1) to (3), wherein the peripheral feature point extraction unit extracts the peripheral feature point group by excluding the feature point group of the subject area from the feature point group included in the image. (7) The three-dimensional data generation device according to any one of (1) to (6) above, wherein the camera posture estimation unit sets a specific camera as a reference camera, sets a camera located in a position where it is more difficult to detect the group of peripheral feature points than the reference camera as an inferior camera, estimates the camera posture of the reference camera based on the group of peripheral feature points extracted from the image of the reference camera, and estimates the camera posture of the inferior camera by converting the camera posture of the reference camera based on the positional relationship between the reference camera and the inferior camera.(8) A computer-executed method for generating three-dimensional data, the method comprising: detecting a subject area from an image captured by a camera; extracting a group of feature points peripheral to the subject area from a group of feature points included in the image as a group of peripheral feature points; and estimating a camera posture by selectively using the group of peripheral feature points. (9) The method for generating three-dimensional data described in (8) above, the method comprising: correcting the camera posture estimated from the group of peripheral feature points using a conversion table for calibrating an error between the camera posture estimated from the group of feature points of the subject area and the camera posture estimated from the group of peripheral feature points. (10) The method for generating three-dimensional data described in (9) above, the method comprising: placing a calibration jig in the center of an imaging stand, and capturing an image, and obtaining, as the conversion table, a correspondence between the camera posture estimated from the feature point group of the calibration jig and the camera posture estimated from the group of peripheral feature points on the periphery of the imaging stand. (11) The method for generating three-dimensional data according to (10) above, comprising: placing a subject in the center of the imaging stand and imaging the subject; correcting the camera attitude estimated from the group of peripheral feature points of the imaging stand using the conversion table; and linking the corrected camera attitude and an image of the subject and storing them in a database. (12) The method for generating three-dimensional data according to any one of (8) to (11) above, extracting a group of feature points of a marker indicating a camera attitude detection area as the group of peripheral feature points. (13) The method for generating three-dimensional data according to (12) above, wherein the marker is placed in an area in the image that does not overlap with the subject area. (14) The method for generating three-dimensional data according to any one of (8) to (11) above, extracting the group of peripheral feature points by excluding the group of feature points of the subject area from the group of feature points included in the image. (15) A program that causes a computer to detect a subject area from an image captured by a camera, extract a group of feature points in the periphery of the subject area from a group of feature points included in the image as a group of peripheral feature points, and estimate a camera posture by selectively using the group of peripheral feature points. (16) A computer-readable non-transitory storage medium that stores the program according to (15) above.(17) A computer-executed method for generating three-dimensional data, the method comprising: detecting a subject area from an image captured by a camera; extracting a group of feature points in a peripheral portion of the subject area from a group of feature points included in the image as a group of peripheral feature points; estimating a camera posture by selectively using the group of peripheral feature points; and reconstructing three-dimensional data of the subject based on information about the estimated camera posture and information about the subject area. (18) A method for manufacturing a three-dimensional data storage medium, the method comprising: detecting a subject area from an image captured by a camera; extracting a group of feature points in a peripheral portion of the subject area from a group of feature points included in the image as a group of peripheral feature points; estimating a camera posture by selectively using the group of peripheral feature points; reconstructing three-dimensional data of the subject based on information about the estimated camera posture and information about the subject area; and writing the reconstructed three-dimensional data to a non-transitory storage medium. (19) A method for manufacturing a three-dimensional data reproduction device comprising a non-transitory storage medium for storing three-dimensional data and a reproduction device for reproducing the three-dimensional data stored in the non-transitory storage medium, the method comprising: detecting a subject area from an image of a camera; extracting a group of feature points in the periphery of the subject area as a group of peripheral feature points from a group of feature points included in the image; estimating the camera posture by selectively using the group of peripheral feature points; reconstructing three-dimensional data of the subject based on information on the estimated camera posture and information on the subject area; and writing the reconstructed three-dimensional data to the non-transitory storage medium.

[0089] 1 Information processing device (three-dimensional data generating device) 30 Subject area recognition unit 40 Peripheral feature point extraction unit 50 Camera attitude estimation unit CJ Calibration jig CM Camera CM A Inferior Camera CM B Reference camera FT Feature point group IM Image MK Marker PD Camera attitude detection area PS Camera attitude SA Subject area SF Surrounding feature point group ST Photography stand

Claims

1. A three-dimensional data generation device having: a subject area recognition unit that detects a subject area from a camera image; a peripheral feature point extraction unit that extracts a group of feature points in the periphery of the subject area as a group of peripheral feature points from a group of feature points included in the image; and a camera posture estimation unit that selectively uses the group of peripheral feature points to estimate the camera posture.

2. The 3D data generation device according to claim 1, wherein the camera posture estimation unit corrects the camera posture estimated from the group of peripheral feature points using a conversion table for calibrating the error between the camera posture estimated from the group of feature points in the subject area and the camera posture estimated from the group of peripheral feature points.

3. The three-dimensional data generation device according to claim 2, wherein the conversion table defines the correspondence between the camera posture estimated from the feature point group of a calibration jig installed in the center of the imaging table and the camera posture estimated from the peripheral feature point group of the periphery of the imaging table.

4. The three-dimensional data generating device according to claim 1, wherein the peripheral feature point extraction unit extracts a group of feature points of a marker indicating a camera posture detection area as the group of peripheral feature points.

5. The three-dimensional data generating device according to claim 4, wherein the markers are arranged in an area of ​​the image that does not overlap with the subject area.

6. The three-dimensional data generating device according to claim 1, wherein the peripheral feature point extraction unit extracts the group of peripheral feature points by excluding the group of feature points in the subject region from the group of feature points included in the image.

7. The three-dimensional data generation device of claim 1, wherein the camera posture estimation unit: sets a specific camera as a reference camera; sets a camera located in a position where it is more difficult to detect the group of peripheral feature points than the reference camera as an inferior camera; estimates the camera posture of the reference camera based on the group of peripheral feature points extracted from the image of the reference camera; and estimates the camera posture of the inferior camera by converting the camera posture of the reference camera based on the positional relationship between the reference camera and the inferior camera.

8. A computer-executable method for generating three-dimensional data, comprising: detecting a subject area from a camera image; extracting a group of feature points in the periphery of the subject area from a group of feature points included in the image as a group of peripheral feature points; and estimating the camera posture by selectively using the group of peripheral feature points.

9. The method for generating three-dimensional data according to claim 8, further comprising correcting the camera posture estimated from the group of peripheral feature points using a conversion table for calibrating an error between the camera posture estimated from the group of feature points of the subject area and the camera posture estimated from the group of peripheral feature points.

10. A method for generating three-dimensional data according to claim 9, comprising: placing a calibration jig in the center of an imaging stand, taking an image, and obtaining, as the conversion table, a correspondence between the camera posture estimated from the feature point group of the calibration jig and the camera posture estimated from the peripheral feature point group of the periphery of the imaging stand.

11. A method for generating three-dimensional data according to claim 10, comprising: placing a subject in the center of the imaging stand and taking an image; correcting the camera attitude estimated from the group of peripheral feature points of the imaging stand using the conversion table; and linking the corrected camera attitude with an image of the subject and storing them in a database.

12. The three-dimensional data generating method according to claim 8, wherein a group of feature points of a marker indicating a camera posture detection area is extracted as the group of peripheral feature points.

13. The three-dimensional data generating method according to claim 12, wherein the markers are arranged in an area of ​​the image that does not overlap with the subject area.

14. The three-dimensional data generating method according to claim 8, wherein the group of peripheral feature points is extracted by excluding the group of feature points in the subject region from the group of feature points included in the image.

15. A program that causes a computer to perform the following operations: detect a subject area from a camera image; extract a group of feature points in the periphery of the subject area from a group of feature points contained in the image as a group of peripheral feature points; and estimate the camera posture by selectively using the group of peripheral feature points.

16. A computer-readable non-transitory storage medium storing the program according to claim 15.

Citation Information

Patent Citations

  • Image evaluation device

    JP2011223265A

  • Three-dimensional reconfiguration system and three-dimensional reconfiguration method

    WO2023223916A1