Photographing position / posture estimation apparatus, system, method and program
Patent Information
- Application Number
- JP2024026437
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-26
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies for estimating the shooting position and orientation of a captured image using 3D mesh data without color information result in decreased accuracy due to the reliance on texture information.
A method that calculates shape features from 3D data, generates virtual image information based on virtual camera information, and calculates similarity between this information and the input image without using color information, enabling accurate estimation of the shooting position and orientation.
Maintains the estimation accuracy of the shooting position and orientation using 3D data without color information, reducing implementation costs by utilizing inexpensive systems like LiDAR or CAD data, and improving matching efficiency with depth-compromised representations.
Smart Images

Figure 2025129661000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a photographing position and orientation estimation device, system, method, and program. [Background technology]
[0002] Patent Document 1 discloses a technology related to a matching device that matches a captured image with 3D mesh data given texture information. The matching device according to Patent Document 1 converts input 3D mesh data into a 2D image captured from a reference camera posture, and calculates feature amounts of the converted 2D image and the input captured image (2D image). The matching device then compares the calculated feature amounts to calculate a similarity, and if the similarity is equal to or greater than a threshold, it estimates that the input captured image was captured from a shooting position and posture (hereinafter referred to as "shooting position and posture") at the reference camera posture. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-153664 Summary of the Invention [Problem to be solved by the invention]
[0004] The technology disclosed in Patent Document 1 is based on the premise that texture information (color information) is provided to each surface of the 3D mesh data. The technology disclosed in Patent Document 1 calculates the similarity with a captured image by converting the 3D mesh data into a 2D image using the color information and calculating feature amounts. Therefore, when using 3D data without color information, the technology disclosed in Patent Document 1 may result in a decrease in the accuracy of estimating the shooting position and orientation of the input captured image.
[0005] In view of the above-mentioned problems, an object of the present disclosure is to provide a shooting position and orientation estimation device, system, method, and program for maintaining the estimation accuracy of the shooting position and orientation of a captured image when using three-dimensional data without color information. [Means for solving the problem]
[0006] The imaging position and orientation estimation device according to the present disclosure comprises: a feature amount calculation means for calculating a shape feature amount from predetermined three-dimensional data; a generating means for generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position of the image area; a similarity calculation means for calculating a similarity between the virtual image information and an input image; Equipped with.
[0007] The imaging position and orientation estimation system according to the present disclosure includes: a photographing terminal and a photographing position and orientation estimation device communicably connected to the photographing terminal; the imaging position and attitude estimation device, Calculate shape features from the 3D data of a given object, generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on virtual camera information with each pixel position of the image area; When an image of the object photographed by the photographing terminal is received as an input image, a similarity between the virtual image information and the input image is calculated; estimating a photographing position and orientation in the input image based on the virtual camera information and the similarity; The estimated photographing position and orientation is returned to the photographing terminal.
[0008] The imaging position and orientation estimation method according to the present disclosure includes: The computer Calculate shape features from the specified 3D data, generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position of the image area; The similarity between the virtual image information and the input image is calculated.
[0009] The imaging position and orientation estimation program according to the present disclosure includes: A feature amount calculation process for calculating a shape feature amount from predetermined three-dimensional data; a generation process for generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position of the image area; a similarity calculation process for calculating a similarity between the virtual image information and an input image; to be executed by the computer. [Effects of the Invention]
[0010] According to the present disclosure, it is possible to maintain the estimation accuracy of the shooting position and orientation of a captured image when using three-dimensional data without color information. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a block diagram illustrating a configuration of a photographing position and orientation estimation device according to the present disclosure. [Figure 2] 1 is a flowchart illustrating a flow of a photographing position and orientation estimation method according to the present disclosure. [Figure 3] FIG. 1 is a block diagram illustrating a configuration of a photographing position and orientation estimation device according to the present disclosure. [Figure 4] FIG. 2 is a diagram for explaining the concept of the data structure of virtual image information according to the present disclosure. [Figure 5] 10 is a flowchart showing the flow of a photographing position and orientation estimation process according to the present disclosure. [Figure 6] 1 is a diagram for explaining the concept of a method for visualizing virtual image information according to the present disclosure. [Figure 7]FIG. 10 is a diagram illustrating an example display of pixel correspondences estimated by local matching according to the present disclosure. [Figure 8] 1 is a flowchart showing the flow of a learning process for an estimation model according to the present disclosure. [Figure 9] 1 is a flowchart showing the flow of a learning process for an estimation model according to the present disclosure. [Figure 10] 1 is a flowchart showing the flow of a learning process for an estimation model according to the present disclosure. [Figure 11] 10 is a flowchart showing the flow of a photographing position and orientation estimation process when a global matching process and a local matching process according to the present disclosure are used in combination. [Figure 12] FIG. 1 is a block diagram illustrating a hardware configuration of a photographing position and orientation estimation apparatus according to the present disclosure. [Figure 13] 10 is a flowchart showing the flow of a shooting position and orientation estimation process for two input images according to the present disclosure. [Figure 14] 1 is a block diagram showing a configuration of a photographing position and orientation estimation system according to the present disclosure. [Figure 15] 10 is a sequence chart showing the flow of a photographing position and orientation estimation process according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, the same or corresponding elements are designated by the same reference numerals, and for clarity of explanation, duplicate explanations will be omitted as necessary.
[0013] <Embodiment 1> FIG. 1 is a block diagram showing the configuration of a photographing position and orientation estimation device 1. The photographing position and orientation estimation device 1 is an information processing device for estimating a photographing position and orientation of a predetermined target object in 3D data created in advance from an image of the target object photographed by a camera. The photographing position and orientation estimation device 1 may also be an information processing device that learns a model for estimating the photographing position and orientation. The target object may be, for example, a structure such as a bridge or a collection of objects arranged in an indoor space. The photographing position and orientation estimation device 1 may also be used for inspecting the structure. The photographing position and orientation estimation device 1 includes a feature amount calculation unit 11, a generation unit 12, and a similarity calculation unit 13. The feature amount calculation unit 11, the generation unit 12, and the similarity calculation unit 13 may be used as means for calculating feature amounts and similarities, and means for generating virtual image information, respectively.
[0014] The feature calculation unit 11 calculates shape features from predetermined three-dimensional data. "Three-dimensional data" is data that represents the three-dimensional structure of a target object. For example, the three-dimensional data may be a set of data points (point cloud data) that represent the structure (at least the external shape) of the target object in a predetermined three-dimensional space using three-dimensional coordinates. Alternatively, the three-dimensional data may be mesh data, Computer Aided Design (CAD) data, Building Information Modeling (BIM) / Construction Information Modeling (CIM) data, implicit function representation data that can be created using Neural Radiance Fields (NeRF) technology, etc. Note that the three-dimensional data is not limited to these, as long as it represents the three-dimensional structure of the target object. In particular, the three-dimensional data according to the present disclosure may not include color information. "Shape features" are vector information that represents the shape characteristics of each data point in multiple dimensions based on its relationship with surrounding data points.
[0015] The generation unit 12 generates virtual image information by associating shape features projected onto an image area when imaging three-dimensional data based on predetermined virtual camera information with each pixel position of the image area. Here, "virtual camera information" refers to information about a virtual camera virtually installed at a predetermined position and orientation (photography angle) in three-dimensional space to photograph a target object in three-dimensional space. The virtual camera information includes at least three-dimensional coordinates (position) and orientation in three-dimensional space. The virtual camera information also includes the frame size of the image to be photographed, i.e., the number of pixels (number of vertical and horizontal pixels). The virtual camera information may also include the angle of view in three-dimensional space.
[0016] "Virtual image information" refers to information in which shape features are associated with each pixel position of an image generated when a target object is photographed based on virtual camera information. That is, when the generation unit 12 visualizes three-dimensional data based on predetermined virtual camera information, i.e., performs 3D rendering, the generation unit 12 uses shape features at data points corresponding to each pixel among the data points of the three-dimensional data as information equivalent to pixel values. In other words, when the generation unit 12 visualizes three-dimensional data based on virtual camera information, instead of projecting color information onto an image region, the generation unit 12 projects shape features of each data point of the three-dimensional data. The generation unit 12 then associates the projected shape features with each pixel position of the image region to generate virtual image information. An "image region" refers to a two-dimensional region (planar region) that is imaged when a target object is photographed with a virtual camera at a predetermined position and orientation in three-dimensional space. Shape features may be converted from shape features corresponding to multiple data points. Furthermore, shape features may be information obtained by performing a predetermined transformation on shape features calculated by the feature calculation unit 11.
[0017] The similarity calculation unit 13 calculates the similarity between the virtual image information and the input image. The similarity calculation unit 13 may calculate the similarity between the virtual image information and the input image by (M1) global matching. For example, the similarity calculation unit 13 may calculate the similarity between a set of pixel positions in the virtual image information (entire frame size) and the entire frame size of the input image. Alternatively, the similarity calculation unit 13 may calculate the similarity between the virtual image information and the input image by determining the correspondence between pixels by (M2) local matching. That is, the similarity calculation unit 13 may calculate the degree to which specific pixel positions in the virtual image information and the input image represent the same part of the object (the same data point in the three-dimensional data) as the similarity. Then, the similarity calculation unit 13 may calculate the similarity between the virtual image information and the input image by combining the similarities on a pixel-by-pixel basis.
[0018] 2 is a flowchart showing the flow of the imaging position and orientation estimation method. First, feature calculation unit 11 calculates shape feature values from predetermined three-dimensional data (S1). Next, generation unit 12 generates virtual image information by associating shape feature values projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position in the image area (S2). Then, similarity calculation unit 13 calculates the similarity between the virtual image information and the input image (S3).
[0019] Therefore, the photographing position and orientation estimation device 1 can estimate the photographing position and orientation of the input image from among a plurality of pieces of virtual camera information (candidate photographing positions and orientations) based on the similarity between the virtual image information and the input image. Here, the similarity calculation unit 13 uses shape features associated with each pixel position of the virtual image information, but does not use color information, when calculating the similarity. Therefore, the photographing position and orientation can be estimated using three-dimensional data without color information. Therefore, the same estimation accuracy as when color information is used can be maintained. In other words, the photographing position and orientation estimation device 1 according to the present disclosure can maintain the estimation accuracy of the photographing position and orientation of the photographed image when three-dimensional data without color information is used.
[0020] The photographing position and orientation estimation device 1 includes a processor, a memory, and a storage device (not shown). The storage device stores a computer program that implements the processing of the photographing position and orientation estimation method shown in Fig. 2, for example. The processor then loads the computer program from the storage device into the memory and executes the computer program. This allows the processor to implement the functions of a feature amount calculation unit 11, a generation unit 12, and a similarity calculation unit 13.
[0021] Alternatively, each component of the image capture position and orientation estimation device 1 may be realized by dedicated hardware. Furthermore, some or all of the components of each device may be realized by general-purpose or dedicated circuits, processors, etc., or a combination of these. These may be configured by a single chip, or by multiple chips connected via a bus. Some or all of the components of each device may be realized by a combination of the above-mentioned circuits, etc., and programs. Furthermore, a CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), quantum processor (quantum computer control chip), etc., may be used as the processor.
[0022] Furthermore, when some or all of the components of the photographing position and orientation estimation device 1 are realized by a plurality of information processing devices, circuits, etc., the plurality of information processing devices, circuits, etc. may be centrally or decentralized. For example, the information processing devices, circuits, etc. may be realized as a client-server system, a cloud computing system, or the like, in a form in which each is connected via a communication network. Furthermore, the functions of the photographing position and orientation estimation device 1 may be provided in a SaaS (Software as a Service) format.
[0023] <Embodiment 2> 3 is a block diagram showing the configuration of the photographing position and orientation estimation device 100. The photographing position and orientation estimation device 100 is an example of the above-mentioned photographing position and orientation estimation device 1. The photographing position and orientation estimation device 1 includes a storage unit 110, an acquisition unit 121, a feature calculation unit 122, a rendering unit 123, a matching unit 124, an estimation unit 125, a display unit 126, and a learning unit 127.
[0024] The storage unit 110 includes, for example, a non-volatile storage device such as a flash memory and a memory such as a RAM (Random Access Memory), i.e., a volatile storage device. The storage unit 110 stores three-dimensional data 111 and virtual camera information 112. As described above, the three-dimensional data 111 is data (3D (three dimensions) data) that represents the three-dimensional structure of a target object. The three-dimensional data 111 may be, for example, 3D data captured by a LiDAR (Light Detection and Ranging) system or 3D-CAD data (without color information) created in the design stage of a structure. The virtual camera information 112 is information similar to the virtual camera information in the first embodiment described above. The storage unit 110 stores two or more pieces of virtual camera information 112.
[0025] The acquisition unit 121 acquires, as an input image, an image obtained by capturing a target object corresponding to the three-dimensional data 111. The input image is an image that allows the imaging position and orientation estimation device 100 to estimate the imaging position and orientation on the three-dimensional data 111.
[0026] The feature calculation unit 122 is an example of the feature calculation unit 11 described above. The feature calculation unit 122 calculates a shape feature for each of multiple data points on the three-dimensional data 111. Here, the "shape feature" is not limited to a specific one as long as it is a feature that expresses the distribution of the three-dimensional data 111 around the data point. The shape feature is assumed to be higher-dimensional vector information compared to general color information (e.g., three-dimensional RGB). Therefore, the feature calculation unit 122 calculates the shape feature for each data point so as to express the distribution of other data points around a specific data point in the three-dimensional data 111. This allows the shape features of each data point to be accurately expressed using multidimensional vector information. For example, the feature calculation unit 122 may calculate the shape feature by quantifying the shape or direction of the distribution of data points in the three-dimensional data 111. Specifically, the feature calculation unit 122 may apply principal component analysis to the distribution of the three-dimensional data 111 around the data points, and quantify the shape or direction of the distribution of the data points in the three-dimensional data 111 based on the three calculated eigenvectors and three eigenvalues.
[0027] Alternatively, the feature calculation unit 122 may calculate the normal vector of each data point in the three-dimensional data 111 as the shape feature. Alternatively, the feature calculation unit 122 may calculate the shape feature using a predetermined trained model. For example, the first model is an AI (Artificial Intelligence) model that inputs three-dimensional data around a specific data point in the three-dimensional data 111 and outputs the shape feature of the data point. For example, PointNet or the like may be used as the first model. The first trained model may be machine-learned (e.g., deep learning) so that shape features are similar when the distributions of data points in the three-dimensional data 111 are similar to those of the first model. In this way, metric learning, which is an example of unsupervised learning, can efficiently improve the calculation accuracy of the shape feature. Alternatively, the second model may be an AI model that inputs three-dimensional data of a specific object, virtual camera information, and a captured image of the object and outputs the shooting position and orientation. In this case, it can be said that the second model internally calculates the shape feature of each data point. In this case, the second trained model may be an AI model that has been machine-learned using, as training data, three-dimensional data of a predetermined object, captured images of the object, and the shooting positions and orientations of the captured images. In this way, supervised learning can also efficiently improve the calculation accuracy of shape features. The AI model may also be called a deep learning model.
[0028] The rendering unit 123 is an example of the generating unit 12. The rendering unit 123 generates virtual image information by rendering the three-dimensional data 111 using shape features based on predetermined virtual camera information. In other words, the rendering unit 123 identifies a set of two-dimensional coordinates of an image area (plane) when a target object (the three-dimensional data 111) in three-dimensional space is assumed to be photographed from an arbitrary virtual camera shooting position and orientation, and generates virtual image information by mapping shape features to each point (pixel position) of the identified set of two-dimensional coordinates. The rendering unit 123 also generates multiple pieces of virtual image information corresponding to each piece of virtual camera information 112 based on each of the multiple pieces of virtual camera information. Here, the virtual image information can be expressed as a three-dimensional array of the height H and width W of the image area when visualized and the number of shape feature dimensions D. FIG. 4 is a diagram for explaining the concept of the data structure of the virtual image information d0. Here, the height H is the number of pixels in the height direction of the image area, and the width W is the number of pixels in the width direction of the image area. Note that the number of pixels in the height H and width W in FIG. 4 is merely an example. The number of shape feature dimensions D is the number of dimensions of the shape feature (feature vector) calculated by the feature calculation unit 122. The number of shape feature dimensions D may be, for example, 4 or more. In other words, the shape feature is a feature vector with 4 or more dimensions. Therefore, the virtual image information is information in which a shape feature (without color information) is associated with each pixel position (a pair of pixel position in the height H direction and pixel position in the width W direction). Therefore, the rendering unit 123 may improve the technology for rendering 3D data with color information and may implement the information by referring to the shape feature of the corresponding pixel position instead of referring to the color information of the 3D data.
[0029] Furthermore, the rendering unit 123 may convert the shape feature based on the shooting position and orientation included in the virtual camera information 112, and generate virtual image information using the converted shape feature. For example, if the shape feature calculated by the feature calculation unit 122 is a feature (e.g., a normal vector) that depends on the rotation or translation of the three-dimensional data 111, rendering the normal direction as is would reflect the absolute direction of the normal, such as westward or eastward. In this case, the direction cannot be determined from the input image that is the matching target for similarity calculation. Therefore, the rendering unit 123 converts the shape feature based on the shooting position and orientation of the virtual camera information 112 so as to remove the absolute shooting position and orientation information, and generates virtual image information using the converted shape feature. This allows for the generation of more accurate virtual image information that eliminates the influence of the absolute direction of the normal.
[0030] The matching unit 124 is an example of the above-mentioned similarity calculation unit 13. The matching unit 124 calculates the similarity between each of the plurality of virtual image information and the input image. When an image of an object corresponding to the three-dimensional data 111 is acquired as the input image, the matching unit 124 may calculate the similarity between the virtual image information and the input image.
[0031] The matching unit 124 performs either the above-described (M1) global matching or (M2) local matching, or both (M1) global matching and (M2) local matching. (M1) global matching globally compares the features of the entire image area of each of the multiple virtual image information with the input image to determine the similarity between each virtual image information and the input image. (M2) local matching compares the shape feature amount of each virtual image information with the color information of the input image on a pixel-by-pixel basis (i.e., locally) to calculate the similarity on a pixel-by-pixel basis and determine the correspondence between pixel positions. At this time, the matching unit 124 appropriately corrects the correspondence between pixel positions when performing matching. Here, (M1) global matching roughly estimates the shooting position and orientation of the input image compared to (M2) local matching. In other words, (M2) local matching can improve the estimation accuracy of the shooting position and orientation of the input image compared to (M1) global matching. On the other hand, (M1) global matching is faster than (M2) local matching. In other words, (M2) local matching requires more processing cost than (M1) global matching.
[0032] The matching unit 124 can be realized using a matching model that performs (M1) global matching and (M2) local matching. A typical matching model can be said to be an AI model that inputs two images having color information and outputs the similarity between the two images. In contrast, the matching unit 124 according to the present disclosure can use an AI model that can input virtual image information by changing the channel dimension of one input image to the dimension of shape features instead of color information, and leaves the channel dimension of the other input image (captured image) unchanged.
[0033] The estimation unit 125 estimates the shooting position and orientation in the input image based on the virtual camera information and the similarity. Specifically, the estimation unit 125 selects virtual image information with a higher similarity from among a plurality of pieces of virtual image information, and estimates the shooting position and orientation corresponding to the selected virtual image information as the shooting position and orientation in the input image.
[0034] The display unit 126 displays display information including the estimation result of the shooting position and orientation of the input image estimated based on the similarity. For example, the display unit 126 displays the display information by outputting it to a display device (not shown) built into the shooting position and orientation estimation device 100 or connected to the shooting position and orientation estimation device 100. Alternatively, the display unit 126 may transmit the display information to the shooting terminal that captured the input image, thereby displaying it on the screen of the shooting terminal. The display unit 126 may also display a virtual image in which shape features associated with each pixel of the virtual image information are converted into color information by dimensional compression. This allows the user to visually recognize the virtual image information without color information as a visualized image. The display unit 126 may also display the virtual image and the input image in comparison. This makes it easier for the user to visually recognize the correspondence between the virtual image and the input image, allowing the user to more accurately grasp the shooting position and orientation. The display unit 126 may also display the virtual image together with the estimation result. The display unit 126 may also display the estimation result, the virtual image, and the input image for comparison. This allows the user to more accurately grasp the shooting position and orientation. The display unit 126 may also function as an output unit that outputs the estimation result, virtual image information, the virtual image, the input image, etc.
[0035] The learning unit 127 performs machine learning on AI models such as a shape feature calculation model, a similarity calculation model, a first estimation model, or a second estimation model, and updates the model parameters. The shape feature calculation model is used in the processing of the feature calculation unit 122. The similarity calculation model is used in the processing of the matching unit 124. The first estimation model is used in the processing of the matching unit 124 and the estimation unit 125. The second estimation model may be used in the processing from the feature calculation unit 122 to the estimation unit 125. Note that the photographing position and orientation estimation device 100 according to the present disclosure may use any one of the above models. Alternatively, the photographing position and orientation estimation device 100 may use a combination of the shape feature calculation model and any one of the similarity calculation model and the first estimation model. Alternatively, the photographing position and orientation estimation device 100 may use the second estimation model.
[0036] FIG. 5 is a flowchart showing the flow of the photographing position and orientation estimation process. First, the feature amount calculation unit 122 calculates shape feature amounts at each data point of the three-dimensional data 111 (S11). Next, the rendering unit 123 renders the three-dimensional data 111 based on the virtual camera information 112 and associates the shape feature amounts with each pixel position to generate virtual image information d0 (S12). Then, the matching unit 124 calculates the similarity between the virtual image information d0 and the photographed image d2 (S13). Thereafter, the estimation unit 125 estimates the photographing position and orientation in the photographed image d2 based on the virtual camera information 112 and the similarity (and the three-dimensional data 111) (S14). Then, the display unit 126 outputs the estimation result of the photographing position and orientation (S15). Furthermore, the display unit 126 converts the virtual image information d0 into a visualized virtual image d1 (S16). Then, the display unit 126 displays the virtual image d1 and the photographed image d2 for comparison (S17).
[0037] <Virtualization method for virtual image information> FIG. 6 is a diagram illustrating the concept of a method for visualizing virtual image information d0. As described above, in the virtual image information d0, a feature vector with shape feature dimension number D is associated with each pixel position of a set of pixels with height H and width W. As described above, the height H and width W are merely examples, and the shape feature dimension number D is four or more. The display unit 126 converts the shape features into a virtual image d01 by performing a dimension reduction process. Here, the dimension reduction process is a process of converting a set of high-dimensional vectors into a set of low-dimensional vectors. Specifically, the dimension reduction process converts vectors with similar values in a high-dimensional space into similar values in a low-dimensional space. For example, principal component analysis (PCA), t-SNE (t-distributed stochastic neighbor embedding), etc. can be used for the dimension reduction process. The dimension reduction process may also be called a dimension compression process.
[0038] In the example of FIG. 6, the display unit 126 interprets the virtual image information d0 as consisting of D-dimensional vectors with a height of H and a width of W, applies dimension reduction processing, and converts it into a virtual image d01, which is a three-dimensional vector with a height of H and a width of W. Here, the converted three-dimensional vector represents color information, for example, RGB. This allows the virtual image information d0 to be converted into a color image format. However, the color information may be a three-dimensional vector other than RGB, or a vector of color information other than three-dimensional. At least the number of dimensions D of the shape feature is greater than the number of dimensions of the converted color information. The display unit 126 then displays the converted virtual image d1 on the screen. This allows the virtual image information d0 to be visualized. In this case, portions with similar colors in the visualized virtual image d1 by the dimension reduction processing represent similar values as the original shape feature. Note that the color information converted into a three-dimensional vector by the dimension reduction processing may represent the height direction of the pixel positions in the three-dimensional data using color or shade. For example, the color information may indicate that red is relatively high, blue is relatively low, green is intermediate between red and blue, and gray indicates some kind of object.
[0039] When dimension reduction is applied to each piece of virtual image information, images may be converted into images with different colors even if the rendering viewpoints are similar. For example, a wall may be green in one visualized virtual image, while it may be red in another virtual image. This is because the dimension reduction process attempts to maintain the distance between vectors before and after reduction, but the order of dimensions after reduction (the order of R, G, B) is indefinite.
[0040] Therefore, the image capture position and orientation estimation device 100 may generate a dimension reduction function using the 3D vector data after calculating the high-dimensional shape feature amount, and apply the function to the dimension reduction process for multiple pieces of virtual image information. This makes the virtual images generated and converted based on the virtual camera information from the close viewpoint similar in color.
[0041] That is, the photographing position and orientation estimation device 100 performs dimension reduction from D dimensions to 3 dimensions on 3D data consisting of N points of data each having a D-dimensional shape feature, and generates a function that "converts one D-dimensional data into one 3D data." At this time, the photographing position and orientation estimation device 100 can generate the function using PCA or t-SNE techniques. Then, the display unit 126 applies the function that "converts one D-dimensional data into one 3D data" to the virtual image information d0, converting the "height H × width W" D-dimensional vectors into "height H × width W" 3D vectors, and can generate the virtual image d1.
[0042] <Example of comparison between visualized virtual image and input image> 7 is a diagram showing an example of pixel correspondences estimated by local matching. Correspondence R1 indicates that pixel P11 in virtual image d1 and pixel P21 in captured image d2 are likely to be in the same location (high similarity). Correspondence R2 indicates that pixel P12 in virtual image d1 and pixel P22 in captured image d2 are likely to be in the same location (high similarity).
[0043] <About learning 1 (training) of estimation model> For example, the matching unit 124 may calculate the similarity using a third trained model that is machine-learned using the three-dimensional data of a predetermined object, the captured image of the object, and the shooting position and orientation of the captured image as training data. The third trained model is a machine-learned version of either the first estimation model or the second estimation model described above.
[0044] Fig. 8 is a flowchart showing the flow of the learning process for the first estimation model. The training data T1 includes training three-dimensional data T11, a captured image T12, and a position and orientation T13. The position and orientation T13 is correct data for the shooting position and orientation when the captured image T12 was captured. The learning process in Fig. 8 can also be applied to the second estimation model described above.
[0045] First, the feature calculation unit 122 calculates shape feature amounts at each data point of the teacher three-dimensional data T11 (S11). Next, the rendering unit 123 renders the three-dimensional data T11 based on the virtual camera information 112 and associates the shape feature amounts with each pixel position to generate virtual image information d0 (S12). Then, the matching unit 124 calculates the similarity between the virtual image information d0 and the captured image T12 using a first estimation model (S13). Thereafter, the estimation unit 125 estimates the shooting position and orientation in the captured image d2 based on the virtual camera information 112 and the similarity (and the three-dimensional data T11) using the first estimation model (S14). Thereafter, the learning unit 127 learns the first estimation model using the teacher position and orientation T13 (S18). In other words, the learning unit 127 evaluates the shooting position and orientation estimated in step S14 using the teacher position and orientation T13 and updates the parameters of the first estimation model. Specifically, the learning unit 127 updates the parameters of the first estimation model so that the shooting position and orientation estimated in step S14 approaches the teacher position and orientation T13. Note that the learning unit 127 may repeat steps S13, S14, and S18 until the learning result satisfies a predetermined condition. The predetermined condition may be, for example, but is not limited to, the number of repetitions or the error between the estimation result and the teacher data being equal to or less than a threshold.
[0046] If step S13 is (M1) global matching, the learning unit 127 updates the parameters of the first estimation model so that a high similarity is calculated when there is a large degree of overlap between the shooting ranges of the position and orientation of the virtual camera information 112 and the teacher's position and orientation T13. If step S13 is (M2) local matching, the learning unit 127 calculates a correspondence relationship R01 between pixels representing the same location in the virtual image information d0 and the teacher's captured image T12, based on the virtual camera information 112, the teacher's position and orientation T13, and the three-dimensional data T11. Then, the learning unit 127 updates the parameters of the first estimation model so that the pixel correspondence relationship R02 calculated in step S13 approaches the correspondence relationship R01.
[0047] <About learning 2 (training) of estimation models> In general, preparing high-quality 3D training data is expensive or difficult. Therefore, the positional relationship between two or more virtual cameras may be reflected in the training of an estimation model.
[0048] In this case, when imaging the three-dimensional data 111, the rendering unit 123 may generate virtual image information for an area corresponding to a common shooting range of the first virtual camera information and the second virtual camera information. In other words, when generating virtual image information from the first virtual camera information, the rendering unit 123 sets the common shooting range of the first virtual camera information and the second virtual camera information as the area to be rendered. This allows for pseudo-reproduction of loss of added image information due to occlusion when creating the three-dimensional data, thereby improving the robustness of the estimation model.
[0049] 9 is a flowchart showing the flow of the learning process of the first estimation model. First, the feature calculation unit 122 calculates shape feature values at each data point of the training three-dimensional data T11 (S11). Next, the rendering unit 123 specifies a common shooting range of the two virtual camera information 1121 and 1122 as a rendering target area (S12a). Then, the rendering unit 123 renders the three-dimensional data T11 for the rendering target area and associates the shape feature values with each pixel position to generate virtual image information d0 (S12b). Thereafter, the shooting position and orientation estimation device 100 performs the processes of steps S13, S14, and S18, similar to those of FIG. 8 described above.
[0050] <About learning 3 (training) of estimation models> A distance learning approach may be used to train the estimation model. FIG. 10 is a flowchart showing the flow of the learning process for the first estimation model. As a premise, the position and orientation of the virtual camera information 112a are assumed to have a shooting range similar to that of the teacher position and orientation T13. Furthermore, the position and orientation of the virtual camera information 112b are assumed to have a shooting range dissimilar to that of the teacher position and orientation T13. First, the feature calculation unit 122 calculates shape feature values at each data point of the teacher three-dimensional data T11 (S11).
[0051] Next, the rendering unit 123 generates first virtual image information from the three-dimensional data T11 based on the virtual camera information 112a as described above (S121). Then, the matching unit 124 calculates the similarity between the first virtual image information and the captured image T12 using the first estimation model (S131). Thereafter, the estimation unit 125 estimates the first shooting position and orientation in the captured image d2 based on the virtual camera information 112a and the similarity (and the three-dimensional data T11) using the first estimation model (S141).
[0052] In parallel with steps S121, S131, and S141, the rendering unit 123 generates second virtual image information from the three-dimensional data T11 based on the virtual camera information 112b as described above (S122). Then, the matching unit 124 calculates the similarity between the second virtual image information and the captured image T12 using the first estimation model (S132). Thereafter, the estimation unit 125 estimates the second shooting position and orientation in the captured image d2 based on the virtual camera information 112b and the similarity (and the three-dimensional data T11) using the first estimation model (S142).
[0053] Thereafter, the learning unit 127 learns a first estimation model using the first shooting position and orientation and the second shooting position and orientation (S181). In other words, the learning unit 127 evaluates the teacher captured image d2 in the feature space so that the similarity with the first virtual image information is greater than the similarity with the second virtual image information, and updates the parameters of the first estimation model. That is, the learning unit 127 learns so that the first virtual image information is closer to the teacher captured image d2 than the second virtual image information.
[0054] As described above, in the technology disclosed in Patent Document 1, when estimating the shooting position and orientation of a target object based on pre-created 3D data from an image of the target object captured by a camera, it is assumed that the 3D data contains color information (texture information, etc.). However, preparing 3D data containing color information in advance requires expensive equipment, such as a camera equipped with a 3D sensor. In contrast, the present disclosure maintains the estimation accuracy of the shooting position and orientation by using shape features even when using 3D data that does not contain color information. Therefore, 3D data that does not contain color information can be prepared using an inexpensive system, thereby reducing implementation costs. For example, the shooting position and orientation estimation device disclosed herein can use 3D data captured by an inexpensive LiDAR system (without a camera) used in MMS (Mobile Mapping System). Alternatively, the shooting position and orientation estimation device disclosed herein can use 3D-CAD data (without color information) created during the design phase of a structure. Furthermore, the input image of the present disclosure can be a photograph taken with an infrared camera, i.e., image data without color information.
[0055] Furthermore, since the imaging position and orientation estimation device according to the present disclosure visualizes 3D data based on virtual camera information, it can generate virtual image information using shape features that include depth-compromised representations, just like the input image, making it easier to match the virtual image information with the input image.
[0056] <Example 2-1> 5, the matching unit 124 may perform (M1) global matching processing, and may not perform (M2) local matching processing. In this case, in step S14, the estimation unit 125 may select, from among the plurality of pieces of virtual camera information, virtual camera information having the highest overall similarity between the virtual image information and the input image, and estimate the shooting position and orientation of the selected virtual camera information as the shooting position and orientation of the input image.
[0057] <Example 2-2> In step S13 of FIG. 5 described above, the matching unit 124 may perform (M2) local matching processing instead of (M1) global matching processing. In this case, in step S14, the estimation unit 125 may associate pixel positions of the input image with coordinates of the three-dimensional data based on the correspondence relationship of pixel positions obtained in M2, and estimate the shooting position and orientation of the input image using PnP (Perspective-n-Point) or the like. For example, if it is estimated that the shooting ranges of the virtual image information and the input image overlap, it is also possible to estimate the detailed position and orientation at which the input image was shot. Furthermore, when performing (M2) local matching processing, the matching unit 124 may use shooting position information of the input image together with virtual camera information. In this case, processing costs can be reduced without performing (M1) global matching processing.
[0058] <Example 2-3> The photographing position and orientation estimation apparatus 100 may use both (M1) global matching processing and (M2) local matching processing. For example, in step S13 of FIG. 5 described above, the matching unit 124 may perform (M1) global matching and then (M2) local matching. As a result, Example 2-3 can improve processing efficiency and estimation accuracy compared to Examples 2-1 and 2-2.
[0059] Furthermore, when generating multiple pieces of virtual camera information, it is possible to generate virtual camera information with a focus on the three-dimensional data portion within the shooting range of the virtual camera information that has a distinctive shape, thereby further improving processing efficiency and estimation accuracy.
[0060] 11 is a flowchart showing the flow of the imaging position and orientation estimation process when global matching process and local matching process are used together. First, the feature amount calculation unit 122 calculates shape feature amounts for each data point of the three-dimensional data 111 (S11). Then, the rendering unit 123 generates multiple pieces of virtual camera information (S123). Note that the order of steps S11 and S123 may be reversed or may be performed in parallel. In step S123, the rendering unit 123 generates multiple pieces of virtual camera information so that data points of the three-dimensional data 111 that have a higher degree of shape feature than other data points are included in the imaging range.
[0061] Here, methods for determining the degree of geometrical distinctiveness of the 3D data portion within the shooting range of the virtual camera information include, but are not limited to, the following. For example, the method may be a method of determining whether the distance from the shooting position and orientation of the virtual camera information to the target 3D data surface is equal to or greater than a predetermined value. Alternatively, the method may be a method of determining whether the degree of dispersion of normal distribution within the shooting range is equal to or greater than a predetermined value. Alternatively, the method may be a method of determining whether the degree to which contour information is included in the 3D data within the shooting range is equal to or greater than a predetermined value.
[0062] Then, the rendering unit 123 generates a plurality of pieces of virtual image information corresponding to each piece of virtual camera information based on the plurality of pieces of virtual camera information (S124). Then, the matching unit 124 calculates the similarity between each piece of virtual image information and the input image as in steps S133 to S135.
[0063] That is, the matching unit 124 calculates the similarity between each piece of virtual image information and the captured image d2 on an image-by-image basis (S133). That is, the matching unit 124 performs (M1) global matching processing. Then, the matching unit 124 selects virtual image information with a high similarity from among the plurality of pieces of virtual image information (S134). Then, the matching unit 124 calculates the similarity by identifying the correspondence between the selected virtual image information and the captured image on a pixel-by-pixel basis (S135). That is, the matching unit 124 performs (M2) local matching processing. Thereafter, the estimation unit 125 estimates the shooting position and orientation in the captured image d2 based on the virtual camera information 112 and the similarity (and the three-dimensional data 111) calculated in step S135 (S14). Thereafter, the shooting position and orientation estimation device 100 performs the processes of steps S15, S16, and S17, as in FIG. 5 described above.
[0064] 12 is a block diagram showing the hardware configuration of the image capture position and orientation estimation apparatus 100. The image capture position and orientation estimation apparatus 100 includes a memory 101, a processor 102, and a network interface 103.
[0065] The memory 101 is configured by a combination of volatile memory and non-volatile memory. The volatile memory is, for example, a volatile storage device such as RAM, and is a storage area for temporarily holding information when the processor 102 is operating. The non-volatile memory is, for example, a non-volatile storage device such as a hard disk or flash memory. The memory 101 stores at least a computer program that implements the processing of the image capture position and orientation estimation apparatus 100 according to the present disclosure. Note that the memory 101 may include storage located away from the processor 102. In this case, the processor 102 may access the memory 101 via an I / O (Input / Output) interface, not shown.
[0066] The processor 102 is a control device that controls each component of the image capture position and orientation estimation device 100. The processor 102 reads and executes software (computer programs) from the memory 101. As a result, the processor 102 realizes the functions of an acquisition unit 121, a feature calculation unit 122, a rendering unit 123, a matching unit 124, an estimation unit 125, a display unit 126, and a learning unit 127. That is, the processor 102 performs processing of the image capture position and orientation estimation method according to the present disclosure. The processor 102 may be, for example, a microprocessor, an MPU (Multi Processing Unit), or a CPU (Central Processing Unit). The processor 102 may also include multiple processors.
[0067] The network interface 103 may be used to communicate with a network node. The network interface 103 may include, for example, a network interface card (NIC) conforming to the IEEE 802.3 series. IEEE stands for Institute of Electrical and Electronics Engineers. The network interface 103 may also include a wireless local area network (LAN), a wired LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0068] <Embodiment 3> Here, when image data is captured close to a target object, there is often little change in shape within the captured range. Therefore, it may be difficult to estimate the shooting position and orientation using shape features. Therefore, it is preferable to automatically select a more distant image from multiple input images of the same target object, and estimate the shooting position and orientation for a closer image using the estimation result of the shooting position and orientation for the more distant image. Specifically, the shooting position and orientation estimation device according to the third embodiment further includes a selection unit in addition to the configuration of the shooting position and orientation estimation device 100 in FIG. 3.
[0069] The selection unit selects one of two input images captured at different distances from an object corresponding to the three-dimensional data as a first input image, which contains more diverse shape information, and the other image as a second input image. Here, the selection unit may select an image by determining the shape information contained in the image based on, for example, the number of line segments contained in the image data or the magnitude of change in depth value as a result of monocular depth estimation.
[0070] Furthermore, the estimation unit 125 estimates a first shooting position and orientation in the first input image based on the virtual image information and the similarity between the virtual image information and the first input image. Then, the estimation unit 125 estimates a relative shooting position and orientation between the first input image and the second input image. Here, the estimation process for the relative shooting position and orientation can use techniques such as homography transformation and Visual-SLAM. Then, the estimation unit 125 estimates a second shooting position and orientation in the second input image based on the first shooting position and orientation and the relative shooting position and orientation. Note that other configurations of the shooting position and orientation estimation device according to the third embodiment are the same as those shown in FIG. 3 above, and therefore, redundant explanations and illustrations will be omitted as appropriate.
[0071] FIG. 13 is a flowchart showing the flow of the photographing position and orientation estimation process for two input images. After steps S11 and S12, the selection unit selects one of the two photographed images that includes more diverse shape information as image A (e.g., photographed image d21) and the other as image B (e.g., photographed image d22) (S130). Then, the matching unit 124 calculates the similarity between the virtual image information d0 and image B (S13). Then, the estimation unit 125 estimates the photographing position and orientation A in image A based on the virtual camera information 112 and the similarity (S143). Then, the estimation unit 125 estimates the relative photographing position and orientation between images A and B (S144). Thereafter, the estimation unit 125 estimates the photographing position and orientation B in image B based on the photographing position and orientation A and the relative photographing position and orientation (S145). Thereafter, the photographing position and orientation estimation device 100 performs the processes of steps S15, S16, and S17, similar to FIG. 5 described above.
[0072] In this way, in the third embodiment, when a plurality of captured images are given, an image containing more diverse shape information is automatically selected as image A. This improves the estimation accuracy of the capturing position and orientation for closer captured images.
[0073] <Embodiment 4> FIG. 14 is a block diagram showing the configuration of a photographing position and orientation estimation system 1000. The photographing position and orientation estimation system 1000 includes a photographing terminal 200 and a photographing position and orientation estimation device 100a. The photographing terminal 200 and the photographing position and orientation estimation device 100a are communicably connected via a network N. Here, the network N is a wired or wireless communication line. The photographing terminal 200 is an information processing device equipped with a camera, a display screen, and a wireless communication function. The photographing terminal 200 is, for example, a mobile terminal such as a smartphone or a tablet terminal. The photographing terminal 200 photographs a target object 300 from an arbitrary photographing position and orientation in response to an operation by a user who carries the photographing terminal 200, and transmits a photographing position and orientation estimation request including the photographed image to the photographing position and orientation estimation device 100a. The photographing terminal 200 also receives an estimation result from the photographing position and orientation estimation device 100a and displays it on the display screen. The photographing terminal 200 may also receive a comparison result between a virtual image and a photographed image from the photographing position and orientation estimation device 100a and display it on the display screen.
[0074] The photographing position and orientation estimation device 100a has the same configuration as the photographing position and orientation estimation device 100 in FIG. 3 described above, and therefore will not be illustrated or described again. However, the acquisition unit 121 of the photographing position and orientation estimation device 100a receives a photographing position and orientation estimation request including a photographed image from the photographing terminal 200, and acquires the photographed image included in the estimation request as an input image. Furthermore, when the matching unit 124 of the photographing position and orientation estimation device 100a receives an image of an object photographed by the photographing terminal 200 as an input image, it calculates the similarity between virtual image information and the input image. Furthermore, the display unit 126 of the photographing position and orientation estimation device 100a transmits the estimation result to the photographing terminal 200 and displays it on the photographing terminal 200. Furthermore, the display unit 126 may transmit a comparison result between the virtual image and the photographed image to the photographing terminal 200 and display it on the photographing terminal 200.
[0075] Fig. 15 is a sequence chart showing the flow of the imaging position and orientation estimation process. First, similar to step S11 in Fig. 5, the feature amount calculation unit 122 calculates shape feature amounts at each data point of the three-dimensional data 111 (S41). Next, similar to step S12 in Fig. 5, the rendering unit 123 generates virtual image information based on the virtual camera information 112 (S42).
[0076] The photographing terminal 200 also photographs the target object 300 (S43). Then, the photographing terminal 200 transmits a photographing position and orientation estimation request including the photographed image to the photographing position and orientation estimation device 100a via the network N (S44). In response to this, the acquisition unit 121 of the photographing position and orientation estimation device 100a receives the estimation request from the photographing terminal 200 via the network N, and acquires the photographed image included in the estimation request as an input image.
[0077] Next, the matching unit 124 calculates the similarity between the virtual image information and the captured image (acquired from the photographing terminal 200) (S45). Then, the estimation unit 125 estimates the photographing position and orientation in the captured image based on the virtual camera information and the similarity (S46). Then, the display unit 126 converts the virtual image information into a visualized virtual image as described above (S47). Thereafter, the display unit 126 transmits the estimation result, the virtual image, the correspondence between the virtual image and the captured image, etc. to the photographing terminal 200 via the network N (S38). In response to this, the photographing terminal 200 receives the estimation result, the virtual image, the correspondence between the virtual image and the captured image, etc. from the photographing position and orientation estimation device 100a via the network N. Then, the photographing terminal 200 displays the estimation result, the virtual image, the correspondence between the virtual image and the captured image, etc., in comparison with each other (S49).
[0078] The photographing position and orientation estimation device 100a may output the estimation results, etc. to another information processing device together with the photographing terminal 200. Alternatively, the photographing position and orientation estimation device 100a may output the estimation results, etc. to another information processing device instead of transmitting them to the photographing terminal 200. Alternatively, the photographing position and orientation estimation device 100a may register (output) the estimation results, etc. to the storage unit 110 or an external storage device.
[0079] Furthermore, the photographing position and orientation estimation device 100a and the photographing terminal 200 in the photographing position and orientation estimation system 1000 may have functions distributed differently from the above. For example, the photographing terminal 200 may include a storage unit 110, an estimation unit 125, and a display unit 126. In this case, the photographing position and orientation estimation device 100a may transmit the similarity (or the correspondence between pixel positions) calculated by the matching unit 124 and the virtual image to the photographing terminal 200. Then, the estimation unit 125 of the photographing terminal 200 may estimate the photographing position and orientation based on the received similarity. Thereafter, the photographing terminal 200 may perform display in step S49. Alternatively, the photographing terminal 200 may include a storage unit 110, a matching unit 124, an estimation unit 125, and a display unit 126. Furthermore, the photographing position and orientation estimation device 100a may transmit virtual image information to the photographing terminal 200 instead of a virtual image. In this case, the photographing terminal 200 may convert the received virtual image information into a visualized virtual image and display it.
[0080] In this way, this embodiment provides the same effects as those of the second or third embodiment described above.
[0081] <Other embodiments> Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0082] Each drawing is merely an example for describing one or more embodiments. Each drawing may relate not only to one particular embodiment, but also to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessary to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.
[0083] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. (Appendix A1) a feature amount calculation means for calculating a shape feature amount from predetermined three-dimensional data; a generating means for generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position of the image area; a similarity calculation means for calculating a similarity between the virtual image information and an input image; A photographing position and orientation estimation device comprising: (Appendix A2) The generating means converts the shape feature amounts based on the shooting position and posture included in the virtual camera information, and generates the virtual image information using the converted shape feature amounts. The imaging position and orientation estimation device according to Appendix A1. (Appendix A3) The feature amount calculation means calculates the shape feature amount of each data point so as to represent the distribution of other data points around a specific data point of the three-dimensional data. The photographing position and orientation estimation device according to appendix A1 or A2. (Appendix A4) The generating means generating a plurality of virtual camera information so that a data point having a higher degree of shape characteristic than other data points among the data points of the three-dimensional data is included in the imaging range; generating a plurality of pieces of virtual image information corresponding to each piece of virtual camera information based on the plurality of pieces of virtual camera information; The similarity calculation means Calculating the similarity between each of the plurality of virtual image information and the input image The imaging position and orientation estimation device according to any one of appendices A1 to A3. (Appendix A5) The generating means When the three-dimensional data is visualized, the virtual image information is generated for an area corresponding to a common imaging range of the first virtual camera information and the second virtual camera information. The imaging position and orientation estimation device according to any one of appendices A1 to A4. (Appendix A6) a selection means for selecting, from two input images taken at different distances from an object corresponding to the three-dimensional data, one image containing more diverse shape information as a first input image and the other image as a second input image; estimating a first photographing position and orientation in the first input image based on the virtual image information and a similarity between the virtual image information and the first input image; Estimating a relative photographing position and orientation between the first input image and the second input image; an estimation means for estimating a second photographing position and orientation in the second input image based on the first photographing position and orientation and the relative photographing position and orientation; Further equipped The imaging position and orientation estimation device according to any one of appendices A1 to A5. (Appendix A7) The virtual image information may further include a display unit for displaying a virtual image in which the shape feature quantities associated with each pixel of the virtual image information are converted into color information by dimensional compression. The imaging position and orientation estimation device according to any one of appendices A1 to A6. (Appendix A8) The display means displays the virtual image and the input image in comparison with each other. The imaging position and orientation estimation device according to Appendix A7. (Appendix A9) The display means displays the virtual image together with the estimation result of the shooting position and orientation in the input image estimated based on the similarity. The imaging position and orientation estimation device according to appendix A7 or A8. (Appendix A10) The feature amount calculation means calculates the shape feature amount by quantifying the shape or direction of the distribution of data points in the three-dimensional data. The imaging position and orientation estimation device according to Appendix A3. (Appendix A11) The feature amount calculation means calculates a normal vector of each data point of the three-dimensional data as the shape feature amount. The imaging position and orientation estimation device according to Appendix A3. (Appendix A12) The feature amount calculation means A first model inputs three-dimensional data around a specific data point in the three-dimensional data and outputs the shape feature of the data point, and calculates the shape feature using a first trained model that has been machine-trained so that the shape feature will be similar when the distribution of data points in the three-dimensional data is similar. The imaging position and orientation estimation device according to Appendix A3. (Appendix A13) The feature amount calculation means The shape feature quantity is calculated using a second trained model that has been machine-trained using the three-dimensional data of the predetermined object, the photographed image of the object, and the photographing position and orientation of the photographed image as training data. The imaging position and orientation estimation device according to Appendix A3. (Appendix A14) The similarity calculation means The similarity is calculated using a third trained model that is machine-trained using the three-dimensional data of the predetermined object, the photographed image of the object, and the photographing position and orientation of the photographed image as training data. The imaging position and orientation estimation device according to any one of appendices A1 to A13. (Appendix A15) the similarity calculation means calculates the similarity when an image of an object corresponding to the three-dimensional data is acquired as the input image; The image capturing apparatus further includes an estimation unit for estimating a photographing position and orientation in the input image based on the virtual camera information and the similarity. The imaging position and orientation estimation device according to any one of appendices A1 to A14. (Appendix B1) a photographing terminal and a photographing position and orientation estimation device communicably connected to the photographing terminal; the imaging position and attitude estimation device, Calculate shape features from the 3D data of a given object, generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on virtual camera information with each pixel position of the image area; When an image of the object photographed by the photographing terminal is received as an input image, a similarity between the virtual image information and the input image is calculated; estimating a photographing position and orientation in the input image based on the virtual camera information and the similarity; The estimated photographing position and orientation are returned to the photographing terminal. Shooting position and orientation estimation system. (Appendix C1) The computer Calculate shape features from the specified 3D data, generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position of the image area; Calculating the similarity between the virtual image information and the input image Shooting position and orientation estimation method. (Appendix D1) A feature amount calculation process for calculating a shape feature amount from predetermined three-dimensional data; a generation process for generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position of the image area; a similarity calculation process for calculating a similarity between the virtual image information and an input image; A shooting position and orientation estimation program that causes a computer to execute the above.
[0084] Some or all of the elements (e.g., configurations and functions) described in Appendix A2 to Appendix A15 that are dependent on Appendix A1 {e.g., device} may also be dependent on Appendix B1 {e.g., system}, Appendix C1 {e.g., method}, and Appendix D1 {e.g., program} in the same dependency relationship as Appendix A2 to Appendix A15. Some or all of the elements described in any appendix may be applied to various hardware, software, recording means for recording software, systems, and methods. [Explanation of symbols]
[0085] 1. Shooting position and orientation estimation device 11 Feature calculation unit 12 Generation part 13 Similarity calculation unit 100 Shooting position and orientation estimation device 110 Storage section 111 3D data 112 Virtual Camera Information 121 Acquisition Department 122 Feature calculation unit 123 Rendering Department 124 Matching Section 125 Estimation part 126 Display section 127 Learning Department d0 Virtual image information H Height W width D: Number of dimensions of shape features d01 Virtual image d1 Virtual Image d2 Captured images P11 pixels P12 pixels P21 pixels P22 pixels R1 Correspondence R2 correspondence T1 training data T11 3D data T12 image T13 Position and Posture 1121 Virtual Camera Information 1122 Virtual Camera Information 101 Memory 102 processors 103 Network Interface d21 Shooting image d22 Shooting image 1000 Shooting position and orientation estimation system 100a Shooting position and orientation estimation device 200 Photographic Devices 300 Target Objects N Network
Claims
1. a feature amount calculation means for calculating a shape feature amount from predetermined three-dimensional data; a generating means for generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position of the image area; a similarity calculation means for calculating a similarity between the virtual image information and an input image; A photographing position and orientation estimation device comprising:
2. The generating means converts the shape feature amounts based on the shooting position and posture included in the virtual camera information, and generates the virtual image information using the converted shape feature amounts. The imaging position and orientation estimation device according to claim 1 .
3. The feature amount calculation means calculates the shape feature amount of each data point so as to represent the distribution of other data points around a specific data point of the three-dimensional data. The photographing position and orientation estimation device according to claim 1 or 2.
4. The generating means generating a plurality of virtual camera information so that a data point having a higher degree of geometrical characteristic than other data points among the data points of the three-dimensional data is included in the imaging range; generating a plurality of pieces of virtual image information corresponding to each piece of virtual camera information based on the plurality of pieces of virtual camera information; The similarity calculation means Calculating the similarity between each of the plurality of virtual image information and the input image The photographing position and orientation estimation device according to claim 1 or 2.
5. The generating means When the three-dimensional data is visualized, the virtual image information is generated for an area corresponding to a common imaging range of the first virtual camera information and the second virtual camera information. The photographing position and orientation estimation device according to claim 1 or 2.
6. a selection means for selecting, from two input images taken at different distances from an object corresponding to the three-dimensional data, one image containing more diverse shape information as a first input image and the other image as a second input image; estimating a first photographing position and orientation in the first input image based on the virtual image information and a similarity between the virtual image information and the first input image; Estimating a relative photographing position and orientation between the first input image and the second input image; an estimation means for estimating a second photographing position and orientation in the second input image based on the first photographing position and orientation and the relative photographing position and orientation; Further equipped The photographing position and orientation estimation device according to claim 1 or 2.
7. The virtual image information may further include a display unit for displaying a virtual image in which the shape feature quantities associated with each pixel of the virtual image information are converted into color information by dimensional compression. The photographing position and orientation estimation device according to claim 1 or 2.
8. a photographing terminal and a photographing position and orientation estimation device communicably connected to the photographing terminal; the imaging position and attitude estimation device, Calculating shape features from three-dimensional data of a predetermined object; generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on virtual camera information with each pixel position of the image area; When an image of the object photographed by the photographing terminal is received as an input image, a similarity between the virtual image information and the input image is calculated; estimating a photographing position and orientation in the input image based on the virtual camera information and the similarity; The estimated photographing position and orientation are returned to the photographing terminal. Shooting position and orientation estimation system.
9. The computer Calculating shape features from predetermined three-dimensional data; generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position of the image area; Calculating the similarity between the virtual image information and the input image Shooting position and orientation estimation method.
10. A feature amount calculation process for calculating a shape feature amount from predetermined three-dimensional data; a generation process for generating virtual image information by associating the shape feature amount projected onto an image area when imaging the three-dimensional data based on predetermined virtual camera information with each pixel position of the image area; a similarity calculation process for calculating a similarity between the virtual image information and an input image; A shooting position and orientation estimation program that causes a computer to execute the above.
Citation Information
Patent Citations
Matching device, method and program
JP2023153664A