Information processing device and method implemented by the information processing device

By classifying three-dimensional points based on normal directions and calculating group accuracy, the method improves the precision of three-dimensional model generation, addressing the challenge of low accuracy in existing technologies.

JP2026086918APending Publication Date: 2026-05-26PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2026-03-05
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies face challenges in generating accurate three-dimensional models from sensor data, particularly when the accuracy of three-dimensional points is low, leading to incomplete or inaccurate models, necessitating repeated detection of the target space.

Method used

A calculation method and device that classify three-dimensional points into groups based on their normal directions, using sensors like LiDAR or image sensors, and calculate the accuracy of each group to improve the precision of three-dimensional model generation.

Benefits of technology

Enhances the accuracy of three-dimensional model generation by classifying points and displaying areas of low accuracy, allowing users to efficiently adjust sensor direction for improved detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086918000001_ABST
    Figure 2026086918000001_ABST
Patent Text Reader

Abstract

The accuracy of each group to which multiple three-dimensional points belong is calculated. [Solution] An information processing device according to one aspect of the present disclosure comprises a processor and a memory connected to the processor, wherein the processor uses the memory to acquire a plurality of three-dimensional points (S151), generates a plurality of groups based on the normal direction associated with each three-dimensional point (S152), generates region information for each of the plurality of groups including information indicating the normal direction corresponding to the group and information indicating the evaluation value corresponding to the group, and the evaluation value corresponding to the group is generated based on the evaluation values ​​corresponding to one or more three-dimensional points belonging to the group (S153).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a calculation method and a calculation device.

Background Art

[0002] Patent Document 1 discloses a technique for representing a structure by a three-dimensional point cloud.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide a calculation method and the like that can calculate the accuracy for each group to which a plurality of three-dimensional points belong.

Means for Solving the Problems

[0005] An information processing apparatus according to an aspect of the present disclosure includes a processor and a memory connected to the processor. The processor uses the memory to acquire a plurality of three-dimensional points, generates a plurality of groups based on the normal directions associated with each three-dimensional point, and for each of the plurality of groups, generates region information including information indicating the normal direction corresponding to the group and information indicating an evaluation value corresponding to the group. The evaluation value corresponding to the group is generated based on the evaluation values corresponding to one or more three-dimensional points belonging to the group.

Effects of the Invention

[0006] The present disclosure can provide a calculation method and the like that can calculate the accuracy for each group to which a plurality of three-dimensional points belong.

Brief Description of the Drawings

[0007] [Figure 1] Figure 1 is a block diagram showing the configuration of a three-dimensional reconstruction system according to an embodiment. [Figure 2] Figure 2 is a block diagram showing the configuration of the imaging device according to the embodiment. [Figure 3] Figure 3 is a flowchart showing the operation of the imaging device according to the embodiment. [Figure 4] Figure 4 is a flowchart showing the position and orientation estimation process according to the embodiment. [Figure 5] Figure 5 is a diagram illustrating a method for generating three-dimensional points according to an embodiment. [Figure 6] Figure 6 is a flowchart showing the position and orientation integration process according to the embodiment. [Figure 7] Figure 7 is a plan view showing the shooting process in the target space according to the embodiment. [Figure 8] Figure 8 is a diagram illustrating an example of an image and an example of a comparison process according to the embodiment. [Figure 9] Figure 9 is a flowchart showing the accuracy calculation process according to the embodiment. [Figure 10] Figure 10 is a diagram illustrating a method for calculating the normal direction of a three-dimensional point according to an embodiment. [Figure 11] Figure 11 is a diagram illustrating a method for calculating the normal direction of a three-dimensional point according to an embodiment. [Figure 12] Figure 12 is a diagram illustrating a method for classifying three-dimensional points according to an embodiment. [Figure 13] Figure 13 is a diagram illustrating a method for classifying three-dimensional points according to an embodiment. [Figure 14] Figure 14 is a diagram illustrating a method for classifying three-dimensional points according to an embodiment. [Figure 15] Figure 15 is a diagram illustrating the method for calculating the accuracy of a group according to an embodiment. [Figure 16] Figure 16 is a flowchart showing the display process according to the embodiment. [Figure 17]FIG. 17 is a diagram showing an example of displaying voxels according to an embodiment. [Figure 18] FIG. 18 is a diagram showing an example of displaying a UI screen according to an embodiment. [Figure 19] FIG. 19 is a diagram showing an example of displaying a UI screen according to an embodiment. [Figure 20] FIG. 20 is a diagram showing an example of displaying a UI screen according to an embodiment. [Figure 21] FIG. 21 is a diagram showing an example of displaying a UI screen according to an embodiment. [Figure 22] FIG. 22 is a flowchart showing a calculation method according to an embodiment.

Mode for Carrying Out the Invention

[0008] (Summary of the Present Disclosure) A calculation method according to one aspect of the present disclosure represents an object in a space on a computer and obtains a plurality of three-dimensional points each indicating the position of the object. Each three-dimensional point of the plurality of three-dimensional points is classified into a plurality of groups based on the normal direction of the three-dimensional point. If the first accuracy of at least one three-dimensional point belonging to each group of the plurality of groups is high, the second accuracy of the group is calculated so as to be high. Each three-dimensional point of the plurality of three-dimensional points is generated by detecting light from the object in a plurality of different directions from a plurality of different positions by a sensor, and the normal direction is obtained based on the plurality of different directions used for generating the three-dimensional point having the normal direction.

[0009] According to this, according to the direction in which the sensor detects light from the object, a plurality of three-dimensional points are classified into a plurality of groups, and according to the accuracy of the three-dimensional points belonging to each group of the plurality of groups, the accuracy for each group to which the three-dimensional points belong can be calculated. That is, according to the calculation method according to one aspect of the present disclosure, since a plurality of three-dimensional points are classified based on the normal direction of each three-dimensional point, the accuracy for each normal direction used for classification can be calculated.

[0010] Further, for example, in the classification, the space is divided into a plurality of small spaces, and each of the plurality of three-dimensional points is classified into the plurality of groups based on the normal direction of the three-dimensional point and the small space among the plurality of small spaces that contains the three-dimensional point.

[0011] According to this, the accuracy for each normal direction and each location can be calculated.

[0012] Further, for example, the first small space among the plurality of small spaces contains the first three-dimensional point and the second three-dimensional point among the plurality of three-dimensional points, and the line extending in the normal direction of the first three-dimensional point passes through the first surface among the plurality of planes defining the first small space. In the classification, when the line extending in the normal direction of the second three-dimensional point passes through the first surface, the first three-dimensional point and the second three-dimensional point are classified into the same group.

[0013] According to this, a plurality of three-dimensional points can be classified into a plurality of groups for each plane according to the normal direction of the three-dimensional point.

[0014] Further, for example, the sensor is any one of a LiDAR sensor, a depth sensor, an image sensor, or a combination thereof.

[0015] Further, for example, the composite direction of the plurality of different directions used for generating the three-dimensional point having the normal direction and the normal direction are opposite.

[0016] Further, for example, each of the plurality of groups corresponds to one of the plurality of planes defining the plurality of small spaces. In the classification, each three-dimensional point of the plurality of three-dimensional points is classified into a group corresponding to the passing plane through which the line extending in the normal direction of the three-dimensional point passes among the plurality of planes defining the small space that contains the three-dimensional point, and the passing plane is displayed in a color according to the accuracy of the group corresponding to the passing plane.

[0017] According to this, the passing surface is displayed in a color corresponding to the accuracy of the group that corresponds to that passing surface, thus providing a display that helps the user adjust the direction of the sensor they are operating.

[0018] Furthermore, for example, the passing surface is displayed in a different color depending on whether the accuracy of the group corresponding to the passing surface is above a predetermined accuracy or below a predetermined accuracy.

[0019] According to this, the passing surface is displayed in a different color depending on whether the accuracy of the group corresponding to that passing surface is above a predetermined accuracy or below a predetermined accuracy, thus providing a display that helps the user adjust the direction of the sensor they are operating.

[0020] Furthermore, for example, each of the multiple groups corresponds to one of the multiple subspaces, and the multiple subspaces are displayed in a color corresponding to the precision of the group that corresponds to that subspace.

[0021] According to this, multiple small spaces are displayed in colors corresponding to the accuracy of the group that the small space belongs to, thus providing a display that helps the user adjust the direction of the sensor they are operating.

[0022] Furthermore, for example, among the multiple subspaces, only the subspaces corresponding to the group whose accuracy is less than a predetermined accuracy are displayed.

[0023] According to this, among multiple small spaces, only the small space corresponding to the group whose accuracy is below a predetermined accuracy is displayed, thus providing a display that helps the user adjust the direction of the sensor they are operating.

[0024] Furthermore, for example, in the calculation described above, for each of the multiple groups, a predetermined number of three-dimensional points are extracted from the one or more three-dimensional points belonging to that group, and the accuracy of the group is calculated based on the accuracy of each of the extracted predetermined number of three-dimensional points.

[0025] This method allows for reducing processing load while appropriately calculating the accuracy of multiple groups.

[0026] Furthermore, for example, in the above calculation, the accuracy of each of the multiple groups is calculated based on the reprojection error that indicates the accuracy of one or more three-dimensional points belonging to that group.

[0027] According to this method, the accuracy of multiple three-dimensional points can be calculated appropriately.

[0028] Furthermore, for example, a second three-dimensional model having a lower resolution than the first three-dimensional model composed of the plurality of three-dimensional points is used, and the accuracy of the plurality of groups is superimposed and displayed on the second three-dimensional model composed of at least a portion of the plurality of three-dimensional points.

[0029] According to this, the accuracy of multiple groups is superimposed on the second three-dimensional model, providing a display that helps the user adjust the direction of the sensor they are operating.

[0030] Furthermore, for example, the accuracy of the multiple groups is superimposed and displayed on the overhead view of the second three-dimensional model.

[0031] According to this, the accuracy of multiple groups is superimposed on an overhead view of the second three-dimensional model, providing a display that helps users adjust the direction of the sensors they are operating.

[0032] Furthermore, for example, the sensor is an image sensor, and the accuracy of the plurality of groups is superimposed and displayed on the image captured by the image sensor, which was used to generate the plurality of three-dimensional points.

[0033] According to this, the accuracy of multiple groups is superimposed on the image captured by the imaging device, providing a display that helps the user adjust the direction of the sensor they are operating.

[0034] Furthermore, for example, the accuracy of the multiple groups is superimposed and displayed on a map of the target space in which the object is located.

[0035] According to this, the accuracy of multiple groups is superimposed on the map, providing a display that helps users adjust the direction of the sensor they are operating.

[0036] Furthermore, for example, the accuracy of the multiple groups is displayed while the sensor is detecting light from the object.

[0037] According to this, a display can be provided to help the user adjust the direction of the sensor they are operating while they are detecting light from an object.

[0038] Furthermore, for example, the opposite direction of the combined direction of the normal directions of two or more three-dimensional points belonging to the first group included in the plurality of groups is calculated.

[0039] According to this, the direction corresponding to each group within multiple groups can be calculated.

[0040] Furthermore, a calculation device according to one aspect of the present disclosure comprises a processor and a memory, wherein the processor uses the memory to acquire a plurality of three-dimensional points that represent an object in computer space and each point indicates the position of the object, classifies each of the plurality of three-dimensional points into a plurality of groups based on the normal direction of the three-dimensional point, calculates a second accuracy such that if the first accuracy of at least one three-dimensional point belonging to each of the plurality of groups is high, the second accuracy of the group is high, and each of the plurality of three-dimensional points is generated by detecting light from the object with a sensor from a plurality of different positions in a plurality of different directions, and the normal direction is determined based on the plurality of different directions used to generate the three-dimensional point having the normal direction.

[0041] According to this, the same effect as the calculation method relating to one aspect of this disclosure is achieved.

[0042] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.

[0043] The embodiments will be described below in detail with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement positions and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, components in the following embodiments that are not described in an independent claim will be described as optional components.

[0044] (Embodiment) [composition] Conventionally, there is a technology that uses sensors such as cameras, depth sensors, or LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) sensors to detect a target space (target space) such as people and buildings, and uses the resulting information (detection result) as input to generate a three-dimensional model of the target space based on the position and orientation of the sensor. For example, by taking an image generated by photographing a subject with a camera as input, multiple three-dimensional points are generated based on the position and orientation of the camera, and a three-dimensional model of the subject composed of these multiple three-dimensional points is generated. The generation of such three-dimensional models is used in surveying at construction sites, etc.

[0045] Here, the accuracy (certainty) of multiple three-dimensional points generated from detection results (detection data) obtained by the user detecting the target space using a sensor may be low, such as images taken by the user using a camera to capture the target space, or more specifically, image data obtained by capturing the target space using a camera. When the accuracy of multiple three-dimensional points (more specifically, the positional information of multiple three-dimensional points) is low in this way, it may not be possible to generate a three-dimensional model composed of these multiple three-dimensional points. In such cases, the user needs to detect the target space again using the sensor. Therefore, in order to obtain detection results that generate highly accurate three-dimensional points, there is a need for technology that enables the user to efficiently detect the target space using a sensor. In the three-dimensional reconstruction system according to the embodiment, in order to obtain detection results that generate highly accurate three-dimensional points, the positions that should be detected by the sensor are displayed to the user so that they can understand them, enabling the user to efficiently detect the target space using a sensor.

[0046] Figure 1 is a block diagram showing the configuration of a three-dimensional reconstruction system according to an embodiment.

[0047] The three-dimensional reconstruction system according to this embodiment is a system that generates a three-dimensional model of the target space, such as a three-dimensional map, using the detection results of sensors in the target space (hereinafter referred to as the target space).

[0048] Here, a three-dimensional model is a computer representation of an object (detection target) in the target space detected by the sensor. The three-dimensional model, for example, contains positional information for each three-dimensional point on the detection target. In this embodiment, the sensor is a camera, and it detects the target space (more specifically, the detection target located in the target space) by capturing an image of the detection target and generating an image.

[0049] The three-dimensional reconstruction system comprises an imaging device 101 and a reconstruction device 102.

[0050] The imaging device 101 is a terminal device used by the user, such as a tablet, smartphone, or notebook personal computer. The imaging device 101 has a camera function, a function to estimate the position and orientation of the camera (hereinafter referred to as position and orientation), and a function to display the accuracy of a three-dimensional model composed of multiple three-dimensional points (also called a three-dimensional point cloud) generated based on the results of the camera's imaging.

[0051] Here, the accuracy of a three-dimensional model refers to the degree to which the three-dimensional model accurately reproduces the actual target space. For example, the accuracy of a three-dimensional model is calculated based on the error between the positional information of the three-dimensional points that make up the three-dimensional model and their actual positions. For example, the higher the accuracy of the three-dimensional model, the smaller the error between the positional information of the three-dimensional points that make up the three-dimensional model and their actual positions. Conversely, the lower the accuracy of the three-dimensional model, the larger the error between the positional information of the three-dimensional points that make up the three-dimensional model and their actual positions. Also, for example, if it is difficult to generate a three-dimensional model (for example, if a three-dimensional model cannot be generated), the accuracy of the three-dimensional model may be judged to be low.

[0052] The imaging device 101 transmits the captured image (image data) and position / orientation (more specifically, position / orientation information, which is information indicating the position / orientation) to the reconstruction device 102 during and after imaging.

[0053] Here, "image" refers to, for example, a video, but "image" may also refer to multiple still images.

[0054] Furthermore, the imaging device 101, for example, estimates the position and orientation of the camera during imaging, and uses the position and orientation and at least one of the three-dimensional point cloud to determine the areas that have been imaged (also called the imaged areas) and / or areas that have not been imaged, and presents at least one of these areas to the user.

[0055] Here, areas that have not been photographed refer to areas that were not photographed at the time the target space was being photographed (for example, areas hidden by other objects) and areas that were photographed but for which no three-dimensional points could be obtained. Areas that have not been photographed are areas where the accuracy of the three-dimensional model of the target space is low. On the other hand, areas that have been photographed are areas where three-dimensional points of the target space have been obtained. In other words, areas that have been photographed are areas where the accuracy of the three-dimensional model of the target space is high.

[0056] Furthermore, the imaging device 101 may estimate the position and orientation of the imaging unit 111 while photographing the subject and perform three-dimensional shape reconstruction (generation of a three-dimensional model).

[0057] The reconstruction device 102 is, for example, a server connected to the imaging device 101 via a network or the like. The reconstruction device 102 acquires images captured by the imaging device 101 and generates a three-dimensional model using the acquired images. For example, the reconstruction device 102 may use the camera position and orientation estimated by the imaging device 101, or it may estimate the camera position and orientation from the acquired images.

[0058] Furthermore, data transfer between the imaging device 101 and the reconstruction device 102 may be performed offline via an HDD (hard disk drive) or the like, or it may be performed continuously via a network.

[0059] The three-dimensional model generated by the reconstruction device 102 may be a densely reconstructed three-dimensional point cloud or a collection of three-dimensional meshes. The three-dimensional point cloud generated by the imaging device 101 is a collection of three-dimensional points that sparsely reconstruct feature points such as corners of objects in space. In other words, the three-dimensional model composed of the three-dimensional point cloud (i.e., multiple three-dimensional points) generated by the imaging device 101 has a lower spatial resolution than the three-dimensional model generated by the reconstruction device 102. To put it another way, the three-dimensional model (three-dimensional point cloud) generated by the imaging device 101 is a simpler model than the three-dimensional model generated by the reconstruction device 102.

[0060] A simple model is, for example, a model with a small amount of information, a model that is easy to generate, or a model with low accuracy. For example, the three-dimensional model generated by the imaging device 101 is a sparser three-dimensional point cloud than the three-dimensional model generated by the reconstruction device 102.

[0061] Figure 2 is a block diagram showing the configuration of the imaging device 101 according to the embodiment.

[0062] The imaging device 101 includes an imaging unit 111, a position and orientation estimation unit 112, a position and orientation integration unit 113, a region detection unit 114, a UI unit 115, a control unit 116, an image storage unit 117, a position and orientation storage unit 118, and a region information storage unit 119.

[0063] The imaging unit 111 is an imaging device such as a camera, which generates (acquires) an image (moving image) of the target space by photographing the target space. The imaging unit 111 stores the acquired image in the image storage unit 117.

[0064] The following examples primarily describe the use of moving images, but multiple still images may be used instead.

[0065] Furthermore, the imaging unit 111 may add information to the image storage unit 117 that identifies the camera that took the image (for example, the imaging unit 111) as header information for moving or still images.

[0066] Furthermore, the imaging unit 111 may be a single camera that, for example, is moved by the user while capturing images of the subject (target space). Alternatively, the imaging unit 111 may consist of multiple cameras. Alternatively, for example, the imaging unit 111 may have a camera and other sensors such as a three-dimensional measurement sensor.

[0067] The position and orientation estimation unit 112 is a processing unit that estimates (calculates) the three-dimensional position and orientation of the imaging unit 111 that captured the image, using the image stored in the image storage unit 117. The position and orientation estimation unit 112 also stores the estimated position and orientation (more specifically, position and orientation information) in the position and orientation storage unit 118. For example, the position and orientation estimation unit 112 uses image processing such as Visual SLAM (Simultaneous Localization and Mapping) or SfM (Structure from Motion) to estimate the position and orientation. Alternatively, the position and orientation estimation unit 112 may estimate the position and orientation of the imaging unit 111 using information obtained from various sensors (GPS (Global Positioning System) or acceleration sensors) provided by the imaging device 101. In the former case, position and orientation estimation can be achieved from information from the imaging unit 111. In the latter case, image processing can be achieved with low processing power.

[0068] Furthermore, the position and orientation estimation unit 112 generates a three-dimensional point cloud using, for example, the position and orientation and image from the imaging unit 111. The position and orientation integration unit 113 generates three-dimensional data, which is position information (also called map information) indicating the position of each of the multiple three-dimensional points.

[0069] The position and orientation integration unit 113 integrates the position and orientation of the imaging unit 111 estimated from each imaging session when multiple imaging sessions are performed in the same environment, and calculates the position and orientation that can be handled in the same space. Specifically, the position and orientation integration unit 113 uses the three-dimensional coordinate axes of the position and orientation obtained from the first imaging session as the reference coordinate axes. Then, the position and orientation integration unit 113 converts the coordinates of the position and orientation obtained from the second and subsequent imaging sessions into coordinates in the space of the reference coordinate axes.

[0070] The region detection unit 114 is a processing unit that calculates the accuracy of multiple three-dimensional points. For example, the region detection unit 114 uses the image stored in the image storage unit 117 and the position and orientation information stored in the position and orientation storage unit 118 to detect regions in the target space where three-dimensional reconstruction is possible (i.e., regions with high accuracy) and regions where three-dimensional reconstruction is not possible (i.e., regions with low accuracy).

[0071] Furthermore, the accuracy of a three-dimensional point refers to the accuracy of the positional information of that three-dimensional point.

[0072] Regions in the target space where three-dimensional reconstruction is not possible are, for example, regions where no images have been taken, or regions where the accuracy of three-dimensional reconstruction is low, and where the number of images taken of that region is small (less than a predetermined number). Furthermore, regions with low accuracy are regions where, when three-dimensional points are generated, the error between the generated three-dimensional points and the actual position is large. The region detection unit 114 also stores information about the detected region in the region information storage unit 119.

[0073] More specifically, the region detection unit 114 first acquires a plurality of three-dimensional points that represent an object in computer space and each point indicates the position of the object. Next, the region detection unit 114 classifies each of the plurality of three-dimensional points into a plurality of groups based on the normal direction of the three-dimensional point. Next, the region detection unit 114 calculates a second accuracy such that if the first accuracy of at least one three-dimensional point belonging to each of the plurality of groups is high, the second accuracy of that group will also be high.

[0074] Here, each of the multiple three-dimensional points is generated by the imaging unit 111 capturing images of an object from multiple different shooting positions and in multiple different shooting directions. The normal direction is determined based on the multiple different shooting directions used to generate the three-dimensional point having that normal direction. Specifically, the normal direction is the direction determined by combining the shooting positions and shooting directions of the imaging unit 111 when two or more shooting results (i.e., images) are obtained that were used to generate the three-dimensional point from among the multiple shooting results.

[0075] The specific processing performed by the region detection unit 114 will be described later.

[0076] The UI unit 115 is a user interface (UI) that presents the captured image and information indicating the accuracy calculated by the region detection unit 114 (also called region information) to the user. The UI unit 115 also has an input function for the user to input instructions to start and end shooting. For example, the UI unit 115 is a display with a touch panel.

[0077] The control unit 116 is a processing unit that controls the entire process of the imaging device 101, including the imaging process.

[0078] The processing units of the imaging device 101, such as the position and orientation estimation unit 112, the position and orientation integration unit 113, the region detection unit 114, and the control unit 116, are realized, for example, by a memory in which a control program is stored and a processor that executes the control program.

[0079] The image storage unit 117 is a memory device that stores images generated by the imaging unit 111.

[0080] The position and orientation storage unit 118 is a storage device that stores information indicating the position and orientation of the imaging unit 111 at the time of shooting, which is generated by the position and orientation estimation unit 112.

[0081] The region information storage unit 119 is a memory device that stores the information generated by the region detection unit 114.

[0082] The information stored in the region information storage unit 119 may be two-dimensional information superimposed on the image, or it may be three-dimensional information such as three-dimensional coordinate information.

[0083] The image storage unit 117, the position and orientation storage unit 118, and the area information storage unit 119 are implemented by, for example, flash memory or an HDD. The image storage unit 117, the position and orientation storage unit 118, and the area information storage unit 119 may be implemented by a single storage device, or they may be implemented by different storage devices.

[0084] [Operation] Next, we will explain the operation of the imaging device 101.

[0085] Figure 3 is a flowchart showing the operation of the imaging device 101 according to the embodiment.

[0086] In the imaging device 101, imaging is started and stopped according to the user's instructions. Specifically, imaging starts when the "Start Imaging" button on the UI (image) displayed by the UI unit 115 is pressed. If an instruction to start imaging is entered (Yes in S101), the imaging unit 111 starts imaging. The captured image is saved in the image storage unit 117.

[0087] Next, the position and orientation estimation unit 112 calculates the position and orientation each time an image is added (S102). The calculated position and orientation are stored in the position and orientation storage unit 118. At this time, if there is a three-dimensional point cloud generated by, for example, SLAM in addition to the position and orientation, the generated three-dimensional point cloud is also stored.

[0088] Next, the position and orientation integration unit 113 integrates the position and orientation data (S103). Specifically, the position and orientation integration unit 113 uses the position and orientation estimation results and the image to determine if it is possible to integrate the three-dimensional coordinate spaces of the position and orientation of the previously captured images and the position and orientation of the newly captured images. If possible, it integrates them. In other words, the position and orientation integration unit 113 converts the coordinates of the position and orientation of the newly captured image to the coordinate system of the previously captured position and orientation data. As a result, multiple position and orientation data are represented in a single three-dimensional coordinate space. Therefore, it becomes possible to use data obtained from multiple captures in common, improving the accuracy of position and orientation estimation.

[0089] Next, the region detection unit 114 calculates the accuracy of each of the multiple three-dimensional points (S104). Specifically, the region detection unit 114 uses the accuracy of each of the multiple three-dimensional points to classify the multiple three-dimensional points in a predetermined way and calculates the accuracy for each group. The region detection unit 114 also stores the region information indicating the calculated accuracy in the region information storage unit 119.

[0090] Next, the UI unit 115 displays the region information obtained through the above process (S105). For example, the UI unit 115 displays the accuracy of multiple groups calculated by the region detection unit 114 while the imaging unit 111 is capturing the target space.

[0091] Furthermore, this series of processes is repeated until the end of shooting (S106). For example, these processes are repeated each time one or more frames of images are acquired.

[0092] Figure 4 is a flowchart illustrating the position and orientation estimation process according to the embodiment. Specifically, Figure 4 is a flowchart detailing step S102.

[0093] First, the position and orientation estimation unit 112 acquires an image from the image storage unit 117 (S111).

[0094] Next, the position and orientation estimation unit 112 calculates the position and orientation of the imaging unit 111 in each image using the acquired images (S112). For example, the position and orientation estimation unit 112 calculates the position and orientation using image processing such as SLAM or SfM.

[0095] Furthermore, if the imaging device 101 has a sensor such as an IMU (Inertial Measurement Unit), the position and orientation estimation unit 112 may use the information obtained from the sensor to estimate the position and orientation.

[0096] Furthermore, the position and orientation estimation unit 112 may use results obtained by pre-calibration as camera parameters such as the focal length of the lens. Alternatively, the position and orientation estimation unit 112 may calculate camera parameters simultaneously with the estimation of position and orientation.

[0097] Next, the position and orientation estimation unit 112 stores the calculated position and orientation information in the position and orientation storage unit 118 (S113).

[0098] If the calculation of position and orientation information fails, information indicating the failure may be stored in the position and orientation storage unit 118.

[0099] This makes it possible to know the location and time of the failure, as well as the type of image that failed, and this information can be used when reshooting, etc.

[0100] Furthermore, the position and orientation estimation unit 112 generates three-dimensional points in the target space based on the image and the calculated position and orientation.

[0101] Figure 5 is a diagram illustrating a method for generating three-dimensional points according to an embodiment.

[0102] As shown in Figure 5(a), for example, first the position and orientation estimation unit 112 extracts feature points of the subject contained in each of the multiple images.

[0103] Next, the position and orientation estimation unit 112 performs feature point matching between multiple images for the extracted feature points.

[0104] Next, as shown in Figure 5(b), the position and orientation estimation unit 112 generates three-dimensional points (map information) by performing triangulation using camera geometry. For example, the three-dimensional point cloud is optimized along with the position and orientation of the imaging unit 111. For example, the three-dimensional points are used as initial values ​​for similarity point matching search.

[0105] Furthermore, a three-dimensional point (map information) may include not only the positional information of the three-dimensional point, but also the color of the three-dimensional point, information representing the surface shape around the three-dimensional point, and information indicating which frame (image) generated the three-dimensional point.

[0106] Figure 6 is a flowchart of the position and orientation integration process according to the embodiment. Specifically, Figure 6 is a flowchart detailing step S103.

[0107] First, the position and orientation integration unit 113 acquires an image from the image storage unit 117 (S121).

[0108] Next, the position and attitude integration unit 113 acquires the current position and attitude (S122).

[0109] Next, the position and orientation integration unit 113 acquires at least one image and position / orientation of a previously captured path other than the current shooting path (S123). A previously captured path can be generated, for example, from the time-series information of the position and orientation of the imaging unit 111 obtained by SLAM. The information of the previously captured path is stored, for example, in the position and orientation storage unit 118. Specifically, the SLAM results are saved for each shooting trial, and in the Nth (N is a natural number) shooting (the current path), the three-dimensional coordinate axes of the Nth shooting are integrated with the three-dimensional coordinate axes of the results from the 1st to N-1th trials (past paths).

[0110] Alternatively, location information obtained via GPS or Bluetooth may be used instead of the SLAM results.

[0111] Next, the position and orientation integration unit 113 determines whether integration is possible (S124). Specifically, the position and orientation integration unit 113 determines whether the acquired position and orientation and images of the captured route are similar to the current position and orientation and images. If they are similar, it determines that integration is possible; otherwise, it determines that integration is not possible. More specifically, the position and orientation integration unit 113 calculates feature quantities that represent the characteristics of the entire image from each image and determines whether the images have similar viewpoints by comparing them. Furthermore, if the shooting device 101 has GPS and the absolute position of the shooting device 101 is known, the position and orientation integration unit 113 may use that information to determine whether an image was taken at the same or close position as the current image.

[0112] If integration is possible (Yes in S124), the position and orientation integration unit 113 performs path integration processing (S125). Specifically, the position and orientation integration unit 113 calculates the three-dimensional relative position between the current image and a reference image taken of a similar region to the current image. The position and orientation integration unit 113 calculates the coordinates of the current image by adding the calculated three-dimensional relative position to the coordinates of the reference image.

[0113] Figure 7 is a plan view showing the shooting process in the target space according to the embodiment. Path C is the path of camera A, which has already captured images and whose position and orientation have been estimated. The figure also shows the case where camera A is located at a predetermined position on path C. Path D is the path of camera B, which is currently capturing images. At this time, camera B and camera A are capturing images with similar fields of view. While this example shows two images obtained from different cameras, two images captured at different times by the same camera may also be used. For example, camera A and camera B may each be imaging units 111.

[0114] Figure 8 is a diagram illustrating an example of an image and comparison process according to the embodiment. Specifically, Figure 8 is a diagram showing an example of an image and comparison process when captured by cameras A and B shown in Figure 7.

[0115] As shown in Figure 8, the position and orientation integration unit 113 extracts features such as ORB (Oriented Fast and Rotated Brief) features from each image and extracts the features of the entire image based on their distribution or number. For example, the position and orientation integration unit 113 clusters the features that appear in the image, such as a Bag of Words, and uses the histogram for each class as features.

[0116] The position and orientation integration unit 113 compares the overall feature quantities of images between them, and if it determines that the images depict the same location, it calculates the relative three-dimensional position between each camera by performing feature point matching between the images. In other words, the position and orientation integration unit 113 searches for images with similar overall feature quantities from multiple images of the captured path. Based on this relative positional relationship, the position and orientation integration unit 113 converts the three-dimensional position of path D into the coordinate system of path C. This makes it possible to represent multiple paths in a single coordinate system. In this way, the position and orientation of multiple paths can be integrated.

[0117] Furthermore, if the imaging device 101 has a sensor capable of detecting absolute position, such as a GPS, the position and attitude integration unit 113 may perform integration processing using the detection results. For example, the position and attitude integration unit 113 may perform processing using the sensor detection results without performing the above-mentioned image processing, or it may use the sensor detection results in addition to the image. For example, the position and attitude integration unit 113 may use GPS information to narrow down the images to be compared. Specifically, the position and attitude integration unit 113 may set images with position and attitudes within a range of ±0.001 degrees or less from the current camera position's latitude and longitude using GPS as the images to be compared. This reduces the amount of processing required.

[0118] Figure 9 is a flowchart showing the accuracy calculation process according to the embodiment. Specifically, Figure 9 is a flowchart detailing step S104.

[0119] First, the region detection unit 114 acquires an image from the image storage unit 117, acquires the position and orientation from the position and orientation storage unit 118, acquires a three-dimensional point cloud from the position and orientation storage unit 118 that shows the three-dimensional positions of feature points generated by SLAM or the like, and calculates the normal direction of each three-dimensional point in the three-dimensional point cloud (S131). In other words, the region detection unit 114 calculates the normal direction of each three-dimensional point using the position and orientation information, feature point matching information, and the generated three-dimensional point cloud.

[0120] In this context, the normal direction of a three-dimensional point refers to a predetermined orientation that corresponds to each three-dimensional point in a three-dimensional point cloud.

[0121] Figures 10 and 11 illustrate a method for calculating the normal vector of a three-dimensional point according to an embodiment. Specifically, Figures 10 and 11 illustrate a method for calculating the normal direction of a three-dimensional point P generated based on images A and B.

[0122] As shown in Figure 10, first, the region detection unit 114 calculates the normal vector of the three-dimensional point P for each image, based on the position and orientation of the imaging unit 111 when generating each image used to generate the three-dimensional point P.

[0123] Specifically, the region detection unit 114 calculates the normal vector A of a three-dimensional point P based on image A. The normal direction of normal vector A is, for example, the opposite direction to the normal direction of image A. More specifically, the normal direction of normal vector A is the opposite direction to the imaging direction of the imaging unit 111 when image A was generated.

[0124] Similarly, the region detection unit 114 calculates the normal vector B of the three-dimensional point P based on image B. The normal direction of normal vector B is, for example, the opposite direction to the normal direction of image B. Specifically, the normal direction of normal vector B is the opposite direction to the imaging direction of the imaging unit 111 when image B was generated.

[0125] Thus, the normal direction of the three-dimensional point corresponding to each image is the direction obtained by reversing the imaging direction of the imaging unit 111 (in other words, the optical axis direction).

[0126] The image used to calculate the normal direction of a three-dimensional point is, for example, the image used for feature point matching of that three-dimensional point based on a covisibility graph or the like.

[0127] Next, the region detection unit 114 determines a single normal vector for a three-dimensional point based on the normal vectors of the three-dimensional points corresponding to each image. Specifically, as shown in Figure 11, the region detection unit 114 calculates the normal vector P of the three-dimensional point P as the composite vector of normal vector A and normal vector B. Specifically, the region detection unit 114 determines the normal direction of the three-dimensional point P by calculating the normal vector P of the three-dimensional point P by averaging the normal direction vectors of the imaging unit 111 corresponding to the three-dimensional point P (i.e., the vector of the imaging direction of the imaging unit 111) using a covisibility graph in Visual SLAM or SfM.

[0128] As a result, the region detection unit 114 calculates the normal direction (predetermined orientation) of each three-dimensional point in the three-dimensional point cloud. More specifically, the region detection unit 114 generates orientation information indicating the normal direction of the three-dimensional point.

[0129] Furthermore, as described above, for example, multiple three-dimensional points are generated based on multiple images of the imaging unit 111 obtained at different positions of the imaging unit 111. Also, for example, the region detection unit 114 calculates the normal direction of the three-dimensional point from the synthesis of the respective imaging directions when the imaging unit 111 captures multiple images. For example, the normal direction is obtained by synthesizing the directions in the opposite direction of each of the aforementioned imaging directions.

[0130] The sensor used to obtain the detection results used to generate multiple three-dimensional points does not have to be the imaging unit 111, i.e., an image sensor. For example, a LiDAR sensor, a depth sensor, an image sensor, or a combination thereof may be used as the sensor. For example, each of the multiple three-dimensional points is generated by detecting light from an object in multiple different directions from multiple different positions using a sensor. For example, the direction of the composite of the multiple different directions used to generate a three-dimensional point with a normal direction is opposite to the normal direction.

[0131] Generally, sensor specifications include the measurable range (the distance at which a three-dimensional point can be generated). The measurable range may differ for each sensor. Therefore, it is conceivable to generate multiple three-dimensional points by using different sensors depending on the distance from the measurement position. In this case, when presenting the area to be photographed to the user, the system may also present to the user which sensor to use depending on the distance to the area to be photographed. Specifically, the imaging device 101 may store information regarding the sensor specifications in advance, and the UI unit 115 may display areas where the three-dimensional point generation accuracy is low and the sensor suitable for those areas.

[0132] Furthermore, the synthesis of the shooting direction can be determined, for example, by treating the shooting direction as a vector based on a three-dimensional point. When calculating the synthesis, the vector may be a unit (magnitude) vector, or its magnitude may be the distance between the shooting position and the three-dimensional point.

[0133] Figure 9 is calculated again, and after step S131, the region detection unit 114 divides the three-dimensional space (the virtual space on the computer where three-dimensional points are located and a three-dimensional model is generated) into multiple voxels (small spaces) and references the three-dimensional points for each voxel (S132). In other words, the region detection unit 114 references the three-dimensional points contained within each voxel.

[0134] Next, the region detection unit 114 classifies the three-dimensional points for each voxel based on the normal direction of the three-dimensional point (S133). Specifically, the region detection unit 114 classifies the three-dimensional points contained within each voxel based on the position of the surface that restricts the voxel, the position information of the three-dimensional point, and the orientation information.

[0135] Figures 12, 13, and 14 are diagrams illustrating a method for classifying three-dimensional points according to an embodiment.

[0136] First, the region detection unit 114 reflects reprojection error information to each three-dimensional point in the three-dimensional point cloud, for example. Three-dimensional points are generated from feature points in the image. When the generated three-dimensional points are projected onto the image, the region detection unit 114 calculates the amount of deviation between the projected three-dimensional points and the reference feature points. This amount of deviation is the reprojection error, and the region detection unit 114 reflects the reprojection error information to each three-dimensional point by calculating the reprojection error for each three-dimensional dimension. Accuracy can be evaluated using the reprojection error. Specifically, the larger the reprojection error, the lower the accuracy is judged to be, and the smaller the reprojection error, the higher the accuracy can be judged to be.

[0137] Next, as shown in Figure 12, the region detection unit 114 divides the three-dimensional space into multiple voxels. As a result, each three-dimensional point is contained within one of the multiple voxels.

[0138] The size of the voxels can be arbitrarily changed depending on the arrangement of the three-dimensional points. For example, in areas where three-dimensional points are densely clustered, the voxel size may be reduced to subdivide the area more finely compared to areas with fewer three-dimensional points.

[0139] Next, as shown in Figure 13, the region detection unit 114 classifies the multiple three-dimensional points contained in the voxel into multiple groups. The region detection unit 114 performs this classification for each voxel.

[0140] In the example shown in Figure 13, multiple three-dimensional points are classified into Group A, Group B, and Group C, but the number of groups can be arbitrary.

[0141] The region detection unit 114 classifies the three-dimensional points contained within a voxel into multiple groups based, for example, on the position of the surface that restricts the voxel, the position information of the three-dimensional points, and the orientation information.

[0142] As shown in Figure 14, for example, the region detection unit 114 classifies three-dimensional points into the same group if the voxel faces located in the normal direction of the three-dimensional point are the same. In this way, for example, there is a one-to-one correspondence between groups and voxel faces. In the example shown in Figure 14, three-dimensional points belonging to group A are grouped as a group corresponding to the top face of the voxel (the face located at the top of the paper), and three-dimensional points belonging to group B are grouped as a group corresponding to the front face of the voxel (the face located at the front of the paper).

[0143] Thus, for example, in classifying three-dimensional points, the region detection unit 114 divides the three-dimensional space into multiple voxels and classifies each of the multiple three-dimensional points into multiple groups based on the normal direction of each three-dimensional point and the voxel in which the three-dimensional point is contained. Specifically, for example, in classifying three-dimensional points, the region detection unit 114 classifies three-dimensional points that are contained in the same voxel and whose faces among the multiple faces that define the same voxel are located in the normal direction of the three-dimensional point into the same group. For example, suppose that the first voxel among the multiple voxels contains the first three-dimensional point and the second three-dimensional point among the multiple three-dimensional points. Also, suppose that the line extending in the normal direction of the first three-dimensional point passes through the first face among the multiple planes that define (form) the first voxel. In other words, suppose that the first face is located in the first normal direction as seen from the first three-dimensional point. In this case, for example, if the line extending in the normal direction of the second three-dimensional point passes through the first plane, the region detection unit 114 classifies the first three-dimensional point and the second three-dimensional point into the same group. On the other hand, for example, if the line extending in the normal direction of the second three-dimensional point does not pass through the first plane, the region detection unit 114 does not classify the first three-dimensional point and the second three-dimensional point into the same group, but classifies them into different groups.

[0144] Furthermore, the method for classifying multiple three-dimensional points is not limited to a classification method using voxels. For example, the region detection unit 114 may calculate the dot product of the normal vectors of two three-dimensional points and classify the two three-dimensional points into the same group if the dot product is greater than or equal to a predetermined value, or into different groups if the dot product is less than the predetermined value. Also, the method for classifying multiple three-dimensional points is not limited to the result of comparing the normal directions of two three-dimensional points. For example, the region detection unit 114 may classify multiple three-dimensional points based on which of these predetermined directions the normal direction of the three-dimensional point aligns with, such as upward, downward, leftward, or rightward.

[0145] Referring again to Figure 9, after step S133, the region detection unit 114 generates region information indicating the calculated accuracy by calculating the accuracy for each classified group (S134). For example, the region detection unit 114 calculates the accuracy for each group based on the reprojection error or density of the three-dimensional points belonging to each group. For example, the region detection unit 114 calculates a group to which three-dimensional points whose reprojection error is lower than a predetermined error belong as a high-accuracy group. On the other hand, for example, the region detection unit 114 calculates a group to which three-dimensional points whose reprojection error is greater than or equal to a predetermined error belong as a low-accuracy group.

[0146] Figure 15 is a diagram illustrating the method for calculating the accuracy of a group according to an embodiment.

[0147] In the example shown in Figure 15, when three-dimensional points (in this example, three three-dimensional points) belonging to a group corresponding to a voxel face are viewed from a predetermined position (for example, a location located in the direction of the normal to the face), the three-dimensional points represented in darker colors have low reprojection errors (i.e., high accuracy), while the three-dimensional points represented in lighter colors have high reprojection errors (i.e., low accuracy). In this case, for example, since more than half of the multiple three-dimensional points belonging to the group have an accuracy of or higher than the predetermined level, the accuracy of the group when viewed from a predetermined position in the voxel space shown in Figure 17 can be determined to be high.

[0148] Thus, for example, in calculating accuracy, the region detection unit 114 calculates the accuracy of each of the multiple groups based on the reprojection error that indicates the accuracy of one or more three-dimensional points belonging to that group.

[0149] As described above, the group to which a three-dimensional point belongs is determined based on, for example, positional information, orientation information, and the position of the voxel faces. The accuracy calculated for each group reflects the accuracy when viewing one or more three-dimensional points located inside the voxel from outside the voxel, through each voxel face corresponding to the group. In other words, the accuracy of a group can be calculated based on the position and orientation from which the three-dimensional model is viewed.

[0150] The region detection unit 114 may calculate accuracy based on all three-dimensional points belonging to the classified group, or it may calculate accuracy based on some of the three-dimensional points belonging to the classified group. For example, in calculating accuracy, the region detection unit 114 extracts a predetermined number of three-dimensional points from one or more three-dimensional points belonging to each of the multiple groups, and calculates the accuracy of the group based on the accuracy of each of the extracted predetermined number of three-dimensional points. Information indicating the predetermined number, etc., is stored in advance in a storage device such as an HDD provided by the imaging device 101. The predetermined number can be arbitrarily determined and is not particularly limited.

[0151] Furthermore, the accuracy of a group may be calculated in two stages, such as being above a predetermined accuracy or below a predetermined accuracy, or it may be determined in three or more stages.

[0152] Furthermore, if multiple three-dimensional points belong to the same group, the accuracy calculation may use the maximum reprojection error of the multiple three-dimensional points, the minimum reprojection error of the multiple three-dimensional points, the average reprojection error of the multiple three-dimensional points, or the median reprojection error of the multiple three-dimensional points. Also, for example, if the average reprojection error of multiple three-dimensional points is used, a predetermined number of three-dimensional points may be extracted (sampled) from the multiple three-dimensional points, and the average reprojection error of the extracted number of three-dimensional points may be calculated.

[0153] Furthermore, if there are no three-dimensional points corresponding to the faces of a voxel, the accuracy of the corresponding points to those faces may be calculated to be low, or it may be determined that a three-dimensional model cannot be generated when viewing a three-dimensional model located inside the voxel from outside the voxel through the face in question.

[0154] Furthermore, the region detection unit 114 may calculate accuracy based on information obtained from a depth sensor or LiDAR, etc., in addition to three-dimensional points. For example, the region detection unit 114 may calculate accuracy based on information indicating the number of overlaps of three-dimensional points obtained from the depth sensor or accuracy information such as reflection intensity. For example, the region detection unit 114 may calculate accuracy using a depth image obtained from an RGB-D sensor, etc. For example, the accuracy of the generated three-dimensional points tends to be higher the closer the distance from the sensor is, and lower the accuracy as the distance increases. Therefore, for example, the region detection unit 114 may calculate accuracy according to depth. Also, for example, the region detection unit 114 may calculate (determine) accuracy according to the distance to the region. For example, the region detection unit 114 may determine that accuracy is higher the closer the distance is. For example, the relationship between distance and accuracy may be defined linearly, or other definitions may be used.

[0155] Furthermore, if the imaging unit 111 is a stereo camera, the region detection unit 114 can generate a depth image from the disparity image, and may calculate the accuracy in the same way as when using a depth image. In this case, the region detection unit 114 may determine that the accuracy is low for three-dimensional points located in regions where depth values ​​could not be calculated from the disparity image. Alternatively, the region detection unit 114 may estimate the depth value from surrounding pixels for three-dimensional points located in regions where depth values ​​could not be calculated. For example, the region detection unit 114 calculates the average value of 5x5 pixels centered on the target pixel.

[0156] Next, the region detection unit 114 outputs the generated region information (S135). This region information may be, for example, an image in which information indicating the accuracy of each group corresponding to each voxel and each voxel face is superimposed on an image captured by the imaging unit 111, or it may be information in which each piece of information is arranged on a three-dimensional model such as a three-dimensional map.

[0157] The region detection unit 114 may also calculate the opposite direction of the combined direction of the normal directions of two or more three-dimensional points belonging to the first group, which is included in multiple groups. For example, the region detection unit 114 may include the calculation result in the region information as the normal direction of the group.

[0158] Figure 16 is a flowchart of the display process according to the embodiment. Specifically, Figure 16 is a flowchart detailing step S105.

[0159] First, the UI unit 115 checks if there is any information to display (display information) (S141). Specifically, the UI unit 115 checks if there are any newly added images in the image storage unit 117, and if there are, it determines that there is display information. The UI unit 115 also checks if there is any newly added information in the area information storage unit 119, and if there are, it determines that there is display information.

[0160] If there is display information (Yes in S141), the UI unit 115 acquires the display information, such as the image and area information (S142).

[0161] Next, the UI unit 115 displays the acquired display information (S143).

[0162] Next, we will explain specific examples of the display information (more specifically, information indicating the accuracy for each group) that is displayed in the UI section 115.

[0163] Figure 17 shows an example of voxel display according to the embodiment. Figure 17 shows an example in which the accuracy is displayed for each surface that controls the voxel.

[0164] As shown in Figure 17, for example, the UI unit 115 displays the accuracy for each group on the surface that controls the voxels using a predetermined color or density (in Figure 17, the density of the hatching) according to the calculated accuracy. For example, the surface corresponding to a group with an accuracy of a predetermined level or higher and the surface corresponding to a group with an accuracy of less than a predetermined level are displayed in different colors.

[0165] Thus, for example, each of the multiple groups corresponds to one of the multiple faces that define a voxel. The UI unit 115 classifies each of the multiple three-dimensional points into a group corresponding to a face (passing face) through which a line extending in the normal direction of the three-dimensional point passes, from among the multiple faces that define the voxel containing the three-dimensional point among the multiple voxels. The UI unit 115 also displays the passing face in a color corresponding to the accuracy of the group corresponding to the passing face. Specifically, for example, the UI unit 115 displays the passing face in a different color depending on whether the accuracy of the group corresponding to the passing face is above a predetermined accuracy or below a predetermined accuracy. Threshold information indicating the predetermined accuracy, etc., is stored in advance in a storage device such as an HDD provided by the imaging device 101. The predetermined accuracy can be arbitrarily determined in advance and is not particularly limited.

[0166] Figures 18, 19, 20, and 21 show examples of UI screen displays according to the embodiment.

[0167] In the examples shown in Figures 18-21, it is assumed that multiple three-dimensional points are classified into groups based on the voxels they contain.

[0168] Furthermore, in Figures 18 and 19, information indicating accuracy is superimposed on one of the multiple images used to generate the three-dimensional points.

[0169] Thus, for example, the UI unit 115 may superimpose and display the accuracy of multiple groups calculated by the region detection unit 114 onto an image captured by the imaging unit 111 and used to generate multiple three-dimensional points. For example, the UI unit 115 may superimpose and display the accuracy calculated by the region detection unit 114 onto an image being captured by the imaging unit 111. Alternatively, for example, the UI unit 115 may superimpose and display the accuracy of multiple groups calculated by the region detection unit 114 onto a map of the target space. Map information showing the map of the target space is pre-stored in a storage device such as an HDD provided by the imaging device 101.

[0170] Furthermore, in Figures 20 and 21, information indicating accuracy is superimposed on an image showing an overhead view (a top-down view of the three-dimensional model) of the three-dimensional model generated by the position and orientation estimation unit 112.

[0171] Thus, for example, the UI unit 115 displays a three-dimensional model (second three-dimensional model) with lower resolution than the three-dimensional model (first three-dimensional model) composed of multiple three-dimensional points, by superimposing the accuracy of multiple groups calculated by the region detection unit 114 onto the three-dimensional model composed of at least some of the multiple three-dimensional points. In other words, the UI unit 115 may display voxels as a so-called 3D viewer. The second three-dimensional model is, for example, a three-dimensional model generated by the position and orientation estimation unit 112. Alternatively, for example, the UI unit 115 displays an overhead view of the three-dimensional model generated by the position and orientation estimation unit 112 by superimposing the accuracy of multiple groups calculated by the region detection unit 114 onto it.

[0172] In the example shown in Figure 18, the UI unit 115 displays only the voxels corresponding to the group where three-dimensional points have been generated but the calculated accuracy is less than a predetermined accuracy, by adding color (hatching in Figure 18) or the like. On the other hand, for example, for the group where the calculated accuracy is equal to or greater than the predetermined accuracy, only the edges of the voxels are displayed to show the shape of the voxels. Areas where no voxels are displayed are, for example, areas where no three-dimensional points have been generated.

[0173] According to this, users can easily identify areas with high accuracy and areas with low accuracy.

[0174] Furthermore, in the example shown in Figure 19, the UI unit 115 displays voxels corresponding to groups where the calculated accuracy is a predetermined second accuracy or higher by adding color (hatching in Figure 18), etc. On the other hand, for example, voxels corresponding to groups where the calculated accuracy is less than a predetermined second accuracy display only the edges of the voxels to indicate their shape. The UI unit 115 also displays voxels corresponding to groups where the calculated accuracy is a predetermined first accuracy or higher with a different color or intensity (hatching intensity in Figure 19) than voxels corresponding to groups where the calculated accuracy is less than a predetermined first accuracy.

[0175] Thus, for example, the UI unit 115 has multiple groups, each corresponding to one of multiple voxels. For example, each group corresponds to a voxel containing a three-dimensional point belonging to that group. For example, the UI unit 115 displays multiple voxels in a color corresponding to the precision of the group that the voxel corresponds to.

[0176] According to this, users can easily determine their priorities regarding the areas they should photograph.

[0177] Furthermore, the UI unit 115 may color-code voxels (or voxel faces) based on thresholds as described above, or it may display voxels with a gradient depending on the accuracy.

[0178] Furthermore, directions that have not been photographed (unphotographed directions) within the target area may be visualized. For example, the UI section 115 may indicate that the back side of a photographed area has not been photographed. In the example shown in Figure 19, information indicating areas that have not been photographed may be displayed, such as the left side of the dark voxel on the left side of the page not being photographed, the right side of the slightly darker voxel on the right side of the page not being photographed, and the back side of the white voxel in the center of the page not being photographed.

[0179] In the example shown in Figure 20, the UI unit 115 displays only the voxels corresponding to the group whose calculated accuracy is equal to or greater than a predetermined accuracy, by adding color (hatching in Figure 19), etc. On the other hand, for example, for the group whose calculated accuracy is less than a predetermined accuracy, only the edges of the voxels are displayed to show the shape of the voxels.

[0180] According to this, users can easily identify areas with high accuracy and areas with low accuracy.

[0181] In the example shown in Figure 21, the UI unit 115 displays only the voxels corresponding to the group whose calculated precision is less than a predetermined precision, adding color (hatching in Figure 19), etc. On the other hand, for example, voxels corresponding to the group whose calculated precision is equal to or greater than a predetermined precision are not displayed.

[0182] Thus, for example, the UI unit 115 displays only the voxels that correspond to the group whose precision is less than a predetermined precision.

[0183] According to this, users can easily identify areas with low accuracy.

[0184] As mentioned above, the information indicating accuracy may be represented by color, text, or symbols. In other words, the information should be such that the user can visually judge each area.

[0185] This allows users to understand where to take photos simply by looking at the image displayed by the UI unit 115, thus avoiding missed shots.

[0186] Furthermore, for areas that you want to draw the user's attention to, such as low-resolution areas, you may use visual cues that are easy for the user to notice, such as flashing lights. Also, the image on which the area information is superimposed may be a past image.

[0187] Furthermore, the UI unit 115 may overlay information indicating accuracy onto the image being captured. In this case, the area required for display can be reduced, making it easier to see even on small devices such as smartphones. Therefore, these display methods may be switched depending on the type of device.

[0188] Furthermore, the UI unit 115 may display information in text or voice, such as how many meters in front of the current position a low-precision area occurred. This distance can be calculated from the position estimation results. Using text information ensures that the user is notified correctly. Using voice notification allows for safe notification as it eliminates the need to look away during recording.

[0189] Alternatively, instead of displaying the information in two dimensions on the image, the region information may be superimposed onto the real world using AR (Augmented Reality) glasses or a HUD (Head-Up Display). This would increase the compatibility with the image as seen from the real world's perspective, making it possible to intuitively present the user with the areas that should be photographed.

[0190] Furthermore, the accuracy display may be changed in real time as the shooting viewpoint shifts.

[0191] This allows the user to easily understand the captured areas, etc., while referring to the image from the current viewpoint of the imaging unit 111.

[0192] Furthermore, the imaging device 101 may notify the user that the position and orientation estimation has failed if it fails to do so.

[0193] This allows users to quickly restart if they make a mistake.

[0194] Furthermore, the imaging device 101 may detect when the user returns to the failure position and notify the user of this fact using text, images, sound, or vibration. For example, the user returning to the failure position can be detected by using the features of the entire image.

[0195] Furthermore, if the imaging device 101 detects a low-resolution area, it may instruct the user to take another image and specify a different imaging method. Here, the imaging method could be, for example, to capture a larger image of that area. For example, the imaging device 101 may give this instruction using text, images, sound, or vibration.

[0196] This improves the quality of the acquired data, thereby improving the accuracy of the generated three-dimensional model.

[0197] Furthermore, the calculated accuracy may be applied not by superimposing it onto the image being captured, but by superimposing the area information onto an image from a third-party perspective, such as a plan view or perspective view. For example, if a three-dimensional map is available, the captured areas may be superimposed onto the three-dimensional map.

[0198] Furthermore, if a CAD or other map of the target space exists at the construction site, or if a 3D model has been created at the same location previously and 3D map information has already been generated, the imaging device 101 may utilize it.

[0199] Furthermore, if GPS or similar technology is available, the imaging device 101 may create map information based on the latitude and longitude information obtained from GPS.

[0200] In this way, by overlaying the estimated position and orientation of the imaging unit 111 with the captured area onto a three-dimensional map and displaying it as an overhead view, users can easily understand which areas have been captured and which paths were taken. For example, it is possible to make users aware of missed shots, such as the back of a pillar not being captured, which would be difficult to see when the image is displayed during capture, thus enabling more efficient shooting.

[0201] Furthermore, even in environments without CAD software, a similar third-person perspective display can be used without a map. In this case, visibility will be reduced because there are no reference objects, but the user can understand the positional relationship between the current shooting location and the low-precision area.

[0202] Furthermore, users can see when areas where objects are expected to be present are not being captured. Therefore, even in this case, the efficiency of the shooting process can be improved.

[0203] Furthermore, while a plan view example is shown here, map information from other viewpoints may also be used. Additionally, the imaging device 101 may have a function to change the viewpoint of the three-dimensional map. For example, a user interface (UI) may be used to allow the user to change the viewpoint.

[0204] Furthermore, for example, the UI unit 115 may display information indicating the normal directions of multiple groups included in the region information. This allows the user to easily understand the shooting direction that should be detected in order to improve the accuracy of the three-dimensional points.

[0205] [Effects, etc.] As described above, the calculation device according to this embodiment performs the processing shown in Figure 22.

[0206] Figure 22 is a flowchart of the calculation method according to the embodiment.

[0207] First, the calculation device acquires multiple three-dimensional points that represent an object in computer space and each point indicates the position of the object (S151). The calculation device is, for example, the imaging device 101 described above. The sensor is, for example, the imaging unit 111 described above. Note that the sensor can be any sensor that can obtain detection results for generating three-dimensional points, such as a LiDAR sensor.

[0208] Next, the calculation device classifies each of the multiple three-dimensional points into multiple groups based on the normal direction of the three-dimensional point (S152).

[0209] Next, the calculation device calculates a second accuracy such that if the first accuracy of at least one three-dimensional point belonging to each of the multiple groups is high, the second accuracy of that group will be high (S153).

[0210] Here, each of the multiple three-dimensional points is generated by detecting light from an object using sensors from multiple different positions and in multiple different directions. The light detected by the sensors can be any wavelength, such as visible light or near-infrared light.

[0211] Furthermore, the normal direction is determined based on multiple different directions used to generate the three-dimensional point having that normal direction.

[0212] According to this, depending on the direction in which the sensor detects light from an object, multiple three-dimensional points can be classified into multiple groups, and the accuracy of each group to which a three-dimensional point belongs can be calculated according to the accuracy of the three-dimensional point belonging to each of the multiple groups. In other words, according to the calculation device according to one aspect of this disclosure, since multiple three-dimensional points are classified based on the normal direction of each three-dimensional point, the accuracy of each normal direction used for classification can be calculated. For example, according to this, by detecting an object in the normal direction corresponding to a low-accuracy group, for example, the opposite direction of the normal direction of any three-dimensional point belonging to that group, or the opposite direction of the average normal direction of multiple three-dimensional points belonging to that group, a highly accurate three-dimensional point (more specifically, position information of a three-dimensional point) can be generated. In other words, according to the calculation method according to one aspect of this disclosure, the detection direction of the sensor that detects information for generating a highly accurate three-dimensional point can be appropriately calculated.

[0213] For example, if the sensor is the imaging unit 111 as described above, the calculation device estimates the position and orientation of the imaging unit 111 while capturing images with the imaging unit 111, determines the captured areas and accuracy when viewing the target space from multiple directions rather than just one direction based on the estimation result, and presents the captured areas and their accuracy in three dimensions during the capture process.

[0214] According to this approach, by presenting failed shots to the user without requiring time-consuming processing such as generating 3D models, stable acquisition of shooting data (images) is possible, significantly reducing the time required for reshoots and other related tasks. Furthermore, users can gain a three-dimensional understanding of the captured space and the accuracy of the generated 3D points within that space.

[0215] Furthermore, for example, in the above classification (S151), the calculation device divides the space on the computer into multiple subspaces and classifies each of the multiple three-dimensional points into multiple groups based on the normal direction of each three-dimensional point and the subspace that contains that three-dimensional point. A subspace is, for example, the voxel described above.

[0216] This allows us to calculate the accuracy for each normal direction and for each location.

[0217] Furthermore, for example, the first of several small spaces contains the first and second three-dimensional points among several three-dimensional points. Also, for example, a line extending in the direction normal to the first three-dimensional point passes through the first plane among several planes defining the first small space. Also, for example, in the above classification, if a line extending in the direction normal to the second three-dimensional point passes through the first plane, the calculation device classifies the first three-dimensional point and the second three-dimensional point into the same group. The multiple planes defining the small space are, for example, the faces of the voxels described above.

[0218] According to this, multiple three-dimensional points can be classified into multiple groups according to the plane corresponding to the normal direction of the three-dimensional point.

[0219] Furthermore, for example, the sensor may be a LiDAR sensor, a depth sensor, an image sensor, or a combination of these.

[0220] For example, the sensor is an image sensor, that is, an imaging device (camera), and multiple three-dimensional points are generated based on multiple images of the imaging device obtained at different positions of the imaging device, and the normal direction is determined by combining the respective shooting directions when the imaging device takes multiple images. The imaging device is, for example, the imaging unit 111 described above.

[0221] The generation of three-dimensional points based on images captured by an imaging device is simpler to process compared to laser measurements such as LiDAR. Therefore, even with easily generated three-dimensional points, the normal direction of the three-dimensional points can be appropriately determined based on the imaging direction of the imaging device.

[0222] Furthermore, for example, the combined direction of multiple different directions used to generate a three-dimensional point with a normal direction is opposite to the normal direction.

[0223] For example, if the sensor is an imaging device, the normal direction can be determined by combining the directions in the opposite direction of each of the aforementioned imaging directions.

[0224] Furthermore, for example, each of the multiple groups corresponds to one of the multiple planes that define the multiple subspaces. Also, for example, in the above classification, the calculation device classifies each of the multiple three-dimensional points into a group corresponding to a plane through which a line extending in the normal direction of the three-dimensional point passes, among the multiple planes that define the subspace containing the three-dimensional point. Furthermore, the calculation device displays the plane through which the point passes in a color corresponding to the accuracy of the group that the plane through which it passes.

[0225] According to this, the passing surface is displayed in a color corresponding to the accuracy of the group that corresponds to that passing surface, thus providing a display that helps the user adjust the direction of the sensor they are operating.

[0226] Furthermore, for example, the calculation device displays the passing surface in a different color depending on whether the accuracy of the group corresponding to that passing surface is above a predetermined accuracy or whether the accuracy of the group corresponding to that passing surface is below a predetermined accuracy.

[0227] According to this, the passing surface is displayed in a different color depending on whether the accuracy of the group corresponding to that passing surface is above a predetermined accuracy or below a predetermined accuracy, thus providing a display that helps the user adjust the direction of the sensor they are operating.

[0228] Furthermore, for example, each of the multiple groups corresponds to one of the multiple subspaces. Also, for example, the calculation device displays the multiple subspaces in a color corresponding to the accuracy of the group that corresponds to that subspace.

[0229] According to this, multiple small spaces are displayed in colors corresponding to the accuracy of the group that the small space belongs to, thus providing a display that helps the user adjust the direction of the sensor they are operating.

[0230] Furthermore, for example, the calculation device displays only the small spaces corresponding to a group whose accuracy is less than a predetermined accuracy, out of a group of small spaces.

[0231] According to this, among multiple small spaces, only the small space corresponding to the group whose accuracy is below a predetermined accuracy is displayed, thus providing a display that helps the user adjust the direction of the sensor they are operating.

[0232] Furthermore, for example, in the calculation (S152) described above, the calculation device extracts a predetermined number of three-dimensional points from among the one or more three-dimensional points belonging to each of the multiple groups, and calculates the accuracy of the group based on the accuracy of each of the extracted predetermined number of three-dimensional points.

[0233] This method allows for reducing processing load while appropriately calculating the accuracy of multiple groups.

[0234] Furthermore, for example, the calculation device calculates the accuracy of each of the multiple groups based on the reprojection error that indicates the accuracy of one or more three-dimensional points belonging to that group.

[0235] According to this method, the accuracy of multiple three-dimensional points can be calculated appropriately.

[0236] Furthermore, for example, the calculation device displays a second three-dimensional model with lower resolution than the first three-dimensional model composed of multiple three-dimensional points, by superimposing the accuracy of multiple groups onto the second three-dimensional model, which is composed of at least some of the multiple three-dimensional points. The first three-dimensional model is, for example, a high-resolution three-dimensional model generated by the reconstruction device 102. The second three-dimensional model is, for example, a three-dimensional model generated by the imaging device 101 described above.

[0237] According to this, the accuracy of multiple groups is superimposed on the second three-dimensional model, providing a display that helps the user adjust the direction of the sensor they are operating.

[0238] Furthermore, for example, the calculation device displays the accuracy of multiple groups superimposed on an overhead view of the second three-dimensional model.

[0239] According to this, the accuracy of multiple groups is superimposed on an overhead view of the second three-dimensional model, providing a display that helps users adjust the direction of the sensors they are operating.

[0240] For example, the sensor is an image sensor, and the calculation device displays the image captured by the image sensor (an image generated when the image sensor detects light from an object), which was used to generate multiple three-dimensional points, by superimposing the accuracy of multiple groups onto the image.

[0241] According to this, the accuracy of multiple groups is superimposed on the image captured by the imaging device, providing a display that helps the user adjust the direction of the sensor they are operating.

[0242] Furthermore, for example, the calculation device displays the accuracy of multiple groups superimposed on a map of the target space.

[0243] According to this, the accuracy of multiple groups is superimposed on the map, providing a display that helps users adjust the direction of the sensor they are operating.

[0244] Furthermore, for example, the calculation device displays the accuracy of multiple groups while the sensor is detecting light from an object.

[0245] According to this, a display can be provided to help the user adjust the direction of the sensor they are operating while they are detecting light from an object.

[0246] Furthermore, for example, the calculation device calculates the opposite direction of the combined direction of two or more normal directions of two or more three-dimensional points belonging to the first group, which is included in multiple groups.

[0247] According to this, the direction corresponding to each group within multiple groups can be calculated.

[0248] Furthermore, the calculation device may calculate the opposite direction of the normal direction of a representative point among the two or more three-dimensional points, rather than calculating the opposite direction of the combined direction of the normal directions of two or more three-dimensional points belonging to the first group, which is included in multiple groups.

[0249] For example, the computing device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0250] (Other embodiments) The calculation methods and other details relating to the embodiments of this disclosure have been described above, but this disclosure is not limited to these embodiments.

[0251] For example, in the above embodiment, the voxel (small space) was a rectangular parallelepiped, but it can be any shape, such as a cone shape. Also, multiple voxels may each have the same shape or different shapes.

[0252] Furthermore, the imaging unit 111 may capture visible light images or invisible light images (for example, infrared images). When using infrared images, it becomes possible to capture images even in dark environments such as at night.

[0253] Furthermore, the imaging unit 111 may be a monocular camera or may have multiple cameras, such as a stereo camera. By using a calibrated stereo camera, the accuracy of the position and orientation estimation of the imaging unit 111 in three dimensions, which is estimated by the position and orientation estimation unit 112, can be improved.

[0254] Furthermore, the imaging unit 111 may be a device capable of capturing depth images, such as an RGB-D sensor. In this case, since depth images, which represent three-dimensional information, can be acquired, the accuracy of estimating the position and orientation of the imaging unit 111 can be improved. In addition, the depth images can be used as alignment information when integrating the three-dimensional orientation, as described later.

[0255] Furthermore, the expressions "greater than or equal to" and "less than" mentioned above are used to indicate a comparison based on a threshold or similar boundary, and may be replaced with "greater than" and "less than or equal to" or similar.

[0256] Furthermore, for example, each processing unit included in the imaging device, etc., according to the above embodiment may be implemented as an integrated circuit, which is a Large Scale Integration (LSI). These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip.

[0257] Furthermore, integrated circuit implementation is not limited to LSIs; it may also be achieved using dedicated circuits or general-purpose processors. Alternatively, an FPGA (Field Programmable Gate Array), which can be programmed after LSI manufacturing, or a reconfigurable processor capable of reconfiguring the connections and settings of circuit cells within the LSI, may be used.

[0258] Furthermore, in each of the above embodiments, each component may be implemented by being composed of dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU (Central Processing Unit) or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0259] Furthermore, for example, this disclosure may be implemented as a calculation device or the like that executes a calculation method, etc. Alternatively, this disclosure may be implemented as a program for a computer to execute a calculation method, or as a non-temporary recording medium on which such program is stored.

[0260] Furthermore, the division of functional blocks in a block diagram is just one example; multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple parts, or some functions can be moved to other functional blocks. Additionally, the functions of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.

[0261] Furthermore, for example, the order in which each step in the flowchart is performed is illustrative for the purpose of specifically illustrating this disclosure, and may be in a different order. Also, some of the above steps may be performed simultaneously (in parallel) with other steps.

[0262] Although the calculation methods and other aspects relating to one or more embodiments have been described above based on the embodiments, this disclosure is not limited to these embodiments. Without departing from the spirit of this disclosure, various modifications that a person skilled in the art could conceive of may be applied to these embodiments, and forms constructed by combining components from different embodiments may also be included within the scope of one or more embodiments. [Industrial applicability]

[0263] This disclosure is applicable to imaging devices. [Explanation of Symbols]

[0264] 101 Imaging device 102 Reconfiguration device 111 Imaging Unit 112 Position and orientation estimation unit 113 Position and orientation integration unit 114 Area detection unit 115 UI section 116 Control Unit 117 Image Storage Section 118 Position and orientation storage section 119 Area information storage section

Claims

1. An information processing device comprising a processor and a memory connected to the processor, wherein the processor uses the memory, By acquiring multiple three-dimensional points, Multiple groups are generated based on the normal direction associated with each three-dimensional point. For each of the aforementioned groups, region information is generated that includes information indicating the normal direction corresponding to the group and information indicating the evaluation value corresponding to the group. The evaluation value corresponding to the group is generated based on the evaluation value corresponding to one or more three-dimensional points belonging to that group. Information processing device.

2. An information processing device according to claim 1, wherein one or more three-dimensional points belonging to a group are included in the same small space among a plurality of small spaces generated by dividing space. Information processing device.

3. The information processing apparatus according to claim 2, wherein one or more three-dimensional points belonging to one group are three-dimensional points contained in the same small space, and which are three-dimensional points whose planes, among a plurality of planes defining the small space, are the same plane through which lines extending in the normal direction of each three-dimensional point pass. Information processing device.

4. An information processing device according to claim 1, wherein the evaluation values ​​corresponding to the one or more three-dimensional points are generated based on the reprojection error. Information processing device.

5. An information processing device according to claim 1, wherein the evaluation value corresponding to the group is generated based on the evaluation values ​​corresponding to a predetermined number of three-dimensional points among one or more three-dimensional points belonging to the group. Information processing device.

6. An information processing apparatus according to claim 1, wherein the normal direction corresponding to the group is the opposite direction to the combined direction obtained by combining the normal directions of two or more three-dimensional points belonging to the group. Information processing device.

7. A method implemented by an information processing device, By acquiring multiple three-dimensional points, Multiple groups are generated based on the normal direction associated with each three-dimensional point. For each of the aforementioned groups, region information is generated that includes information indicating the normal direction corresponding to the group and information indicating the evaluation value corresponding to the group. The evaluation value corresponding to the group is generated based on the evaluation value corresponding to one or more three-dimensional points belonging to that group. method.