Three-dimensional model generation device, three-dimensional model generation method, and program
The three-dimensional model generating device and method address the challenge of accumulated errors by displaying guidance for capturing correction images, thereby enhancing the accuracy of the generated models.
Patent Information
- Application Number
- PCT/JP2024/040379
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-22
AI Technical Summary
Existing three-dimensional model generating devices and methods face challenges in reducing accumulated errors, which affect the accuracy of the generated models.
A three-dimensional model generating device and method that acquires consecutive images, generates a three-dimensional model with positional information, and determines whether to display guidance for capturing an image to correct accumulated errors, thereby improving model accuracy.
The proposed solution supports the correction of accumulated errors by providing guidance for capturing specific images, leading to improved accuracy and further enhancements in three-dimensional model generation.
Smart Images

Figure JP2024040379_22052025_PF_FP_ABST
Abstract
Description
Three-dimensional model generating device, three-dimensional model generating method and program
[0001] The present disclosure relates to a three-dimensional model generation device, a three-dimensional model generation method, and a program.
[0002] A method is known in which the position and orientation of a camera are detected from multiple images using Simultaneous Localization and Mapping (SLAM) and a 3D model is generated, and a method is known in which loop closing is used to reduce accumulated errors (see, for example, Patent Documents 1 and 2).
[0003] JP 2023-69019 A JP 2017-111688 A
[0004] Further improvements are desired in such 3D model generation devices and 3D model generation methods. Therefore, the present disclosure provides a 3D model generation device or a 3D model generation method that can achieve further improvements.
[0005] A three-dimensional model generation device according to one aspect of the present disclosure includes a memory and a circuit connected to the memory, wherein the circuit uses the memory to acquire a plurality of consecutive images, generates a three-dimensional model using the plurality of images including a plurality of three-dimensional points, each having positional information, determines whether or not to display first information on a display unit to indicate a first position and orientation for capturing a first image for correcting accumulated error in the three-dimensional model, and, if it is determined that the first information should be displayed on the display unit, causes the first information to be displayed on the display unit.
[0006] The present disclosure can provide a three-dimensional model generation device or a three-dimensional model generation method that can achieve further improvements.
[0007] FIG. 1 is a diagram illustrating a configuration of a 3D model generation system according to an embodiment. FIG. 2 is a flowchart illustrating the operation of a 3D model generation device according to an embodiment. FIG. 3 is a diagram illustrating a camera position and orientation estimation process according to an embodiment. FIG. 4 is a diagram illustrating a 3D model generation process according to an embodiment. FIG. 5 is a diagram illustrating an accumulated error according to an embodiment. FIG. 6 is a diagram illustrating generation of accumulated error information according to an embodiment. FIG. 7 is a diagram illustrating a display example of accumulated error information according to an embodiment. FIG. 8 is a diagram illustrating a display example of accumulated error information according to an embodiment. FIG. 9 is a diagram illustrating a display example of accumulated error information according to an embodiment. FIG. 10 is a diagram illustrating a display example of accumulated error information according to an embodiment. FIG. 11 is a diagram illustrating a display example of a user interface according to an embodiment. FIG. 12 is a diagram illustrating a display example of a user interface according to an embodiment. FIG. 13 is a diagram illustrating a display example of a user interface according to an embodiment. FIG. 14 is a diagram illustrating an example of guide information according to an embodiment. FIG. 15 is a diagram illustrating an example of guide information according to an embodiment. FIG. 16 is a flowchart of a 3D model generation process according to an embodiment. FIG. 17 is a block diagram illustrating a hardware configuration of a 3D model generation device according to an embodiment.
[0008] A three-dimensional model generation device according to one aspect of the present disclosure includes a memory and a circuit connected to the memory, wherein the circuit uses the memory to acquire a plurality of consecutive images, generates a three-dimensional model using the plurality of images including a plurality of three-dimensional points, each having positional information, determines whether or not to display first information on a display unit to indicate a first position and orientation for capturing a first image for correcting accumulated error in the three-dimensional model, and, if it is determined that the first information should be displayed on the display unit, causes the first information to be displayed on the display unit.
[0009] This allows the 3D model generation device to support the capture of the first image for correcting the accumulated error. Therefore, by appropriately performing the correction process, the accuracy of the generated 3D model can be improved. In this way, the 3D model generation device can achieve further improvements.
[0010] For example, the first information may be a second image included in the plurality of images and captured in the first position and orientation.
[0011] This allows the user to determine the first position and orientation while referring to the second image.
[0012] For example, the circuit may cause the second image to be displayed on the display unit in a state where the second image is superimposed on the image currently being captured.
[0013] This allows the user to capture the first image by capturing an image in a state where the currently captured image matches the second image.
[0014] For example, the circuit may further cause the display unit to display second information indicating the accumulated error.
[0015] This allows the user to refer to the second information and determine whether or not to correct the accumulated error.
[0016] For example, the circuit may display, as the second information, a third image of the three-dimensional model viewed from the current position and orientation on the display unit, superimposed on the currently captured image.
[0017] This allows the user to recognize the difference between the currently captured image and the third image, and based on this difference, the user can determine whether or not correction of the accumulated error is necessary.
[0018] For example, the circuit may further cause the display unit to display, as the second information, an area where there is a difference between the third image and the currently captured image.
[0019] This allows the user to determine the need for correction of the accumulated error based on the third information.
[0020] For example, the circuit may further calculate the amount of accumulated error based on the difference between the third image and the currently captured image, and cause the display unit to display third information indicating the calculated amount.
[0021] This allows the user to recognize the magnitude (degree) of the difference between the currently captured image and the third image, and based on the magnitude of the difference, the user can determine whether or not correction of the accumulated error is necessary.
[0022] For example, the circuit may further display, together with the second information, a user interface on the display unit for the user to instruct whether or not to perform the correction, and when an instruction to perform the correction is received, cause the first information to be displayed on the display unit.
[0023] This allows correction to be performed based on the judgment of the user who has referred to the second information.
[0024] For example, the circuit may further display on the display unit a user interface for allowing a user to select the second image from a plurality of fourth images included in the plurality of images.
[0025] This allows the user to select the second image.
[0026] For example, the circuitry may further select the fourth images from the plurality of images based on feature points contained in each of the plurality of images.
[0027] This improves the accuracy of feature point matching for the second image, thereby improving the accuracy of correction of accumulated error.
[0028] For example, the circuitry may further select the second image from the plurality of images based on feature points contained in each of the plurality of images.
[0029] This improves the accuracy of feature point matching for the second image, thereby improving the accuracy of correction of accumulated error.
[0030] For example, when the first information is displayed on the display unit, (1) if the first image taken from the first position and orientation is acquired, the correction is performed using the acquired first image, and (2) if a sixth image taken from among the multiple images at a position and orientation other than the first position and orientation of a fifth image is acquired, the correction of the accumulated error does not need to be performed using the acquired sixth image.
[0031] This can prevent correction from being performed using an inappropriate image.
[0032] For example, the circuit may use the multiple images to estimate an estimated position and orientation, which is the position and orientation at which each of the multiple images was captured, and generate the three-dimensional model using the estimated position and orientation. In correcting the accumulated error, the circuit may correct the accumulated error based on the error between the estimated position and orientation of the first image and the estimated position and orientation of a second image captured at the first position and orientation, which are included in the multiple images.
[0033] For example, the circuit may display the current three-dimensional model and the past three-dimensional model on the display unit as the second information, and the circuit may further accept an operation to align the positions of the current three-dimensional model and the past three-dimensional model, and may perform the correction using information obtained by the operation to align the positions as auxiliary information.
[0034] This reduces the amount of processing required to correct accumulated errors and improves accuracy.
[0035] A three-dimensional model generation method according to one aspect of the present disclosure acquires a plurality of consecutive images, uses the plurality of images to generate a three-dimensional model including a plurality of three-dimensional points, each having positional information, determines whether or not to display first information on a display unit to indicate a first position and orientation for capturing a first image for correcting accumulated error in the three-dimensional model, and, if it is determined that the first information should be displayed on the display unit, displays the first information on the display unit.
[0036] This allows the 3D model generation method to support the capture of images for correcting accumulated errors. Therefore, by appropriately performing correction processing, the accuracy of the generated 3D model can be improved. In this way, the 3D model generation method can be further improved.
[0037] Furthermore, a program according to one aspect of the present disclosure is a program for causing a computer to execute the three-dimensional model generation method.
[0038] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0039] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. However, more detailed explanation than necessary may be omitted. For example, detailed explanation of well-known matters or redundant explanation of substantially the same configuration may be omitted. This is to avoid unnecessary redundancy in the following explanation and to facilitate understanding by those skilled in the art.
[0040] The inventors have provided the accompanying drawings and the following description to enable those skilled in the art to fully understand the present disclosure, and do not intend for them to limit the subject matter described in the claims.
[0041] (Embodiment) [1. Configuration] First, the configuration of a 3D model generation system according to this embodiment will be described. Fig. 1 is a block diagram showing the configuration of the 3D model generation system according to this embodiment. The 3D model generation system is a system that generates a 3D model from a plurality of images, and includes an imaging device 101, a 3D model generation device 102, and a display device 103.
[0042] The imaging device 101 generates a plurality of images and a plurality of pieces of sensor information, and outputs the generated plurality of images and a plurality of pieces of sensor information to the three-dimensional model generation device 102. The imaging device 101 includes a camera 111 and a sensor 112. The camera 111 captures an image of a subject (object) from different viewpoints (image capture positions and image capture directions), and outputs the obtained plurality of images (frames) to the three-dimensional model generation device 102. For example, the imaging device 101 captures images while moving, thereby capturing a plurality of images.
[0043] The sensor 112 includes, for example, at least one of a distance sensor (depth sensor), a laser sensor (e.g., LiDAR), and an IMU (Inertial Measurement Unit). The IMU includes, for example, at least one of an acceleration sensor, a rotational angular acceleration sensor, and a gyro sensor. The sensor 112 outputs sensor information corresponding to each of the multiple images to the three-dimensional model generating device 102. Here, the sensor information includes, for example, at least a portion of a depth image obtained by the distance sensor, point cloud data obtained by the laser sensor, and acceleration information obtained by the IMU. Note that the imaging device 101 does not necessarily have to include the sensor 112.
[0044] The 3D model generating device 102 generates a 3D model using a plurality of images and a plurality of pieces of sensor information. The 3D model generating device 102 includes an acquisition unit 121, a storage unit 122, a camera position and orientation estimation unit 123, a 3D model generating unit 124, an accumulated error detection unit 125, a user guide unit 126, an output unit 127, and a control unit 128.
[0045] The acquisition unit 121 acquires a plurality of images and a plurality of pieces of sensor information from the imaging device 101. The storage unit 122 stores the plurality of images and a plurality of pieces of sensor information acquired by the acquisition unit 121. The storage unit 122 also stores camera position and orientation information, map information, a three-dimensional model, and the like, which will be described later.
[0046] The camera position and orientation estimation unit 123 uses multiple images to generate camera position and orientation information that indicates the position and orientation (direction) of the image capturing device 101 when each image was captured. Here, the orientation of the image capturing device 101 indicates at least one of the image capturing direction of the image capturing device 101 and the tilt of the image capturing device 101. The image capturing direction of the image capturing device 101 is the direction of the optical axis of the image capturing device 101. The tilt of the image capturing device 101 is the rotation angle around the optical axis of the image capturing device 101 from the reference orientation.
[0047] The camera position and orientation estimation unit 123 also generates map information indicating a three-dimensional point group indicating the three-dimensional position together with the camera position and orientation information.
[0048] The three-dimensional model generating unit 124 generates a three-dimensional model of the subject based on a plurality of images, camera position and orientation information, map information, and sensor information.
[0049] The accumulated error detection unit 125 detects accumulated errors (also called cumulative errors) of the three-dimensional model and generates accumulated error information indicating the accumulated errors. The user guide unit 126 generates guide information to support the user in moving to the shooting position of the correction image used in the loop closing process.
[0050] The output unit 127 outputs the current image, accumulated error information, guide information, etc. to the display device 103. The control unit 128 controls each processing unit. For example, the control unit 128 controls each processing unit in response to a user operation, etc.
[0051] The display device 103 displays the current image, accumulated error information, guide information, etc. sent from the 3D model generating device 102. The display device 103 includes a display unit 131 and an operation reception unit 132. The display unit 131 displays the current image, accumulated error information, guide information, etc. The operation reception unit 132 receives user operations. For example, the operation reception unit 132 is a touch panel, or a keyboard and mouse, etc. The operation reception unit 132 may also be a microphone, etc. that receives voice operations. The display device 103 may also include a speaker that outputs voice, a vibrator that transmits vibrations to the user, etc., as means for presenting information to the user.
[0052] 1 may be realized as a single device. Alternatively, multiple processing units included in each device may be distributed across multiple devices.
[0053] For example, in a use case in which a subject is photographed while moving using a terminal carried by a user (such as a smartphone, tablet terminal, or personal computer), the imaging device 101, the 3D model generation device 102, and the display device 103 may be included in the terminal. In another use case, the imaging device 101 may be included in a moving object (such as a vehicle, robot, or drone) operated by a user, and the 3D model generation device 102 and the display device 103 may be included in a terminal carried by the user. In either case, some of the processing units included in each device may be included in another device, such as a server, connected to the terminal operated by the user via a communication network or the like. For example, some or all of the processing units included in the 3D model generation device 102 may be included in the server.
[0054] 1 may be transmitted and received by any method, such as wired or wireless communication. These communications may be performed directly between the devices, or indirectly via other communication devices or servers. These communications may also be performed within a single device.
[0055] 2. Operation Next, a description will be given of the operation of the 3D model generation device 102. For example, the 3D model generation system generates a 3D model in real time based on images captured by the imaging device 101.
[0056] Fig. 2 is a flowchart showing the operation of the three-dimensional model generating device 102. Here, the process shown in Fig. 2 is repeatedly performed, for example, in units of one frame. Note that the process shown in Fig. 2 may also be repeatedly performed in units of multiple frames.
[0057] The multiple images may be multiple still images captured by the imaging device 101 while it is moving, or multiple images included in a moving image captured by the imaging device 101 while it is moving. When a moving image is used, the multiple images do not need to be all images included in the moving image, but may be key images (key frames) that are part of the multiple images included in the moving image. Here, key images are, for example, images extracted at a predetermined time interval from the multiple images included in the moving image.
[0058] [2-1. Estimation of Camera Position and Orientation] First, the acquisition unit 121 acquires a current image (frame) and multiple pieces of sensor information from the image capture device 101 (S101). Next, the camera position and orientation estimation unit 123 generates camera position and orientation information and map information using multiple images including the current image (S102).
[0059] There are no particular limitations on the method used by the camera position and orientation estimation unit 123 to estimate the positions and orientations of the image capture devices 101. For example, the camera position and orientation estimation unit 123 may estimate the positions and orientations of the multiple image capture devices 101 using, for example, Visual-SLAM (Simultaneous Localization and Mapping), Structure-From-Motion, or ICP (Iterative Closest Point).
[0060] 3 is a diagram for explaining the process of estimating the position and orientation of the camera, which shows an example of the process of estimating the position and orientation of the camera using three frames F0, F1, and F2.
[0061] 3, the camera position and orientation estimation unit 123 performs feature point matching processing on multiple frames F0 to F2 captured by the image capture device 101. That is, the camera position and orientation estimation unit 123 extracts feature points from each of the multiple frames F0 to F2, and extracts sets of similar points (corresponding points) that are similar between the multiple frames from the extracted multiple feature points. Next, the camera position and orientation estimation unit 123 uses the extracted sets of similar points to estimate the position and orientation of the image capture device 101 for each frame.
[0062] The map information indicates a plurality of map points, each of which is a three-dimensional point indicating a three-dimensional position. The camera position and orientation estimation unit 123 generates the map information by performing triangulation using the result of the feature point matching process and the camera position and orientation information. Note that the map information may include, in addition to the position information, information indicating the color of each map point and the surface shape around the map point (e.g., information indicating normals). The map information may also be modified to high-precision information by performing an optimization process on the map information together with the camera position and orientation.
[0063] [2-2. Generation of 3D Model] Next, the 3D model generation unit 124 generates a 3D model of the subject based on multiple images including the current image and the camera position and orientation information (S103). For example, the 3D model generation unit 124 generates the 3D model using map information. FIG. 4 is a diagram for explaining the 3D model generation process in this case. For example, as shown in FIG. 4, the 3D model generation unit 124 generates a mesh model represented by a plane with map points as vertices as the 3D model.
[0064] The three-dimensional model is not limited to a mesh model and may be any three-dimensional model. For example, map information may be used as the three-dimensional model as is. Furthermore, if the imaging device 101 includes a sensor 112, the three-dimensional model may be generated using sensor information obtained by the sensor 112 and camera position and orientation information. For example, the three-dimensional model may be generated using point cloud data obtained by LiDAR or a depth image obtained by a range sensor. Furthermore, a combination of the map information, point cloud data, and depth information may be used. For example, the depth information may be estimated from an image using AI (artificial intelligence).
[0065] Furthermore, the three-dimensional model may include attribute information (information indicating color or normals) in addition to position information.
[0066] [2-3. Detection of Accumulated Error] Next, the accumulated error detection unit 125 detects the accumulated error of the current frame and displays accumulated error information indicating the accumulated error on the display unit 131 (S104). Here, the accumulated error is the difference (error) between a 3D model generated using a past frame and a 3D model generated using the current frame. In other words, the accumulated error is the difference (error) between the estimated camera position and orientation of the current frame and the actual camera position and orientation.
[0067] FIG. 5 is a diagram illustrating this accumulated error. FIG. 5 shows an example in which the imaging device 101 captures images of multiple objects while moving. In this example, the imaging device 101 starts capturing images and moving at time t0, and at time t=x, returns to the same position as at time t=0. Here, in estimating the camera position and orientation and generating a 3D model, the estimated camera position and orientation of a previously captured image are used to estimate the camera position and orientation of a newly captured image, and the 3D model is updated (added). Therefore, as the number of images used increases (as the capturing time increases), there is a possibility that the error between the estimated camera position and orientation and the 3D model and the actual camera position and orientation and the position of the object increases. For example, as shown in FIG. 5, as time progresses, the difference (accumulated error) between the actual camera position and orientation (actual trajectory) and the estimated camera position and orientation (estimated trajectory) increases.
[0068] The loop closing process is a process for correcting this accumulated error. By performing the loop closing process, the estimated camera position and orientation (estimated trajectory) is corrected so as to approach the actual camera position and orientation (actual trajectory). Furthermore, the position information of the 3D model is updated based on the corrected camera position and orientation. For example, using the estimated camera positions and orientations of a first image (image at t = 0) captured in the past and a second image (image at t = x) captured at the same position as the first image, the camera positions and orientations of all images included in the range from t = 0 to t = x are corrected so that the camera position and orientation of the second image matches the camera position and orientation of the first image. In other words, to perform the loop closing process, an image captured at the same camera position and orientation as one of the multiple images captured in the past is required.
[0069] An example of accumulated error information will be described below. Fig. 6 is a diagram for explaining the generation of accumulated error information. For example, as shown in Fig. 6, accumulated error information is generated by projecting a past three-dimensional model onto a current image. Here, the past three-dimensional model is, for example, a three-dimensional model generated using a plurality of images from t = 0 to x - 1.
[0070] FIG. 7 is a diagram showing an example of the display of accumulated error information. For example, as shown in FIG. 7, an image in which a past three-dimensional model is projected onto a current image is displayed as the accumulated error information. This allows the user to determine the need for loop closing processing based on the presence or absence of a difference between the current image shown in the image and the image in which the past three-dimensional model is projected, and the magnitude of the difference. Note that the projected three-dimensional model may be displayed semi-transparently, for example. Alternatively, characteristic parts such as the contours of the subject may be extracted, and only those characteristic parts may be displayed.
[0071] The accumulated error detection unit 125 may also calculate the accumulated error. For example, the accumulated error detection unit 125 uses a first image viewed from the current camera position and orientation based on the current image or the current 3D model, and a second image viewed from the current camera position and orientation of a past 3D model, to calculate the difference in pixel value of each pixel between the first image and the second image. Here, the pixel value may be, for example, a depth value, a color value, a normal vector, or the like. The depth value of the first image may be obtained, for example, from the current 3D model, or if the sensor information includes a depth image, the current depth image may be used as the first image. Alternatively, a depth image generated using multiple images including the current image may be used as the first image. Note that the current 3D model is a 3D model generated from multiple images including the current image (e.g., images from t = x - 2 to t = x). The depth value of the second image is generated from a past 3D model.
[0072] Also, when comparing color values, the first image is a current image and the second image is generated from, for example, a past three-dimensional model.
[0073] The accumulated error information may include information indicating the difference calculated in this manner. FIG. 8 is a diagram showing an example of displaying accumulated error information in this case. For example, as shown in FIG. 8, an area where a difference exists may be highlighted. In the example shown in FIG. 8, the magnitude of the difference is indicated by shading. Note that the highlighting method is not limited to this, and any method such as blinking may be used. Also, highlighting is not necessarily required, and any method may be used as long as it allows the user to recognize the area where a difference exists. Also, the number of gradations indicating the magnitude of the difference may be arbitrary. Also, the presence or absence of a difference (whether the difference is equal to or greater than a threshold) may be indicated without using gradation display.
[0074] 8 shows an area where a difference exists, but the object where a difference exists may be highlighted. The calculation of the difference is not limited to pixel-by-pixel, but may be area-by-area including multiple pixels. In this case, the sum, average, median, or maximum value of the differences (or absolute values of the differences) of multiple pixels included in the area may be calculated as the difference of the area.
[0075] Alternatively, the difference of the current entire image may be calculated as the accumulated error. For example, the sum, average, median, or maximum value of the differences of multiple pixels may be calculated as the difference of the entire image. Alternatively, a weighted sum or weighted average value may be used. In this case, for example, a larger weight may be set for pixels closer to the center of the image.
[0076] Weights may also be set according to the accuracy of the three-dimensional model. For example, a higher weight may be set for an area with higher accuracy. This accuracy may be determined, for example, based on the distance between the camera position of the image used to generate the three-dimensional model and the subject. For example, the closer the distance, the higher the accuracy is determined. Alternatively, the accuracy may be determined based on the distance from a detected feature point, vertex, or edge. For example, the closer the distance, the higher the accuracy is determined. The accuracy may also be determined based on whether or not texture is present. For example, if texture is present, the accuracy may be determined to be high. In addition, when a depth image or point cloud data obtained by a laser sensor is used to generate the three-dimensional model, the accuracy of the three-dimensional model may be determined based on information indicating the accuracy of the sensor or the accuracy of the sensor result, which is included in the additional information of the sensor information. The accuracy may also be determined by combining these multiple factors.
[0077] To facilitate the above determination, when generating a 3D model, information indicating the accuracy of each point of the 3D model may be added to the point, and information identifying the frame used to generate the point may be added to the point.
[0078] The accumulated error information may include information indicating the accumulated error of the entire image calculated in this manner. Fig. 9 is a diagram showing an example of how accumulated error information is displayed in this case. For example, as shown in Fig. 9, information 201 indicating the necessity of loop closing processing may be displayed as information indicating the accumulated error of the entire image. Note that the information 201 may indicate the magnitude of the accumulated error instead of the necessity of loop closing processing. Furthermore, here, the information 201 indicates the degree of necessity of loop closing processing (the magnitude of the accumulated error), but it may also indicate whether loop closing processing is necessary (whether an accumulated error exists (or is greater than a threshold)).
[0079] The necessity of loop closing processing may be determined based on the accumulated error as well as the shooting time during which loop closing processing has not been performed (the number of images used to generate the three-dimensional model). For example, the longer the shooting time, the greater the necessity of loop closing processing may be determined.
[0080] Furthermore, the method of notifying (presenting) the user whether or not loop closing processing is necessary (or the degree of necessity) is not limited to this, and any display method may be used, such as changing the color of the screen frame or flashing it. Furthermore, other notification methods such as sound or vibration may be used instead of display. Furthermore, a plurality of these methods may be combined.
[0081] In the above, an example has been described in which a past three-dimensional model is projected onto a current image as the accumulated error information. However, the current three-dimensional model and the past three-dimensional model may also be displayed as the accumulated error information. FIG. 10 is a diagram showing an example of the display of accumulated error information in this case. For example, the current three-dimensional model and the past three-dimensional model may be displayed in different colors. Furthermore, in FIG. 10, only the past three-dimensional model corresponding to the current three-dimensional model is displayed, but the entire three-dimensional model, including three-dimensional models of subjects not included in the current image, may also be displayed as the past three-dimensional model. In this case, the color, etc., of the past three-dimensional model may be changed depending on the distance from the current three-dimensional model.
[0082] The viewpoint of the displayed three-dimensional model may be changeable. For example, as shown in Fig. 10, a user interface 202 may be displayed for changing (translating, rotating, and scaling) the viewpoint of the displayed current three-dimensional model and past three-dimensional models. Note that the user interface 202 does not necessarily have to be displayed, and the viewpoint may be changed by a predetermined operation on a touch panel (such as dragging, flicking, pinching in and out, etc.).
[0083] 10 may be displayed instead of the accumulated error information shown in FIG. 7 or the like, or may be displayed simultaneously on a split screen, or either one may be selectively displayed in response to a user operation. Also, a difference may be displayed as in the example shown in FIG. 8, or information indicating the necessity of loop closing processing (the magnitude of the accumulated error) may be displayed as in the example shown in FIG. 9.
[0084] 10 , the current 3D model and the previous 3D model may be manually aligned. That is, a user may input an operation to align the current 3D model and the previous 3D model. Information based on this manually input operation may be used as auxiliary information for the loop closing process. For example, this operation may obtain the direction and amount of deviation between the current 3D model and the previous 3D model. By using at least one of the direction and amount of deviation to set the initial value for the search in the matching process in the loop closing process, the amount of processing can be reduced and accuracy can be improved.
[0085] The above describes a method for calculating the accumulated error by comparing the pixel values of each pixel between a first image based on the current image or the current 3D model, viewed from the current camera position and orientation, and a second image based on a past 3D model viewed from the current camera position and orientation. However, other methods may be used. For example, the accumulated error detection unit 125 may detect corresponding points, which are corresponding feature points, between the current 3D model and the past 3D model and calculate the accumulated error based on the distance between the corresponding points. In this case, the accumulated error may be calculated on a point-by-point basis, a region including multiple points, or an image-by-image basis. For example, the accumulated error for each region or image may be the sum, average, median, maximum, weighted sum, or weighted average of multiple differences between multiple corresponding points included in the region or image. Note that the weights may be set, for example, as described above.
[0086] Furthermore, instead of detecting corresponding points between three-dimensional models, corresponding points may be detected between a current image and an image of a past three-dimensional model viewed from the current camera position and orientation.
[0087] 2, next, the accumulated error detection unit 125 determines whether or not to perform the loop closing process (S105). Note that this determination is not a determination as to whether or not to immediately perform the loop closing process, but a determination as to whether or not to display guide information for capturing a correction image used in the loop closing process (described later).
[0088] For example, the accumulated error detection unit 125 displays, together with the accumulated error information, a user interface for the user to instruct the execution of the loop closing process. Fig. 11 is a diagram showing a display example of this user interface 203. If the user instructs the execution of the loop closing process via the user interface 203, the accumulated error detection unit 125 determines to execute the loop closing process (Yes in S105), and if not, determines not to execute the loop closing process (No in S105).
[0089] The method by which the user instructs the execution of the loop closing process is not limited to the above, and any method may be used.
[0090] Furthermore, the user interface 203 is not always displayed, but is displayed when the calculated accumulated error (necessity of loop closing processing) is equal to or greater than a predetermined threshold, and may not be displayed in other cases.
[0091] If it is determined that the loop closing process is not to be performed (No in S105), the process from step S101 onwards is performed on the next image.
[0092] [2-5. Generation of Guide Information] On the other hand, if it is determined that the loop closing process is to be performed (Yes in S105), the user guide unit 126 selects a target image (S106). Here, the target image is an image included in past images, and is an image for determining the camera position and orientation of a correction image to be captured. For example, the camera position and orientation of the target image are determined to be the camera position and orientation of the correction image. Note that the camera position and orientation of the correction image do not need to be exactly the same as the camera position and orientation of the target image. For example, it is sufficient if the difference between the camera position and orientation of the correction image and the camera position and orientation of the target image is within a predetermined range.
[0093] First, the user guide unit 126 selects multiple candidate images from multiple past images. For example, the user guide unit 126 selects multiple candidate images from multiple images used to generate a past 3D model (i.e., a 3D model corresponding to the subject in the current image) projected onto the current image when calculating the accumulated error. Note that the multiple candidate images may be selected from multiple images related to the multiple images. Here, the related images are, for example, images with camera positions and orientations similar to those of the multiple images.
[0094] Specifically, the user guide unit 126 preferentially selects an image with a camera position and orientation that is close to the estimated current camera position and orientation. Alternatively, the user guide unit 126 preferentially selects an image with many feature points. Note that the user guide unit 126 may determine the degree of dispersion (degree of variance) of the feature points in addition to the number of feature points, and preferentially select an image with a high degree of dispersion. Alternatively, the user guide unit 126 may preferentially select an image in which the projected past three-dimensional model is displayed in the center. Note that the user guide unit 126 may use any one of these multiple determination criteria, or may use a combination of two or more determination criteria.
[0095] The feature points may be points detected from a three-dimensional model by feature analysis, or points detected from a two-dimensional image by feature analysis, or points in a three-dimensional model corresponding to points detected from a two-dimensional image by feature analysis.
[0096] Next, the user guide unit 126 displays a user interface for the user to select a target image from the selected plurality of candidate images. Fig. 12 is a diagram showing a display example of this user interface 205. The target image is determined by the user selecting the target image from the displayed plurality of candidate images 206. Note that the method by which the user selects the target image is not limited to this.
[0097] 13 is a diagram showing another example of the display of a user interface that allows a user to select a target image from a plurality of candidate images. The user interface 207 shown in FIG. 13 displays a past 3D model, the camera positions and orientations of the current image, and the camera positions and orientations of a plurality of candidate images. In this case, the candidate image is determined by the user selecting one of the camera positions and orientations of the plurality of candidate images. Note that the user interface 207 shown in FIG. 13 may be used instead of the user interface 205 shown in FIG. 12, or both may be displayed simultaneously.
[0098] Alternatively, instead of the user selecting a target image, the user guide unit 126 may select the target image from a plurality of images. In this case, for example, the user guide unit 126 selects the image with the highest priority as the target image in the above-described method for selecting candidate images.
[0099] Next, the user guide unit 126 generates guide information for the user to capture the target image and displays the generated guide information (S107). Here, the guide information is information for assisting the user in capturing the correction image, and is information for guiding the camera position and orientation of the camera 111 to the camera position and orientation of the correction image (i.e., the camera position and orientation of the target image).
[0100] FIG. 14 is a diagram showing an example of guide information. For example, as shown in FIG. 14, a target image is displayed superimposed on a current image as guide information. For example, a semi-transparent target image is superimposed on the current image. This allows the user to capture a correction image with the same camera position and orientation as the target image by capturing an image with the current image superimposed on the target image. Alternatively, only characteristic portions, such as the outline of a subject in the target image, may be extracted, and only these characteristic portions may be superimposed on the current image.
[0101] The guide information also displays information 208 instructing the user to move to the camera position of the target image (the camera position of the correction image). Note that the guide information may include instructions to adjust the camera posture to that of the target image (for example, "point the camera up"). Note that only one of superimposing the target image on the current image and displaying the information 208 may be performed.
[0102] Although the example described here is one in which the destination image is superimposed on the current image, the current image and the destination image may be displayed separately. In this case, multiple destination images may be displayed. For example, the multiple destination images may include multiple candidate images or portions thereof.
[0103] Fig. 15 is a diagram showing another example of guide information. The guide information 209 shown in Fig. 15 includes a three-dimensional model, the camera position and orientation of the current image, the camera position and orientation of the target image (the camera position and orientation of the correction image), and information 210 instructing the user to move to the camera position of the target image. Note that the guide information 209 shown in Fig. 15 may be used instead of the guide information shown in Fig. 14, or both of them may be displayed simultaneously.
[0104] If a correction image has been captured (Yes in S108), a loop closing process is performed (S109). Specifically, the camera position and orientation estimation unit 123 corrects the camera position and orientation information and map information of a series of images (e.g., images from time t = 0 to t = x shown in FIG. 5 ) using the correction image. The three-dimensional model generation unit 124 corrects the three-dimensional model by regenerating the three-dimensional model using the corrected camera position and orientation information and map information. The corrected camera position and orientation information, map information, and three-dimensional model are stored in the storage unit 122 as information generated from the series of images.
[0105] During the period until the correction image is captured (No in S108), although not shown in Figure 2, at least some of the processing in steps S101 to S104 is performed as necessary, and the guide information is updated to content based on the current image (S107).
[0106] Furthermore, in the above, the three-dimensional model generating device 102 determines whether or not a correction image has been captured, and performs loop closing processing if a correction image has been captured, but it may also determine whether or not an image has been captured that has the same (or similar) camera position and orientation as any of multiple images captured in the past, and perform loop closing processing if that image has been captured.
[0107] Alternatively, the 3D model generation device 102 may perform the loop closing process when a correction image is captured, but may not perform the loop closing process when an image with the same (or similar) camera position and orientation as an image other than the target image is captured among multiple images captured in the past. In other words, the 3D model generation device 102 may determine whether to perform the loop closing process only with respect to the target image among multiple images captured in the past. This reduces the amount of processing and reduces the occurrence of unnecessary loop closing processes.
[0108] [3. Summary] As described above, the three-dimensional model generation device (for example, the three-dimensional model generation device 102) according to this embodiment performs the processing shown in Fig. 16. Fig. 16 is a flowchart of the three-dimensional model generation processing performed by the three-dimensional model generation device according to this embodiment. Fig. 17 is a block diagram showing the hardware configuration of the three-dimensional model generation device 102. For example, the three-dimensional model generation device 102 includes a circuit 11 and a memory 12. The circuit 11 is connected to the memory 12 and performs the following processing using the memory 12.
[0109] First, the three-dimensional model generation device acquires a plurality of consecutive images (S201), and uses the plurality of images to generate a three-dimensional model including a plurality of three-dimensional points, each having position information (S202. The three-dimensional model generation device determines whether to correct accumulated errors in the three-dimensional model (e.g., loop closing processing) (S203). For example, the three-dimensional model generation device determines whether to display, on a display unit (e.g., the display unit 131), first information (e.g., guide information) for instructing a first position and orientation for capturing a first image (e.g., a correction image) for correcting accumulated errors in the three-dimensional model. If the three-dimensional model generation device determines to perform correction (Yes in S203), it causes the display unit (e.g., the display unit 131) to display the first information (e.g., guide information) for instructing a first position and orientation for capturing a first image (e.g., a correction image) for performing correction (S204). In other words, if it determines to display the first information on the display unit, the three-dimensional model generation device causes the first information to be displayed on the display unit.
[0110] This allows the 3D model generation device to support the capture of the first image for correcting the accumulated error. Therefore, by appropriately performing the correction process, the accuracy of the generated 3D model can be improved. In this way, the 3D model generation device can achieve further improvements.
[0111] For example, the first information is a second image (e.g., a target image) included in the plurality of images and captured at the first position and orientation, allowing the user to determine the first position and orientation while referring to the second image.
[0112] For example, the circuit 11 displays the second image on the display unit by superimposing it on the currently captured image (e.g., FIG. 14). This allows the user to capture the first image by capturing the currently captured image in a state where the second image matches the currently captured image.
[0113] For example, the circuit 11 further displays second information (e.g., accumulated error information) indicating the accumulated error on the display unit, allowing the user to refer to the second information and determine whether or not to correct the accumulated error.
[0114] For example, the circuit 11 may display, as the second information, a third image of a 3D model viewed from the current position and orientation on the display unit, superimposed on the currently captured image (e.g., FIG. 7). This allows the user to recognize the difference between the currently captured image and the third image. Based on this difference, the user can determine whether or not correction of accumulated error is necessary.
[0115] For example, the circuit 11 may further display, as the second information, an area where there is a difference between the third image and the currently captured image on the display unit (e.g., FIG. 8 ), allowing the user to determine the need for correction of the accumulated error based on the third information.
[0116] For example, the circuit 11 further calculates the amount of accumulated error based on the difference between the third image and the currently captured image, and displays third information indicating the calculated amount on the display unit (e.g., FIG. 9). This allows the user to recognize the magnitude (degree) of the difference between the currently captured image and the third image. Furthermore, the user can determine the need for correction of the accumulated error based on the magnitude of the difference.
[0117] For example, the circuit 11 further displays a user interface on the display unit (e.g., FIG. 11 ) together with the second information, allowing the user to instruct whether or not to perform correction, and when an instruction to perform correction is received, causes the circuit 11 to display the first information on the display unit, thereby enabling correction to be performed based on the user's decision based on the second information.
[0118] For example, the circuit 11 further displays a user interface on the display unit (e.g., FIG. 12 ) for the user to select the second image from a plurality of fourth images included in the plurality of images. This allows the user to select the second image.
[0119] For example, the circuit 11 further selects a plurality of fourth images from the plurality of images based on feature points contained in each of the plurality of images, thereby improving the accuracy of feature point matching with the second image and therefore improving the accuracy of accumulated error correction.
[0120] For example, the circuit 11 further selects a second image from the plurality of images based on feature points contained in each of the plurality of images, thereby improving the accuracy of feature point matching for the second image and therefore improving the accuracy of correction of accumulated error.
[0121] For example, when the first information is displayed on the display unit, (1) if a first image captured from a first position and orientation is acquired, correction is performed using the acquired first image, and (2) if a sixth image captured from among the multiple images at a position and orientation other than the first position and orientation of a fifth image is acquired, correction of accumulated error is not performed using the acquired sixth image. This makes it possible to prevent correction from being performed using an inappropriate image.
[0122] For example, the circuit 11 uses a plurality of images to estimate an estimated position and orientation, which is the position and orientation at which each of the plurality of images was captured, and generates a three-dimensional model using the estimated position and orientation. In correcting the accumulated error, the circuit 11 corrects the accumulated error based on the error between the estimated position and orientation of a second image captured at a first position and orientation, which are contained in the plurality of images, and the estimated position and orientation of the first image.
[0123] For example, the circuit 11 displays the current 3D model and the past 3D model on the display unit as the second information (e.g., FIG. 10 ). The circuit 11 further accepts an operation for aligning the current 3D model with the past 3D model, and performs the correction of the accumulated error using the information obtained by the operation for aligning the positions as auxiliary information. This reduces the amount of processing required for correcting the accumulated error and improves accuracy.
[0124] For example, a three-dimensional model generation device includes a memory and a circuit connected to the memory, and the circuit uses the memory to acquire a plurality of consecutive images, and uses the plurality of images to generate a three-dimensional model including a plurality of three-dimensional points, each having positional information, and selects a first image from the plurality of images, the first image being one of two images taken at the same position and orientation for correcting accumulated errors in the three-dimensional model, and in selecting the first image, the first image is selected from the plurality of images based on feature points contained in each of the plurality of images.
[0125] For example, the circuitry selects, from the plurality of images, a plurality of second images that include the first image based on the feature points.
[0126] For example, the circuit selects the first image from the plurality of images based on the number of feature points. For example, the circuit preferentially selects an image with a large number of feature points as the first image. For example, the circuit selects the first image from the plurality of images based on the degree of dispersion (dispersion) of the feature points. For example, the circuit preferentially selects an image with a high degree of dispersion of the feature points as the first image.
[0127] For example, the feature points are points detected from the first 3D model by feature analysis, or points detected from the plurality of images by feature analysis, or points in the 3D model corresponding to points detected from the plurality of images by feature analysis.
[0128] Although the 3D model generation system, 3D model generation device, etc. according to the embodiments of the present disclosure have been described above, the present disclosure is not limited to these embodiments.
[0129] Furthermore, each processing unit included in the 3D model generation device according to the above-described embodiments is typically realized as an LSI, which is an integrated circuit. These may be individually implemented as single chips, or some or all of them may be integrated into a single chip.
[0130] Furthermore, the integrated circuit is not limited to an LSI, but may be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after the LSI is manufactured, or a reconfigurable processor that can reconfigure the connections and settings of circuit cells within the LSI may also be used.
[0131] In each of the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for that component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0132] Furthermore, the present disclosure may be realized as a three-dimensional model generation method or the like executed by a three-dimensional model generation system or a three-dimensional model generation device or the like.
[0133] The division of functional blocks in the block diagram is an example, and multiple functional blocks may be realized as a single functional block, one functional block may be divided into multiple blocks, or some functions may be moved to another functional block.Furthermore, the functions of multiple functional blocks having similar functions may be processed in parallel or in time-sharing by a single piece of hardware or software.
[0134] The order in which the steps in the flowchart are executed is merely an example for specifically explaining the present disclosure, and other orders may be used. Also, some of the steps may be executed simultaneously (in parallel) with other steps.
[0135] While the 3D model generation system and 3D model generation device according to one or more aspects have been described based on the embodiments, the present disclosure is not limited to these embodiments. As long as they do not deviate from the spirit of the present disclosure, various modifications conceivable by those skilled in the art to the present embodiments and configurations constructed by combining components of different embodiments may also be included within the scope of one or more aspects.
[0136] The present disclosure can be applied to a three-dimensional model generation system, a three-dimensional model generation device, and the like.
[0137] REFERENCE SIGNS LIST 11 Circuit 12 Memory 101 Imaging device 102 Three-dimensional model generating device 103 Display device 111 Camera 112 Sensor 121 Acquisition unit 122 Storage unit 123 Camera position and orientation estimation unit 124 Three-dimensional model generating unit 125 Accumulated error detection unit 126 User guide unit 127 Output unit 128 Control unit 131 Display unit 132 Operation acceptance unit 201, 208, 210 Information 202, 203, 205, 207 User interface 206 Candidate image 209 Guide information
Claims
1. A three-dimensional model generating device comprising: a memory; and a circuit connected to the memory, wherein the circuit uses the memory to acquire a plurality of successive images; uses the plurality of images to generate a three-dimensional model including a plurality of three-dimensional points, each having positional information; determines whether or not to display first information on a display unit for indicating a first position and orientation for capturing a first image for correcting accumulated error in the three-dimensional model; and when it is determined that the first information should be displayed on the display unit, causes the first information to be displayed on the display unit.
2. The three-dimensional model generating device according to claim 1, wherein the first information is a second image included in the plurality of images and captured at the first position and orientation.
3. The three-dimensional model generating device according to claim 2, wherein said circuitry causes said second image to be displayed on said display unit in a state where it is superimposed on the image currently being captured.
4. The three-dimensional model generating device according to claim 1, wherein said circuit further causes said display unit to display second information indicating said accumulated error.
5. The three-dimensional model generating device according to claim 4, wherein the circuit causes the display unit to display, as the second information, a third image of the three-dimensional model seen from the current position and orientation, superimposed on the currently captured image.
6. The three-dimensional model generating device according to claim 5, wherein said circuit further causes said display unit to display, as said second information, an area in which there is a difference between said third image and said currently captured image.
7. A three-dimensional model generating device as described in claim 5, wherein the circuit further calculates the amount of accumulated error based on the difference between the third image and the currently captured image, and causes the display unit to display third information indicating the calculated amount.
8. The three-dimensional model generating device according to claim 4, wherein the circuit further displays, together with the second information, a user interface on the display unit for allowing the user to instruct whether or not to make the correction, and when an instruction to make the correction is received, causes the first information to be displayed on the display unit.
9. The three-dimensional model generating device according to claim 2, wherein the circuit further displays on the display unit a user interface for allowing a user to select the second image from a plurality of fourth images included in the plurality of images.
10. The three-dimensional model generating device according to claim 9, wherein said circuitry further selects said fourth images from said plurality of images based on feature points contained in each of said plurality of images.
11. The three-dimensional model generating device according to claim 2, wherein said circuitry further comprises: selecting said second image from said plurality of images based on feature points contained in each of said plurality of images.
12. A three-dimensional model generating device as described in claim 1, wherein, when the first information is displayed on the display unit, (1) if the first image taken from the first position and orientation is acquired, the correction is performed using the acquired first image, and (2) if a sixth image taken from among the multiple images at a position and orientation other than the first position and orientation of a fifth image is acquired, the correction of the accumulated error is not performed using the acquired sixth image.
13. The three-dimensional model generating device according to claim 1, wherein the circuit uses the plurality of images to estimate an estimated position and orientation, which is the position and orientation at which each of the plurality of images was captured, and generates the three-dimensional model using the estimated position and orientation; and in correcting the accumulated error, correcting the accumulated error based on the error between the estimated position and orientation of a second image captured in the first position and orientation, which are contained in the plurality of images, and the estimated position and orientation of the first image.
14. A three-dimensional model generating device as described in claim 4, wherein the circuit causes the display unit to display the current three-dimensional model and the past three-dimensional model as the second information, and the circuit further receives an operation for aligning the positions of the current three-dimensional model and the past three-dimensional model, and performs the correction by using information obtained by the operation for aligning the positions as auxiliary information.
15. A three-dimensional model generating method comprising: acquiring a plurality of consecutive images; using the plurality of images to generate a three-dimensional model including a plurality of three-dimensional points, each having positional information; determining whether or not to display first information on a display unit to indicate a first position and orientation for capturing a first image for correcting accumulated error in the three-dimensional model; and, if it is determined that the first information should be displayed on the display unit, displaying the first information on the display unit.
16. A program for causing a computer to execute the three-dimensional model generating method according to claim 15.
Citation Information
Patent Citations
Autonomous mobile device, autonomous mobile method, and program
JP2017111688A
Information processing device, method for controlling information processing device and program
JP2022022525A
Device, system, method, and program for information processing
JP2023069019A
Information processing device, control method for information processing device, information processing method, and program
JP6823403B2