Three-dimensional model generation device, three-dimensional model generation method, and program

By displaying position and orientation information and error information on the display unit, users can correct the accumulated errors of the 3D model, solving the problem of insufficient model accuracy in the prior art and achieving higher generation accuracy.

CN122122908APending Publication Date: 2026-05-29PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2024-11-13
Publication Date
2026-05-29

Smart Images

  • Figure CN122122908A_ABST
    Figure CN122122908A_ABST
Patent Text Reader

Abstract

A three-dimensional model generation apparatus (102) includes a storage (12) and a circuit (11) connected to the storage (12). The circuit (11) uses the storage (12) to acquire a plurality of continuous images (S201), generates a three-dimensional model using the plurality of images, the three-dimensional model including a plurality of three-dimensional points each having position information (S202), determines whether or not to display first information indicating a first position attitude at which a first image is captured on a display portion, the first image being used for correction of cumulative error of the three-dimensional model (S203), and displays the first information on the display portion when it is determined to display the first information on the display portion (YES in S203) (S204).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a three-dimensional model generation apparatus, a three-dimensional model generation method, and a program. Background Technology

[0002] A method is known to use SLAM (Simultaneous Localization and Mapping) to detect the position and pose of a camera from multiple images and generate a 3D model. At this time, methods are known to use loop closing (also known as loop closure) to reduce accumulated errors (e.g., see Patent Documents 1 and 2).

[0003] Existing technical documents Patent documents Patent Document 1: Japanese Patent Application Publication No. 2023-69019 Patent Document 2: Japanese Patent Application Publication No. 2017-111688 Summary of the Invention

[0004] The problem that the invention aims to solve Further improvements are desired in such 3D model generation apparatuses and methods. Therefore, this disclosure provides a 3D model generation apparatus or method capable of achieving further improvements.

[0005] Methods for solving problems One aspect of the present disclosure is a three-dimensional model generation apparatus comprising a memory and a circuit connected to the memory. The circuit uses the memory to acquire a plurality of consecutive images and uses the plurality of images to generate a three-dimensional model, the three-dimensional model comprising a plurality of three-dimensional points each having position information. The circuit determines whether to display first information for indicating a first position attitude on a display unit, the first position attitude being the position attitude at which the first image was captured, the first image being used to correct the cumulative error of the three-dimensional model. If it is determined that the first information should be displayed on the display unit, the first information is displayed on the display unit.

[0006] The effects of the invention This disclosure provides a three-dimensional model generation apparatus or method that can achieve further improvements. Attached Figure Description

[0007] Figure 1 This is a diagram showing the structure of the three-dimensional model generation system for the implementation method.

[0008] Figure 2 This is a flowchart illustrating the operation of the three-dimensional model generation device in the implementation method.

[0009] Figure 3 This is a diagram used to illustrate the estimation process of camera position and attitude in the implementation method.

[0010] Figure 4 This is a diagram used to illustrate the three-dimensional model generation process of the implementation method.

[0011] Figure 5 This is a graph used to illustrate the cumulative error of the implementation method.

[0012] Figure 6 This is a diagram used to illustrate the generation of cumulative error information in the implementation method.

[0013] Figure 7 This is a diagram showing an example of cumulative error information in an implementation method.

[0014] Figure 8 This is a diagram showing an example of cumulative error information in an implementation method.

[0015] Figure 9 This is a diagram showing an example of cumulative error information in an implementation method.

[0016] Figure 10 This is a diagram showing an example of cumulative error information in an implementation method.

[0017] Figure 11 This is a diagram showing an example of a user interface for an implementation method.

[0018] Figure 12 This is a diagram showing an example of a user interface for an implementation method.

[0019] Figure 13 This is a diagram showing an example of a user interface for an implementation method.

[0020] Figure 14 This is a diagram illustrating an example of guidance information for implementing a method.

[0021] Figure 15 This is a diagram illustrating an example of guidance information for implementing a method.

[0022] Figure 16 This is a flowchart of the three-dimensional model generation process in the implementation method.

[0023] Figure 17 This is a block diagram illustrating the hardware structure of the three-dimensional model generation device for the implementation method. Detailed Implementation

[0024] One aspect of the present disclosure is a three-dimensional model generation apparatus comprising a memory and a circuit connected to the memory. The circuit uses the memory to acquire a plurality of consecutive images and uses the plurality of images to generate a three-dimensional model, the three-dimensional model comprising a plurality of three-dimensional points each having position information. The circuit determines whether to display first information for indicating a first position attitude on a display unit, the first position attitude being the position attitude at which the first image was captured, the first image being used to correct the cumulative error of the three-dimensional model. If it is determined that the first information should be displayed on the display unit, the first information is displayed on the display unit.

[0025] Therefore, the 3D model generation apparatus can support the capture of a first image for correcting accumulated errors. Thus, by performing appropriate correction processing, the accuracy of the generated 3D model can be improved. In this way, the 3D model generation apparatus can achieve further improvements.

[0026] For example, the first information could be a second image included in the plurality of images, captured at the first position and pose.

[0027] Therefore, the user can determine the pose of the first position while referring to the second image.

[0028] For example, the circuit may also display the second image on the display unit in an overlapping manner with the currently captured image.

[0029] Therefore, the user can take a picture when the currently captured image is consistent with the second image, thus capturing the first image.

[0030] For example, the circuit may further display a second piece of information representing the accumulated error on the display unit.

[0031] Therefore, users can refer to the second piece of information to determine whether to perform cumulative error correction.

[0032] For example, as the second information, the third image obtained by the circuit from observing the three-dimensional model from the current position and pose can be superimposed on the currently captured image and displayed on the display unit.

[0033] Thus, the user can identify the difference between the currently captured image and the third image. Furthermore, based on this difference, the user can determine the necessity of correcting the accumulated error.

[0034] For example, as the second information, the circuit may further display the areas in the third image that differ from the currently captured image on the display unit.

[0035] Therefore, based on the third piece of information, the user can determine the necessity of correcting the cumulative error.

[0036] For example, the circuit may further calculate the amount of the accumulated error based on the difference between the third image and the currently captured image, and display the third information representing the calculated amount on the display unit.

[0037] Thus, users can identify the magnitude (degree) of the difference between the currently captured image and the third image. Furthermore, based on the magnitude of this difference, users can determine the necessity of correcting the accumulated error.

[0038] For example, the circuit may further display a user interface for the user to indicate whether to perform the correction on the display unit, together with the second information, and display the first information on the display unit when the instruction to perform the correction is received.

[0039] Therefore, corrections can be implemented based on the judgment of users who have referred to the second piece of information.

[0040] For example, the circuit may further display a user interface on the display unit for allowing a user to select the second image from a plurality of fourth images included in the plurality of images.

[0041] Therefore, the user can select the second image.

[0042] For example, the circuit may further select the plurality of fourth images from the plurality of images based on the feature points contained in each of the plurality of images.

[0043] This improves the accuracy of feature point matching for the second image, and thus improves the accuracy of cumulative error correction.

[0044] For example, the circuit may further select the second image from the plurality of images based on the feature points contained in each of the plurality of images.

[0045] This improves the accuracy of feature point matching for the second image, and thus improves the accuracy of cumulative error correction.

[0046] For example, when the first information is displayed on the display unit, (1) if the first image captured from the first position and posture is obtained, the correction is performed using the obtained first image; (2) if the sixth image obtained from the position and posture of the fifth image is obtained, the correction for the cumulative error is not performed using the obtained sixth image, wherein the fifth image was captured from a position and posture other than the first position and posture.

[0047] Therefore, it is possible to suppress the use of inappropriate images for correction.

[0048] For example, the circuit may use the plurality of images to estimate the estimated position and pose, and use the estimated position and pose to generate the three-dimensional model. The estimated position and pose is the position and pose of each of the plurality of images. In the correction of the accumulated error, the accumulated error is corrected based on the error between the estimated position and pose of the second image, which is included in the plurality of images and is taken at the first position and pose, and the estimated position and pose of the first image.

[0049] For example, as the second information, the circuit may display the current 3D model and the past 3D model on the display unit, and the circuit may further perform an operation to align the position of the current 3D model with that of the past 3D model. In the correction, the information obtained by the operation to align the position is used as auxiliary information to perform the correction.

[0050] This reduces the amount of processing required to correct accumulated errors and improves accuracy.

[0051] One aspect of the disclosed method for generating a three-dimensional model involves acquiring multiple consecutive images, using the multiple images to generate a three-dimensional model, the three-dimensional model including multiple three-dimensional points each having position information, determining whether to display first information used to indicate a first position attitude on a display unit, the first position attitude being the position attitude at which the first image was captured, the first image being used to correct the cumulative error of the three-dimensional model, and if it is determined that the first information should be displayed on the display unit, then the first information is displayed on the display unit.

[0052] Therefore, the 3D model generation method can support the capture of images used to correct accumulated errors. Thus, by performing appropriate correction processing, the accuracy of the generated 3D model can be improved. This allows for further improvements to the 3D model generation method.

[0053] Additionally, one aspect of the program disclosed herein is a program for causing a computer to execute the three-dimensional model generation method.

[0054] Furthermore, these general or specific forms can be realized either through systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or through any combination of systems, methods, integrated circuits, computer programs, and recording media.

[0055] Hereinafter, embodiments will be described in detail with appropriate reference to the accompanying drawings. However, sometimes unnecessary detailed descriptions are omitted. For example, detailed descriptions of matters already known or repeated descriptions of substantially the same structures are sometimes omitted. This is to avoid making the following description unnecessarily lengthy and to facilitate understanding by those skilled in the art.

[0056] Furthermore, the inventors have provided the drawings and the following description in order to enable those skilled in the art to fully understand this disclosure, but they are not intended to limit the subject matter of the claims.

[0057] (Implementation Method) [1. Structure] First, the structure of the three-dimensional model generation system of this embodiment will be explained. Figure 1 This is a block diagram illustrating the structure of the three-dimensional model generation system according to this embodiment. The three-dimensional model generation system is a system that generates three-dimensional models based on multiple images, including a camera device 101, a three-dimensional model generation device 102, and a display device 103.

[0058] The imaging device 101 generates multiple images and multiple sensor information, and outputs the generated multiple images and multiple sensor information to the 3D model generation device 102. The imaging device 101 includes a camera 111 and a sensor 112. The camera 111 captures images of the subject (object) from different viewpoints (shooting position and shooting direction), and outputs the resulting multiple images (frames) to the 3D model generation device 102. For example, the imaging device 101 moves while shooting, thereby capturing multiple images.

[0059] Sensor 112 includes, for example, at least one of a distance sensor (depth sensor), a laser sensor (e.g., LiDAR), and an IMU (Inertial Measurement Unit). The IMU may include, for example, at least one of an accelerometer, a rotational angular acceleration sensor, and a gyroscope sensor. Sensor 112 outputs sensor information corresponding to multiple images to the 3D model generation apparatus 102. Here, the sensor information includes, for example, at least a portion of the depth image obtained from the distance sensor, point group data obtained from the laser sensor, and acceleration information obtained from the IMU. Furthermore, the camera apparatus 101 may not include sensor 112.

[0060] The 3D model generation apparatus 102 uses multiple images and multiple sensor information to generate a 3D model. The 3D model generation apparatus 102 includes an acquisition unit 121, a storage unit 122, a camera position and attitude estimation unit 123, a 3D model generation unit 124, a cumulative error detection unit 125, a user guidance unit 126, an output unit 127, and a control unit 128.

[0061] The acquisition unit 121 acquires multiple images and multiple sensor information from the imaging device 101. The storage unit 122 stores the multiple images and multiple sensor information acquired by the acquisition unit 121. In addition, the storage unit 122 stores camera position and pose information, mapping information, and three-dimensional models, etc., which will be described later.

[0062] The camera position and attitude estimation unit 123 uses multiple images to generate camera position and attitude information representing the position and attitude (orientation) of the camera device 101 when each image is captured. Here, the attitude of the camera device 101 represents at least one of the shooting direction of the camera device 101 and the tilt of the camera device 101. The shooting direction of the camera device 101 is the direction of the optical axis of the camera device 101. The tilt of the camera device 101 is the rotation angle of the camera device 101 about the optical axis relative to the reference attitude.

[0063] In addition, the camera position and attitude estimation unit 123 generates mapping information, which together with the camera position and attitude information represents a three-dimensional position and a three-dimensional point group.

[0064] The 3D model generation unit 124 generates a 3D model of the subject based on multiple images, camera position and pose information, mapping information, and sensor information.

[0065] The cumulative error detection unit 125 detects the cumulative error (also called cumulative error) of the 3D model and generates cumulative error information representing the cumulative error. The user guidance unit 126 generates guidance information that supports the user in moving to the shooting position of the correction image used for closed-loop processing.

[0066] The output unit 127 outputs the current image, accumulated error information, and guidance information to the display device 103. The control unit 128 controls each processing unit. For example, the control unit 128 controls each processing unit based on user operations.

[0067] The display device 103 displays the current image, accumulated error information, and guidance information sent from the 3D model generation device 102. The display device 103 includes a display unit 131 and an operation receiving unit 132. The display unit 131 displays the current image, accumulated error information, and guidance information. The operation receiving unit 132 receives user operations. For example, the operation receiving unit 132 may be a touch panel, keyboard, or mouse. Alternatively, the operation receiving unit 132 may be a microphone for receiving voice operations. Furthermore, as a mechanism for providing information to the user, the display device 103 may also include a speaker for voice output and a vibrator for transmitting vibrations to the user.

[0068] also, Figure 1The camera device 101, the 3D model generation device 102, and the display device 103 shown can also be implemented as a single device. Alternatively, the multiple processing units contained in each device can be distributed across multiple devices.

[0069] For example, in a use case where a user-held terminal (such as a smartphone, tablet, or personal computer) is used to move the subject while taking pictures, the camera device 101, the 3D model generation device 102, and the display device 103 may also be included in that terminal. Alternatively, in other use cases, the camera device 101 may be included in a user-operated mobile object (e.g., a vehicle, robot, or drone), and the 3D model generation device 102 and the display device 103 may be included in the user-held terminal. Furthermore, in any case, a portion of the processing unit included in each device may be included in another device, such as a server, connected to the user-operated terminal via a communication network. For example, part or all of the processing unit included in the 3D model generation device 102 may also be included in the server.

[0070] in addition, Figure 1 The data transmission and reception between the devices shown can be carried out through any method, such as wired or wireless communication. Furthermore, this communication can occur directly between the devices or indirectly via other communication devices or servers. Additionally, this transmission and reception can also be the exchange of data within a single device.

[0071] [2. Action] Next, the operation of the 3D model generation device 102 will be explained. For example, the 3D model generation system generates a 3D model in real time based on images captured by the camera device 101.

[0072] Figure 2 This is a flowchart illustrating the operation of the 3D model generation device 102. Figure 2 The processing shown is repeated, for example, in units of one frame. Furthermore, Figure 2 The processing shown can also be repeated in multiple frame units.

[0073] Furthermore, the multiple images can be multiple still images captured by the camera device 101 while it is moving, or multiple images contained in a moving image captured by the camera device 101 while it is moving. Additionally, when using a moving image, the multiple images do not need to be all the images contained in the moving image; they can be a portion of the multiple images contained in the moving image, i.e., the main image (keyframe). Here, the main image is, for example, an image extracted from the multiple images contained in the moving image at predetermined time intervals.

[0074] [2-1. Estimation of Camera Position and Attitude] First, the acquisition unit 121 acquires the current image (frame) and multiple sensor information from the camera device 101 (S101). Next, the camera position and pose estimation unit 123 uses the multiple images, including the current image, to generate camera position and pose information and mapping information (S102).

[0075] Furthermore, the method used by the camera position and attitude estimation unit 123 to estimate the position and attitude of the camera device 101 is not particularly limited. For example, the camera position and attitude estimation unit 123 may use, for example, Visual-SLAM (Simultaneous Localization and Mapping), Structure-From-Motion, or ICP (Iterative Closest Point) to estimate the position and attitude of multiple camera devices 101.

[0076] Figure 3 This diagram illustrates the camera position and pose estimation process. The diagram shows an example of using three frames, F0, F1, and F2, for camera position and pose estimation.

[0077] like Figure 3 As shown, the camera position and attitude estimation unit 123 performs feature point matching processing on multiple frames F0 to F2 captured by the camera device 101. That is, the camera position and attitude estimation unit 123 extracts feature points for each of the multiple frames F0 to F2, and extracts groups of similar points (corresponding points) that are similar across the multiple frames from the extracted feature points. Then, the camera position and attitude estimation unit 123 uses the extracted groups of similar points to estimate the position and attitude of the camera device 101 in each frame.

[0078] Furthermore, the mapping information is represented as multiple mapping points, each representing a 3D point. The camera position and pose estimation unit 123 generates the mapping information by performing triangulation using the result of feature point matching processing and the camera position and pose information. In addition to position information, the mapping information may also include information such as the color of each mapping point and the surface shape surrounding the mapping point (e.g., information representing normals). The mapping information can be further refined into high-precision information through optimization processing together with the camera position and pose.

[0079] [2-2. Generation of 3D Models] Next, the 3D model generation unit 124 generates a 3D model of the subject based on multiple images, including the current image, and camera position and pose information (S103). For example, the 3D model generation unit 124 uses mapping information to generate the 3D model. Figure 4 This is a diagram used to illustrate the 3D model generation process in this case. For example, as shown... Figure 4As shown, the three-dimensional model generation unit 124 generates a mesh model represented by a plane with mapping points as vertices, as a three-dimensional model.

[0080] Furthermore, the 3D model is not limited to a mesh model and can be any 3D model. For example, mapping information can be directly used as a 3D model. Additionally, if the camera device 101 is equipped with a sensor 112, the 3D model can be generated using sensor information and camera position and pose information obtained from the sensor 112. For example, a 3D model can be generated using point cluster data obtained from LiDAR or a depth image obtained from a distance sensor. Alternatively, multiple of these mapping information, point cluster data, and depth information can be used. For example, depth information can also be estimated from the image using AI (artificial intelligence).

[0081] In addition to location information, 3D models can also contain attribute information (information representing color or normals).

[0082] [2-3. Detection of Cumulative Error] Next, the cumulative error detection unit 125 detects the cumulative error of the current frame and displays the cumulative error information, representing the cumulative error, on the display unit 131 (S104). Here, the cumulative error is the difference (error) between the 3D model generated using past frames and the 3D model generated using the current frame. In other words, the cumulative error is the difference (error) between the estimated camera position and pose of the current frame and the actual camera position and pose.

[0083] Figure 5 This is a graph used to illustrate the cumulative error. Figure 5 An example is shown where the camera device 101 moves while capturing images of multiple objects. In this example, the camera device 101 starts capturing images and moves from time t0, returning to the same position as at time t=x. Here, in the estimation of camera position and pose and the generation of the 3D model, the camera position and pose of newly captured images are estimated using previously estimated images, and the 3D model is updated (added). Therefore, as the number of images used increases (the shooting time becomes longer), the positional error between the estimated camera position and pose and the 3D model and the actual camera position and pose and the position of the objects may increase. For example, as... Figure 5 As shown, the further forward the time, the greater the difference (cumulative error) between the actual camera position and attitude (actual trajectory) and the estimated camera position and attitude (estimated trajectory).

[0084] Loop closure processing is the process of correcting this accumulated error. By performing loop closure processing, the estimated camera position and pose (estimated trajectory) is corrected to be close to the actual camera position and pose (actual trajectory). Furthermore, the position information of the 3D model is updated based on the corrected camera position and pose. For example, using the estimated camera position and pose between a first image taken in the past (image at t=0) and a second image taken at the same position as the first image (image at t=x), the camera position and pose of all images within the range of t=0 to t=x are corrected in such a way that the camera position and pose of the second image becomes the camera position and pose of the first image. That is, in order to perform loop closure processing, images taken at the same camera position and pose as any one of the multiple images taken in the past are required.

[0085] The following is an example illustrating cumulative error information. Figure 6 This is a graph used to illustrate the generation of accumulated error information. For example, such as... Figure 6 As shown, cumulative error information is generated by projecting a past 3D model onto the current image. Here, the past 3D model is, for example, a 3D model generated using multiple images from t=0 to x-1.

[0086] Figure 7 This is a graph showing an example of accumulated error information. For example, as shown... Figure 7 As shown, the image obtained by projecting a past 3D model onto the current image is displayed as cumulative error information. Therefore, the user can determine the necessity of loop closure processing based on the presence and magnitude of the difference between the current image and the image projected with the past 3D model. Furthermore, the projected 3D model can be displayed semi-transparently, for example. Alternatively, features such as the outline of the subject can be extracted, and only those features can be displayed.

[0087] Furthermore, the cumulative error detection unit 125 can calculate the cumulative error. For example, the cumulative error detection unit 125 uses a first image, based on the current image or the current 3D model and viewed from the current camera position and pose, and a second image, viewed from the current camera position and pose, of a past 3D model, to calculate the difference in pixel values ​​between the first image and the second image. Here, pixel values ​​are, for example, depth values, color values, or normal vectors. The depth value of the first image can be obtained from the current 3D model, for example, and if the sensor information includes a depth image, the current depth image can also be used as the first image. Alternatively, a depth image generated using multiple images including the current image can be used as the first image. Furthermore, the current 3D model refers to a 3D model generated based on multiple images including the current image (e.g., images from t=x-2 to t=x). Additionally, the depth value of the second image is generated based on a past 3D model.

[0088] In addition, when comparing color values, the first image is the current image, and the second image is generated, for example, based on a past 3D model.

[0089] Cumulative error information can include information representing the difference calculated in this way. Figure 8 This is a diagram showing an example of accumulated error information in this situation. For example, as shown... Figure 8 As shown, areas with differences can also be highlighted. Additionally, in Figure 8 In the example shown, the magnitude of the difference is represented by shades of gray. Furthermore, the method of emphasis is not limited to this; any method such as blinking can be used. Additionally, emphasis is not always necessary; any method that allows the user to identify the area where the difference exists is sufficient. Furthermore, the number of grayscale values ​​representing the magnitude of the difference can be arbitrary. Alternatively, grayscale can be omitted, and the presence or absence of the difference (whether the difference is above a threshold) can be indicated instead.

[0090] In addition, Figure 8 The text indicates that the region containing the difference can be displayed, but it can also emphasize the objects containing the difference. Furthermore, the difference can be calculated not only in pixels but also in regions containing multiple pixels. In this case, the sum, average, median, or maximum value of the differences (or absolute values ​​of the differences) among the pixels within the region can be calculated as the difference for that region.

[0091] Furthermore, the difference across the entire image can be calculated as the cumulative error. For example, the difference across the entire image can be calculated as the sum, average, median, or maximum of the differences among multiple pixels. Alternatively, a weighted sum or weighted average can be used. In this case, for example, pixels closer to the center of the image can be assigned a higher weight.

[0092] Additionally, weights can be set based on the accuracy of the 3D model. For example, regions with higher accuracy can be assigned greater weights. This accuracy can be determined based on the camera position and distance from the subject in the image used to generate the 3D model. For instance, closer distances generally indicate higher accuracy. Alternatively, accuracy can be determined based on the distance to detected feature points, vertices, or edges. Again, closer distances generally indicate higher accuracy. Furthermore, accuracy can be determined based on the presence of texture. For example, the presence of texture indicates higher accuracy. Additionally, when using depth images or point group data obtained from laser sensors in the generation of the 3D model, the accuracy of the 3D model can be determined based on information indicating the accuracy of the sensor or the accuracy of the sensor results contained in the supplementary information of the sensor data. Furthermore, multiple factors can be combined to determine accuracy.

[0093] Furthermore, to facilitate the aforementioned determination, information indicating the precision of each point in the 3D model can be added during its generation. Additionally, information identifying the frame used to generate that point can also be added to each point in the 3D model.

[0094] The cumulative error information can include information representing the cumulative error of the entire image calculated in this way. Figure 9 This is a diagram showing an example of accumulated error information in this situation. For example, as shown... Figure 9 As shown, information 201, representing the cumulative error of the entire image, can also be displayed to indicate the necessity of loop closure processing. Alternatively, information 201 can show the magnitude of the cumulative error instead of the necessity of loop closure processing. Furthermore, while information 201 indicates the degree of necessity for loop closure processing (the magnitude of the cumulative error), it can also indicate whether loop closure processing is required (whether there is a cumulative error (or whether it exceeds a threshold)).

[0095] In addition to considering accumulated errors, the necessity of closed-loop processing can also be determined based on the shooting time (the number of images used in the generation of the 3D model) without closed-loop processing. For example, it is determined that the longer the shooting time, the greater the necessity of closed-loop processing.

[0096] Furthermore, the methods for notifying (prompting) users whether closed-loop processing is needed (or the degree of necessity) are not limited to this; any display method can be used, such as changing the color of the screen frame or making it blink. Additionally, other notification methods such as voice or vibration can be used, beyond just display. Furthermore, multiple methods can be combined.

[0097] In addition, the above text describes an example of projecting a past 3D model onto the current image as accumulated error information, but it is also possible to display both the current 3D model and the past 3D model as accumulated error information. Figure 10 This is a diagram showing an example of accumulated error information in this situation. For example, the current 3D model and past 3D models can also be displayed in different colors. Additionally, in... Figure 10 In this mode, only the past 3D model corresponding to the current 3D model is displayed. However, it is also possible to display the entire 3D model, including subjects not currently included in the image, as the past 3D model. In this case, the color of the past 3D model can be changed based on its distance from the current 3D model.

[0098] Furthermore, the viewpoint of the displayed 3D model can be variable. For example, as Figure 10As shown, a user interface 202 can display the viewpoint of the current 3D model and past 3D models for changing (translation, rotation, and scaling). Furthermore, the user interface 202 does not necessarily need to be displayed; the viewpoint can also be changed through predetermined operations on the touch panel (drag, swipe, pinch, and expand, etc.).

[0099] in addition, Figure 10 The cumulative error information shown can replace Figure 7 The cumulative error information, such as those shown, can be displayed, or it can be displayed simultaneously through screen splitting, or any one of them can be selectively displayed based on user actions. Additionally, it can be combined with... Figure 8 The example shown also illustrates the difference, which can also be compared with... Figure 9 Similarly, it displays information indicating the necessity of closed-loop processing (the magnitude of accumulated error), etc.

[0100] In addition, Figure 10 In this process, the current 3D model can also be manually aligned with past 3D models. That is, the user can input an operation to align the current 3D model with past 3D models. Furthermore, information based on this manually input operation can be used as auxiliary information for closed-loop processing. For example, this operation can obtain the direction and amount of the offset between the current 3D model and the past 3D model. By using at least one of the offset direction and amount in the initial value setting of the matching process search in closed-loop processing, the processing workload can be reduced and accuracy improved.

[0101] Furthermore, as described above, the method for calculating the cumulative error involves comparing the pixel values ​​of each pixel in the first image and the second image, which are obtained from the current camera position and pose based on the current image or the current 3D model, and a second image of the past 3D model observed from the current camera position and pose. However, other methods can also be used. For example, the cumulative error detection unit 125 can detect corresponding points that are feature points corresponding to each other in the current 3D model and the past 3D model, and calculate the cumulative error based on the distance between the corresponding points. In this case, the cumulative error can be calculated in point units, region units containing multiple points, or image units. For example, as the cumulative error in region or image units, the sum, average, median, maximum, weighted sum, or weighted average of multiple differences among multiple corresponding points contained in the region or image can be used. Furthermore, the weights can be set in the same way as described above.

[0102] Furthermore, instead of detecting corresponding points between 3D models, corresponding points can be detected between the current image and an image obtained by observing past 3D models from the current camera position and pose.

[0103] [2-4. Implementation Determination of Closed-Loop Processing] like Figure 2 As shown, the cumulative error detection unit 125 then determines whether to perform closed-loop processing (S105). However, this determination is not about whether to perform closed-loop processing immediately, but rather whether to display guidance information used to capture a correction image for the closed-loop processing described later.

[0104] For example, the cumulative error detection unit 125 displays a user interface along with the cumulative error information for the user to instruct on the implementation of closed-loop processing. Figure 11 This is a diagram showing an example of the user interface 203. If the user instructs the user interface 203 to implement closed-loop processing, the cumulative error detection unit 125 determines that closed-loop processing is to be implemented ("Yes" in S105); otherwise, it determines that closed-loop processing is not to be implemented ("No" in S105).

[0105] Furthermore, the method for implementing user-instructed closed-loop processing is not limited to the methods described above, and can be any method.

[0106] Alternatively, the user interface 203 may not always be displayed, but may be displayed only when the calculated cumulative error (necessity of closed-loop processing) is above a predetermined threshold, and not displayed otherwise.

[0107] If it is determined that closed-loop processing will not be performed ("No" in S105), the next image will be processed according to the steps after S101.

[0108] [2-5. Generation of Guiding Information] On the other hand, if it is determined that closed-loop processing is to be implemented ("Yes" in S105), the user guidance unit 126 selects a target image (S106). Here, the target image is an image included in past images, and is an image used to determine the camera position and orientation of the correction image to be taken subsequently. For example, the camera position and orientation of the target image is determined to be the camera position and orientation of the correction image. Furthermore, the camera position and orientation of the correction image does not need to be exactly the same as the camera position and orientation of the target image. For example, the difference between the camera position and orientation of the correction image and the camera position and orientation of the target image only needs to be within a specified range.

[0109] First, the user guidance unit 126 selects multiple candidate images from a plurality of past images. For example, the user guidance unit 126 selects multiple candidate images from a plurality of past 3D models (i.e., 3D models corresponding to the subject in the current image) used to generate a cumulative error-projected image onto the current image. Alternatively, the multiple candidate images may be selected from a plurality of images associated with the plurality of past images. Here, the associated images are, for example, images with camera positions and poses close to those of the plurality of past images.

[0110] Specifically, the user guidance unit 126 preferentially selects images whose camera position and pose are close to the estimated current camera position and pose. Alternatively, the user guidance unit 126 preferentially selects images with more feature points. Furthermore, in addition to the number of feature points, the user guidance unit 126 can also determine the degree of feature point dispersion (dispersion) and preferentially select images with a high degree of dispersion. Alternatively, the user guidance unit 126 can also preferentially select images in which the projected past 3D model is centered. Furthermore, the user guidance unit 126 can use any one of these multiple determination criteria, or it can combine two or more determination criteria.

[0111] Furthermore, the aforementioned feature points are, for example, points detected based on a 3D model through feature quantity analysis. Alternatively, the aforementioned feature points are points detected based on a 2D image through feature quantity analysis. Or, the aforementioned feature points are points in a 3D model corresponding to points detected based on a 2D image through feature quantity analysis.

[0112] Next, the user guidance unit 126 displays a user interface for the user to select a target image from a plurality of candidate images. Figure 12 This is a diagram illustrating an example display of the user interface 205. The user determines the target image by selecting one from a plurality of candidate images 206 displayed. However, the method by which the user selects the target image is not limited to this.

[0113] Figure 13 This is a diagram showing another example of a user interface for allowing a user to select a target image from multiple candidate images. Figure 13 The user interface 207 shown displays past 3D models, the camera position and pose of the current image, and the camera position and pose of multiple candidate images. In this case, the user selects a candidate image by choosing any one of the camera position and pose of the multiple candidate images. Alternatively, instead... Figure 12 The user interface 205 shown can be used Figure 13 The user interface 207 shown can also display both of these simultaneously.

[0114] Alternatively, instead of the user selecting the target image, the user guidance unit 126 may select the target image from multiple images. In this case, for example, the user guidance unit 126 may select the image with the highest priority as the target image in the candidate image selection method described above.

[0115] Next, the user guidance unit 126 generates guidance information for the user to capture the target image and displays the generated guidance information (S107). Here, the guidance information is information used to assist the user in capturing the correction image, and is information used to guide the camera position and attitude of the camera 111 to the camera position and attitude of the correction image (i.e., the camera position and attitude of the target image).

[0116] Figure 14 This is a diagram representing an example of guiding information. For example, such as... Figure 14 As shown, the target image is overlaid on the current image as guidance information. For example, a semi-transparent target image is overlaid on the current image. Thus, by taking a picture with the current image and the target image overlapping, the user can capture a correction image with the same camera position and pose as the target image. Alternatively, only the contours or other features of the subject within the target image can be extracted and overlaid only on the current image.

[0117] Additionally, the guidance information display 208 provides information 208 indicating to the user to move the camera to the target image (the camera position of the correction image). Furthermore, the guidance information may also include instructions to align the camera orientation with the target image's camera orientation (e.g., "Please turn the camera upwards," etc.). Alternatively, it may be possible to simply overlay either the target image or the display information 208 onto the current image.

[0118] Furthermore, while an example of overlaying the target image onto the current image has been described, the current image and the target image can also be displayed separately. Additionally, in this case, multiple target images can also be displayed. For example, multiple target images could contain multiple candidate images or portions thereof.

[0119] Figure 15 This is another example of a diagram representing guidance information. Figure 15 The guidance information 209 shown includes a 3D model, the camera position and pose of the current image, the camera position and pose of the target image (the camera position and pose of the image for correction), and information 210 regarding the user's instruction to move the camera towards the target image. Additionally, it can be used... Figure 15 The guidance information 209 shown is used instead Figure 14 The guidance information shown can also display both of these simultaneously.

[0120] If a correction image has been captured ("Yes" in S108), loop closure processing is performed (S109). Specifically, the camera position and pose estimation unit 123 uses the correction image to correct a series of images (e.g., Figure 5 The image (shown at times t=0 to t=x) contains camera position and orientation information and mapping information. The 3D model generation unit 124 corrects the 3D model by regenerating the 3D model using the corrected camera position and orientation information and mapping information. Furthermore, the corrected camera position and orientation information, mapping information, and 3D model are stored in the storage unit 122 as information generated from a series of images.

[0121] Furthermore, during the period up to the time it takes to capture the correction image ("No" in S108), although in Figure 2 The illustrations are omitted, but at least a portion of steps S101 to S104 are processed as needed, and the guidance information is updated to the content based on the current image (S107).

[0122] In addition, as mentioned above, the 3D model generation device 102 determines whether a correction image has been captured and performs closed-loop processing if a correction image has been captured. However, it can also determine whether an image has been captured that is the same as (or similar to) any of the multiple images captured in the past, and perform closed-loop processing if such an image has been captured.

[0123] Alternatively, the 3D model generation apparatus 102 may perform loop closure processing only when a correction image is captured, and may not perform loop closure processing when an image with the same (or similar) camera position and pose as the target image is captured from among multiple previously captured images. That is, the 3D model generation apparatus 102 may determine whether to perform loop closure processing only for the target image from among multiple previously captured images. This reduces the processing load and minimizes the generation of unnecessary loop closure processing.

[0124] [3. Conclusion] As described above, the three-dimensional model generation apparatus (e.g., three-dimensional model generation apparatus 102) of this embodiment performs... Figure 16 The processing is shown. Figure 16 This is a flowchart of the three-dimensional model generation process of the three-dimensional model generation device in this embodiment. Figure 17 This is a block diagram showing the hardware structure of the 3D model generation apparatus 102. For example, the 3D model generation apparatus 102 includes a circuit 11 and a memory 12. The circuit 11 is connected to the memory 12, and the memory 12 is used to perform the following processing.

[0125] First, the 3D model generation apparatus acquires multiple consecutive images (S201), and uses these images to generate a 3D model, which includes multiple 3D points, each with positional information (S202). The 3D model generation apparatus determines whether to perform cumulative error correction (e.g., closed-loop processing) on ​​the 3D model (S203). For example, the 3D model generation apparatus determines whether to display first information (e.g., guidance information) used to indicate a first positional attitude on a display unit (e.g., display unit 131), where the first positional attitude is the positional attitude of the first image (e.g., a correction image) used to correct the cumulative error of the 3D model. If the 3D model generation apparatus determines that correction is to be performed ("Yes" in S203), it displays the first information (e.g., guidance information) used to indicate the first positional attitude on the display unit (e.g., display unit 131) (S204), where the first positional attitude is the positional attitude of the first image (e.g., a correction image) used for correction. That is, if the 3D model generation apparatus determines that the first information should be displayed on the display unit, it displays the first information on the display unit.

[0126] Therefore, the 3D model generation apparatus can support the capture of a first image for correcting accumulated errors. Thus, by performing appropriate correction processing, the accuracy of the generated 3D model can be improved. In this way, the 3D model generation apparatus can achieve further improvements.

[0127] For example, the first information is a second image (e.g., the target image) captured in a first position and pose, which is included in a plurality of images. Thus, the user can determine the first position and pose while referring to the second image.

[0128] For example, circuit 11 displays the second image superimposed on the currently captured image on the display unit (e.g., Figure 14 Thus, the user can capture the first image by taking a picture while the currently captured image is in a state consistent with the second image.

[0129] For example, circuit 11 further displays second information indicating accumulated error (e.g., accumulated error information) on the display unit. Thus, the user can refer to the second information to determine whether to perform accumulated error correction.

[0130] For example, as the second piece of information, the circuit 11 displays the third image obtained from observing the three-dimensional model from the current position and orientation on the display unit, overlapping the currently captured image (e.g., Figure 7 Thus, the user can identify the difference between the currently captured image and the third image. Furthermore, based on this difference, the user can determine the necessity of correcting the accumulated error.

[0131] For example, as the second piece of information, circuit 11 further displays the areas in the third image that differ from the currently captured image on the display unit (e.g., Figure 8 Therefore, based on the third piece of information, the user can determine the necessity of correcting the cumulative error.

[0132] For example, circuit 11 further calculates the amount of accumulated error based on the difference between the third image and the currently captured image, and displays the third information representing the calculated amount on the display unit (e.g., Figure 9 Thus, the user can identify the magnitude (degree) of the difference between the currently captured image and the third image. Furthermore, based on the magnitude of this difference, the user can determine the necessity of correcting the accumulated error.

[0133] For example, circuit 11 further displays the user interface for the user to indicate whether to perform correction along with the second information on the display unit (e.g., Figure 11 Upon receiving an instruction to perform correction, the first information is displayed on the display unit. This allows correction to be performed based on the user's judgment, which is then referenced to the second information.

[0134] For example, circuit 11 further displays a user interface on the display unit for allowing a user to select a second image from a plurality of fourth images included in a plurality of images (e.g., Figure 12 Therefore, the user can select the second image.

[0135] For example, circuit 11 further selects a plurality of fourth images from the plurality of images based on the feature points contained in each of the plurality of images. This improves the accuracy of feature point matching for the second image, and thus improves the accuracy of correction of accumulated errors.

[0136] For example, circuit 11 further selects a second image from the multiple images based on the feature points contained in each of the multiple images. This improves the accuracy of feature point matching for the second image, and thus improves the accuracy of correction of accumulated errors.

[0137] For example, when displaying the first information on the display unit, (1) if a first image captured from a first position and attitude is obtained, the obtained first image is used for correction; (2) if a sixth image captured from a fifth image among multiple images is obtained, the obtained sixth image is not used for cumulative error correction, wherein the fifth image was captured from a position and attitude other than the first position and attitude. Thus, correction using an inappropriate image can be suppressed.

[0138] For example, circuit 11 uses multiple images to estimate the estimated position and pose, and uses the estimated position and pose to generate a three-dimensional model. The estimated position and pose is the position and pose of each image in the multiple images. In the correction of cumulative error, the cumulative error is corrected based on the error between the estimated position and pose of the second image taken in the first position and pose and the estimated position and pose of the first image.

[0139] For example, as the second piece of information, circuit 11 displays the current 3D model and the past 3D model on the display unit (e.g., Figure 10 Circuit 11 further handles the operation of aligning the current 3D model with the position of the past 3D model. In the correction of accumulated errors, the information obtained through the operation of aligning the position is used as auxiliary information for correction. As a result, the amount of processing for correcting accumulated errors can be reduced, and the accuracy can be improved.

[0140] For example, a 3D model generation apparatus includes a memory and circuitry connected to the memory. The circuitry uses the memory to acquire a plurality of consecutive images and uses the plurality of images to generate a 3D model, the 3D model including a plurality of 3D points each having positional information. A first image is selected from the plurality of images. The first image is one of two images taken at the same position and pose for correcting the cumulative error of the 3D model. In the selection of the first image, the first image is selected from the plurality of images based on the feature points contained in each of the plurality of images.

[0141] For example, the circuit selects a plurality of second images containing the first image from the plurality of images based on the feature points.

[0142] For example, the circuit selects the first image from the plurality of images based on the number of feature points. For example, the circuit preferentially selects the image with a large number of feature points as the first image. For example, the circuit selects the first image from the plurality of images based on the degree of dispersion (dispersion) of the feature points. For example, the circuit preferentially selects the image with a high degree of dispersion of feature points as the first image.

[0143] For example, the feature point is a point detected by feature quantity analysis based on the first 3D model. Alternatively, the feature point is a point detected by feature quantity analysis based on the plurality of images. Or, the aforementioned feature point is a point in the 3D model corresponding to a point detected by feature quantity analysis based on the plurality of images.

[0144] The above describes the three-dimensional model generation system and three-dimensional model generation apparatus according to the embodiments of the present disclosure, but the present disclosure is not limited to this embodiment.

[0145] Furthermore, the processing units included in the 3D model generation apparatus and the like in the above embodiments are typically implemented as LSIs (Liquid Crystal Sensors) of integrated circuits. They can be implemented as individual chips or as chips that include some or all of them.

[0146] Furthermore, integrated circuitry is not limited to LSIs; it can also be achieved through dedicated circuits or general-purpose processors. Alternatively, FPGAs (Field Programmable Gate Arrays) that can be programmed after LSI manufacturing, or reconfigurable processors that can reconfigure the connections and settings of the internal circuitry units of the LSI, can be used.

[0147] Furthermore, in the above embodiments, each component may be constructed using dedicated hardware, or implemented by executing software programs suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing software programs recorded on a recording medium such as a hard disk or semiconductor memory.

[0148] In addition, this disclosure can also be implemented as a three-dimensional model generation method executed by a three-dimensional model generation system or a three-dimensional model generation device.

[0149] Furthermore, the segmentation of functional blocks in the block diagram is one example. Multiple functional blocks can also be implemented as a single functional block, or a single functional block can be divided into multiple functional blocks, or some functionality can be transferred to other functional blocks. Additionally, the functionality of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.

[0150] Furthermore, the order of the steps in the execution flowchart is illustrative for the purpose of explaining this disclosure, and may be in a different order than described above. Additionally, some of the steps described above may be executed simultaneously (in parallel) with other steps.

[0151] The above description illustrates one or more methods of three-dimensional model generation systems and apparatuses based on embodiments, but this disclosure is not limited to these embodiments. Various modifications conceived by those skilled in the art to these embodiments, and combinations of constituent elements from different embodiments, can also be included within the scope of one or more embodiments, provided they do not depart from the spirit of this disclosure.

[0152] Industrial availability This disclosure can be applied to 3D model generation systems and 3D model generation devices, etc.

[0153] Explanation of reference numerals in the attached figures 11 Circuits 12 Memory 101 Camera Device 102 Three-dimensional model generation device 103 Display Device 111 Camera 112 Sensors 121 Acquisition Department 122 Storage Department 123 Camera position and attitude estimation unit 124 3D Model Generation Department 125 Cumulative Error Detection Department 126 User Guidance Department 127 Output Section 128 Control Department 131 Display Department 132 Operations and Acceptance Department Information 201, 208, 210 User interface for numbers 202, 203, 205, and 207 206 candidate images 209 Guiding Information

Claims

1. A three-dimensional model generation device, wherein, have: Memory; and The circuit is connected to the memory. The circuit uses a memory. Obtain multiple consecutive images. The multiple images are used to generate a three-dimensional model, which includes multiple three-dimensional points, each with its own location information. A determination is made as to whether to display the first information used to indicate the first position and attitude on the display unit. The first position and attitude is the position and attitude at which the first image was captured. The first image is used to correct the cumulative error of the 3D model. If it is determined that the first information should be displayed on the display unit, the first information shall be displayed on the display unit.

2. The three-dimensional model generation apparatus according to claim 1, wherein, The first information is a second image included in the plurality of images, taken at the first position and orientation.

3. The three-dimensional model generation apparatus according to claim 2, wherein, The circuit causes the second image to be displayed on the display unit in an overlapping manner with the currently captured image.

4. The three-dimensional model generation apparatus according to claim 1, wherein, The circuit further, The second piece of information representing the cumulative error is displayed on the display unit.

5. The three-dimensional model generation apparatus according to claim 4, wherein, As the second piece of information, the third image obtained by the circuit from observing the three-dimensional model from the current position and posture is superimposed on the currently captured image and displayed on the display unit.

6. The three-dimensional model generation apparatus according to claim 5, wherein, As the second piece of information, the circuit further displays the areas in the third image that differ from the currently captured image on the display unit.

7. The three-dimensional model generation apparatus according to claim 5, wherein, The circuit further, The amount of the cumulative error is calculated based on the difference between the third image and the currently captured image. The third piece of information representing the calculated quantity is displayed on the display unit.

8. The three-dimensional model generation apparatus according to claim 4, wherein, The circuit further, The user interface for indicating whether the correction should be performed is displayed on the display unit along with the second information. Upon receiving an instruction to perform the correction, the first information is displayed on the display unit.

9. The three-dimensional model generation apparatus according to claim 2, wherein, The circuit further, A user interface for allowing a user to select the second image from a plurality of fourth images included in the plurality of images is displayed on the display unit.

10. The three-dimensional model generation apparatus according to claim 9, wherein, The circuit further, Based on the feature points contained in each of the multiple images, the multiple fourth images are selected from the multiple images.

11. The three-dimensional model generation apparatus according to claim 2, wherein, The circuit further, The second image is selected from the plurality of images based on the feature points contained in each of the plurality of images.

12. The three-dimensional model generation apparatus according to claim 1, wherein, When the first information is displayed on the display unit, (1) Having obtained the first image captured from the first position and pose, the correction is performed using the obtained first image. (2) In the case of obtaining a sixth image obtained from the position and pose of the fifth image among the plurality of images, the sixth image obtained is not used to perform the correction of the cumulative error, wherein the fifth image was obtained from a position and pose other than the first position and pose.

13. The three-dimensional model generation apparatus according to claim 1, wherein, The circuit, The multiple images are used to estimate the estimated position and pose, and the estimated position and pose are used to generate the 3D model. The estimated position and pose are the position and pose of each image among the multiple images. In the correction of the accumulated error, the accumulated error is corrected based on the error between the estimated position and pose of the second image captured at the first position and pose and the estimated position and pose of the first image, which is included in the plurality of images.

14. The three-dimensional model generation apparatus according to claim 4, wherein, As the second piece of information, the circuit displays the current 3D model and the past 3D model on the display unit. The circuit further, Accept operations for aligning the current 3D model with the position of the past 3D model. In the correction, information obtained through operations to align the position is used as auxiliary information for the correction.

15. A method for generating a three-dimensional model, wherein, Obtain multiple consecutive images. The multiple images are used to generate a three-dimensional model, which includes multiple three-dimensional points, each with its own location information. A determination is made as to whether to display the first information used to indicate the first position and attitude on the display unit. The first position and attitude is the position and attitude at which the first image was captured. The first image is used to correct the cumulative error of the 3D model. If it is determined that the first information should be displayed on the display unit, the first information shall be displayed on the display unit.

16. A program for causing a computer to perform the three-dimensional model generation method of claim 15.