Photography method and photography device

By detecting and indicating the photography indicator device that optimizes the photography position and posture in real time, the problem of low three-dimensional model generation accuracy is solved, and efficient and accurate three-dimensional model reconstruction and user-friendly photography instructions are achieved.

CN115336250BActive Publication Date: 2025-07-04PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180023034.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-27
Filing Date
2021-03-24
Publication Date
2025-07-04
Estimated Expiration
2041-03-24

AI Technical Summary

Technical Problem

In the prior art, when generating a three-dimensional model, it is difficult for the model to be improved, and it is difficult for the user to judge the required image, resulting in failure of reconstruction or reduction of accuracy.

Method used

The photography indication device detects areas that are difficult to generate a three-dimensional model in real time, and indicates the optimization of the photography position and posture, generates a three-dimensional model image that is easy to generate a specified area. Combined with SLAM and image processing technology, real-time indication of the photography position and posture are provided.

Benefits of technology

It improves the accuracy and generation efficiency of the three-dimensional model, reduces reconstruction failure, and enhances user convenience and photography effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115336250B_ABST
    Figure CN115336250B_ABST
Patent Text Reader

Abstract

The photographic indication method is executed by a photographic indication device. In order to generate a three-dimensional model of a subject based on the photographic positions and postures of a plurality of images obtained by photographing the subject and the plurality of images, it accepts the designation of a first region (S205), and indicates at least one of the photographic position and the posture to photograph an image for generating the three-dimensional model of the designated first region. For example, the photographic indication method may also detect a second region in which it is difficult to generate a three-dimensional model based on the respective photographic positions and postures and the plurality of images, and indicate at least one of the photographic position and the posture to generate an image in which it is easy to generate the three-dimensional model of the second region. In the indication corresponding to the first region, it indicates at least one of the photographic position and the posture to photograph an image in which it is easy to generate the three-dimensional model of the first region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a photographing instruction method, a photographing method, a photographing instruction device, and a photographing device. Background Art

[0002] A technique is disclosed in Patent Document 1, in which a three-dimensional model of a subject is generated using a plurality of images obtained by photographing the subject from a plurality of viewpoints.

[0003] Prior Art Documents

[0004] Patent Documents

[0005] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2017-130146 Summary of the Invention

[0006] Problems to be Solved by the Invention

[0007] In the generation process of the three-dimensional model, it is desired to improve the accuracy of the three-dimensional model. An object of the present disclosure is to provide a photographing instruction method or a photographing instruction device capable of improving the accuracy of the three-dimensional model.

[0008] Means for Solving the Problems

[0009] The photographing instruction method according to one aspect of the present disclosure is executed by a photographing instruction device. In order to generate a three-dimensional model of a subject based on the photographing positions and postures of a plurality of images obtained by photographing the subject, and the plurality of images, the designation of a first region is accepted, and at least one of the photographing position and the posture is instructed to photograph an image for generating the three-dimensional model of the designated first region.

[0010] The photographing method according to one aspect of the present disclosure is executed by a photographing device. A plurality of first images of an object space are photographed, first three-dimensional position information of the object space is generated based on the plurality of first images and the first photographing positions and postures of the plurality of first images, and for a second region in the object space where it is difficult to generate second three-dimensional position information more detailed than the first three-dimensional position information, the second three-dimensional position information is not generated and the first three-dimensional position information is used for determination.

[0011] Advantages of the Invention

[0012] The present disclosure can provide a photographing instruction method or a photographing instruction device capable of improving the accuracy of the three-dimensional model. Brief Description of the Drawings

[0013] Figure 1 It is a block diagram of a terminal device according to Embodiment 1.

[0014] Figure 2It is a timing chart of the terminal device according to Embodiment 1.

[0015] Figure 3 It is a flowchart of the initial processing according to Embodiment 1.

[0016] Figure 4 It is a diagram showing an example of the initial display according to Embodiment 1.

[0017] Figure 5 It is a diagram showing an example of the method for selecting the priority designation position according to Embodiment 1.

[0018] Figure 6 It is a diagram showing an example of the method for selecting the priority designation position according to Embodiment 1.

[0019] Figure 7 It is a flowchart of the position and attitude estimation processing according to Embodiment 1.

[0020] Figure 8 It is a flowchart of the photographing position candidate determination processing according to Embodiment 1.

[0021] Figure 9 It is a diagram showing the state of the camera and the object according to Embodiment 1 as viewed from above.

[0022] Figure 10 It is a diagram showing an example of the image obtained by each camera according to Embodiment 1.

[0023] Figure 11 It is a schematic diagram for explaining an example of the photographing position candidate determination according to Embodiment 1.

[0024] Figure 12 It is a schematic diagram for explaining an example of the photographing position candidate determination according to Embodiment 1.

[0025] Figure 13 It is a schematic diagram for explaining an example of the photographing position candidate determination according to Embodiment 1.

[0026] Figure 14 It is a flowchart of the three-dimensional reconstruction processing according to Embodiment 1.

[0027] Figure 15 It is a flowchart of the display processing during photographing according to Embodiment 1.

[0028] Figure 16 It is a diagram showing an example of the method for visually presenting the photographing position candidate according to Embodiment 1.

[0029] Figure 17This is a diagram showing an example of a method for visually presenting a candidate for a photographing position according to Embodiment 1.

[0030] Figure 18 This is a diagram showing an example of the display of an alarm according to Embodiment 1.

[0031] Figure 19 This is a flowchart of the photographing instruction process according to Embodiment 1.

[0032] Figure 20 This is a diagram showing the configuration of a three-dimensional reconstruction system according to Embodiment 2.

[0033] Figure 21 This is a block diagram of a photographing device according to Embodiment 2.

[0034] Figure 22 This is a flowchart of the operation of a photographing device according to Embodiment 2.

[0035] Figure 23 This is a flowchart of the position and orientation estimation process according to Embodiment 2.

[0036] Figure 24 This is a flowchart of the position and orientation integration process according to Embodiment 2.

[0037] Figure 25 This is a plan view showing a situation of photographing in an object space according to Embodiment 2.

[0038] Figure 26 This is a diagram showing an example of an image and an example of comparison processing according to Embodiment 2.

[0039] Figure 27 This is a flowchart of the area detection process according to Embodiment 2.

[0040] Figure 28 This is a flowchart of the display process according to Embodiment 2.

[0041] Figure 29 This is a diagram showing an example of the display of a UI screen according to Embodiment 2.

[0042] Figure 30 This is a diagram showing an example of area information according to Embodiment 2.

[0043] Figure 31 This is a diagram showing an example of the display in the case where the estimation of the position and orientation according to Embodiment 2 fails.

[0044] Figure 32 This is a diagram showing an example of the display in the case where a low-precision area is detected according to Embodiment 2.

[0045] Figure 33 This is a diagram showing an example of an instruction to a user according to Embodiment 2.

[0046] Figure 34 This is a diagram showing an example of an instruction (arrow) according to Embodiment 2.

[0047] Figure 35 This is a diagram showing an example of area information according to Embodiment 2.

[0048] Figure 36 This is a plan view showing the photographing state of the object area according to Embodiment 2.

[0049] Figure 37 This is a diagram showing an example of a photographed completion area in the case of using three-dimensional points according to Embodiment 2.

[0050] Figure 38 This is a diagram showing an example of a photographed completion area in the case of using a grid according to Embodiment 2.

[0051] Figure 39 This is a diagram showing an example of a depth image according to Embodiment 2.

[0052] Figure 40 This is a diagram showing an example of a photographed completion area in the case of using a depth image according to Embodiment 2.

[0053] Figure 41 This is a flowchart of a photographing method according to Embodiment 2. Detailed Embodiment

[0054] The photographing instruction method according to one aspect of the present disclosure is executed by a photographing instruction device. In order to generate a three-dimensional model of a subject based on the photographing positions and postures of a plurality of images obtained by photographing the subject and the plurality of images, it accepts the designation of a first area and instructs at least one of the photographing position and posture to photograph an image for generating the three-dimensional model of the designated first area.

[0055] Thereby, it is possible to preferentially improve the accuracy of the three-dimensional model of the area required by the user, and thus it is possible to improve the accuracy of the three-dimensional model.

[0056] For example, based on the respective photographing positions and postures and the plurality of images, a second area in which it is difficult to generate the three-dimensional model is detected, and at least one of the photographing position and posture is instructed to generate an image that is easy to generate the three-dimensional model of the second area. In the instruction corresponding to the first area, at least one of the photographing position and posture is instructed to photograph an image that is easy to generate the three-dimensional model of the first area.

[0057] For example, an image of the subject after the recognition of the executed attribute is displayed, and the designation of the attribute is accepted in the designation of the first region.

[0058] For example, it may also be that in the detection of the second region, (i) on a two-dimensional image, an edge whose angle difference from the epipolar line based on the photographing position and posture is smaller than a predetermined value is obtained, and (ii) a three-dimensional region corresponding to the obtained edge is detected as the second region, and in the instruction corresponding to the second region, at least one of the photographing position and posture is instructed to photograph an image in which the angle difference is larger than the value.

[0059] For example, it may also be that the plurality of images are a plurality of frames included in a moving image being currently photographed and displayed, and the instruction corresponding to the second region is performed in real time.

[0060] Thus, by performing the photographing instruction in real time, the convenience of the user can be improved.

[0061] For example, it may also be that in the instruction corresponding to the second region, the photographing direction is instructed.

[0062] Thus, the user can easily perform appropriate photographing according to the instruction.

[0063] For example, it may also be that in the instruction corresponding to the second region, the photographing area is instructed.

[0064] Thus, the user can easily perform appropriate photographing according to the instruction.

[0065] In addition, the photographing instruction device according to one aspect of the present disclosure includes a processor and a memory, and in order to generate a three-dimensional model of the subject based on the photographing position and posture of each of a plurality of images obtained by photographing the subject and the plurality of images, the designation of the first region is accepted, and at least one of the photographing position and posture is instructed to photograph an image for generating the three-dimensional model of the designated first region.

[0066] Thus, the accuracy of the three-dimensional model of the region required by the user can be preferentially improved, and therefore the accuracy of the three-dimensional model can be improved.

[0067] The photographing instruction method according to one aspect of the present disclosure is based on the photographing position and posture of each of a plurality of images obtained by photographing a subject and the plurality of images, detects a region where it is difficult to generate a three-dimensional model of the subject by using the plurality of images, and instructs at least one of the photographing position and posture to photograph an image that is easy to generate the three-dimensional model of the detected region.

[0068] Thus, the accuracy of the three-dimensional model can be improved.

[0069] For example, the photographing instruction method may further accept the designation of a priority area, and in the instruction, at least one of the photographing position and posture is instructed to photograph an image that can easily generate a three-dimensional model of the designated priority area.

[0070] Thereby, the accuracy of the three-dimensional model of the area required by the user can be preferentially improved.

[0071] The photographing method according to one aspect of the present disclosure is executed by a photographing device, photographs a plurality of first images of an object space, generates first three-dimensional position information of the object space based on the plurality of first images and the first photographing position and posture of each of the plurality of first images, and for a second area in which it is difficult to generate second three-dimensional position information of the object space that is more detailed than the first three-dimensional position information, the second three-dimensional position information is not generated and the first three-dimensional position information is used for determination.

[0072] Thereby, the photographing method can use the first three-dimensional position information to determine the second area in which it is difficult to generate the second three-dimensional position information without generating the second three-dimensional position information, and thus can improve the efficiency of photographing a plurality of images for generating the second three-dimensional position information.

[0073] For example, the second area may be at least one of an area where photographing of an image has not been performed and an area where the accuracy of the second three-dimensional position information is estimated to be lower than a predetermined reference.

[0074] For example, the first three-dimensional position information may include a first three-dimensional point cloud, and the second three-dimensional position information may include a second three-dimensional point cloud that is denser than the first three-dimensional point cloud.

[0075] For example, in the determination, a third area of the object space corresponding to an area around the first three-dimensional point cloud is determined, and an area other than the third area is determined as the second area.

[0076] For example, in the determination, a grid is generated using the first three-dimensional point cloud, and an area other than a third area of the object space corresponding to the area where the grid is generated is determined as the second area.

[0077] For example, in the determination, the second area is determined based on the reprojection error of the first three-dimensional point cloud.

[0078] For example, the first three-dimensional position information may include a depth image, and based on the depth image, an area within a predetermined distance from the photographing viewpoint is determined as a third area, and an area other than the third area is determined as the second area.

[0079] For example, the photography method may further use a plurality of second images that have been photographed, the second photographing positions and postures of the plurality of second images, the plurality of first images, and the plurality of first photographing positions and postures to make the coordinate systems of the plurality of first photographing positions and postures correspond to the coordinate systems of the plurality of second photographing positions and postures.

[0080] Thereby, it is possible to use the information obtained through multiple photographings to determine the second area.

[0081] For example, in photographing the object space, the photography method may further display the second area or a third area outside the second area.

[0082] Thereby, the second area can be presented to the user.

[0083] For example, it may also be to display the information indicating the second area or the third area by overlapping it on one of the plurality of images.

[0084] Thereby, the position of the second area within the image can be presented to the user, so that the user can easily grasp the position of the second area.

[0085] For example, it may also be to display the information indicating the second area or the third area by overlapping it on the map of the object space.

[0086] Thereby, the position of the second area in the surrounding environment can be presented to the user, so that the user can easily grasp the position of the second area.

[0087] For example, it may also be to display the second area and the restoration accuracy of each area included in the second area.

[0088] Thereby, the user can not only grasp the second area but also the restoration accuracy of each area, so that appropriate photography can be performed accordingly.

[0089] For example, the photography method may further give an instruction to the user for photographing the second area.

[0090] Thereby, the user can effectively perform appropriate photography.

[0091] For example, the instruction may include at least one of the direction and distance from the current position to the second area.

[0092] Thereby, the user can effectively perform appropriate photography.

[0093] In addition, a photographic apparatus according to one aspect of the present disclosure includes a processor and a memory. The processor uses the memory to photograph a plurality of first images of a photographic object space, and generates first three-dimensional position information of the object space based on the plurality of first images, and the first photographing position and posture of each of the plurality of first images. For a second region in which it is difficult to generate second three-dimensional position information of the object space that is more detailed than the first three-dimensional position information using the plurality of first images, the second three-dimensional position information is not generated, and the first three-dimensional position information is used for determination.

[0094] Thus, the photographic apparatus can determine a second region in which it is difficult to generate second three-dimensional position information using the first three-dimensional position information without generating the second three-dimensional position information, and thus can improve the efficiency of photographing a plurality of images for generating the second three-dimensional position information.

[0095] In addition, these general or specific aspects can also be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0096] Hereinafter, the embodiments will be specifically described with reference to the drawings. In addition, all of the embodiments described below represent a specific example of the present disclosure. The numerical values, shapes, materials, constituent elements, arrangement positions and connection methods of the constituent elements, steps, order of steps, etc. shown in the following embodiments are examples, and are not intended to limit the present disclosure. In addition, among the constituent elements in the following embodiments, the constituent elements not described in the independent claims are described as optional constituent elements.

[0097] (Embodiment 1)

[0098] By using a plurality of images photographed by a camera to generate a three-dimensional model, it is possible to generate a three-dimensional map or the like more simply than a method of generating a three-dimensional model using laser measurement. Therefore, in the case of measuring distances in construction site management or the like, a method of generating a three-dimensional model by using images is adopted. Here, the three-dimensional model is a model that represents a measured object photographed on a computer. The three-dimensional model has, for example, position information of each three-dimensional position on the measured object.

[0099] However, when generating a three-dimensional model using images, it is necessary to have a plurality of images with parallax that reflect the same position. In addition, the more the number of images, the denser the three-dimensional model can be generated. In addition, the parallax affects the restoration accuracy. Therefore, in the case of measuring distances, it is better to photograph while appropriately moving relative to the subject. However, it is difficult to grasp an appropriate photographing position during photographing while corresponding to the state of the subject.

[0100] In the present embodiment, a UI (user interface) or the like that detects a region where three-dimensional reconstruction (generation of a three-dimensional model) is difficult and indicates a photographing position or a photographing attitude based on the detection result will be described. Thereby, it is possible to suppress the failure of three-dimensional model reconstruction or the reduction of accuracy. In addition, the occurrence of re-photographing can be reduced, so that the work efficiency can be improved.

[0101] First, the configuration of a terminal device 100 as an example of a photographing instruction device according to the present embodiment will be described. Figure 1 It is a block diagram of the terminal device 100 according to the present embodiment. The terminal device 100 has a photographing function, a function of estimating the three-dimensional position and attitude during photographing, a function of determining a candidate for the next photographing position based on the photographed image, and a function of presenting the estimated candidate for the photographing position to the user. In addition, the terminal device 100 may have a function of performing three-dimensional reconstruction using the estimated three-dimensional position and attitude and generating a three-dimensional model (point cloud) of the photographing environment, a function of determining a candidate for the next photographing position using the three-dimensional model, a function of presenting the estimated candidate for the photographing position to the user, and a function of transmitting and receiving at least one of a photographed image, a three-dimensional position and attitude, and a three-dimensional model to and from other terminal devices and a management server.

[0102] The terminal device 100 includes a photographing unit 101, a control unit 102, a position and attitude estimation unit 103, a three-dimensional reconstruction unit 104, an image analysis unit 105, a point cloud analysis unit 106, a communication unit 107, a UI unit 108, an image storage unit 111, a camera attitude storage unit 112, and a three-dimensional model storage unit 113.

[0103] The photographing unit 101 is a photographing device such as a camera and acquires an image (moving image). In addition, although an example of using an image will be mainly described below, a plurality of still images may be used instead of the image. The photographing unit 101 stores the acquired image in the image storage unit 111. In addition, the photographing unit 101 can photograph either a visible light image or an infrared image. When an infrared image is used, photographing can be performed even in a dark environment such as at night. In addition, the photographing unit 101 can be either a monocular camera or a multi-camera such as a stereo camera. By using a calibrated stereo camera, the accuracy of the three-dimensional position and attitude can be improved. Even when the stereo camera is not calibrated, a parallax image with parallax can be acquired.

[0104] The control unit 102 controls the overall imaging process and the like of the terminal device 100. The position and attitude estimation unit 103 estimates the three-dimensional position and attitude of the camera that captured the image using the images stored in the image storage unit 111. In addition, the position and attitude estimation unit 103 stores the estimated three-dimensional position and attitude in the camera attitude storage unit 112. For example, the position and attitude estimation unit 103 uses image processing such as SLAM (Simultaneous Localization and Mapping) to estimate the position and attitude of the camera. Alternatively, the position and attitude estimation unit 103 can also use the information obtained from various sensors (GPS or acceleration sensors) equipped in the terminal device 100 to calculate the position and attitude of the camera. In the former case, the position and attitude can be estimated based on the information from the imaging unit 101. In the latter case, image processing can be achieved with less processing.

[0105] The three-dimensional reconstruction unit 104 performs three-dimensional reconstruction using the images stored in the image storage unit 111 and the three-dimensional position and attitude stored in the camera attitude storage unit 112 to generate a three-dimensional model. In addition, the three-dimensional reconstruction unit 104 stores the generated three-dimensional model in the three-dimensional model storage unit 113. For example, the three-dimensional reconstruction unit 104 uses image processing represented by SfM (Structure from Motion) for three-dimensional reconstruction. Alternatively, when using images obtained from a camera calibrated by a stereo camera or the like, the three-dimensional reconstruction unit 104 can also utilize stereo disparity. In the former case, a high-precision three-dimensional model can be generated by using a large number of images. In the latter case, a three-dimensional model can be generated at high speed with less processing.

[0106] In addition, when SLAM is used in the position and attitude estimation unit 103, a three-dimensional model of the surrounding environment is also generated while estimating the three-dimensional position and attitude. Thus, this three-dimensional model can also be utilized.

[0107] In the image analysis unit 105, in order to perform three-dimensional reconstruction with high precision using the images stored in the image storage unit 111 and the three-dimensional position and attitude stored in the camera attitude storage unit 112, it analyzes which position is better for photography and determines candidate photography positions based on the analysis results. Information indicating the determined candidate photography positions is output to the UI unit 108 and presented to the user.

[0108] The point group analysis unit 106 determines the density of the point group included in the three-dimensional model, using the image stored in the image storage unit 111, the three-dimensional position and attitude stored in the camera attitude storage unit 112, and the three-dimensional model in the three-dimensional model storage unit 113. The point group analysis unit 106 determines candidate photographing positions for mapping sparse regions. In addition, the point group analysis unit 106 detects a region of the point group generated by using the peripheral region of an image that is prone to lens distortion or the like, and determines candidate photographing positions for mapping this region to the center of the camera. The determined candidate photographing positions are output to the UI unit 108 and presented to the user.

[0109] The communication unit 107 communicates and receives the photographed image, the calculated three-dimensional attitude, and the three-dimensional model with a cloud server or other terminal device via communication.

[0110] The UI unit 108 presents the photographed image and the candidate photographing positions determined by the image analysis unit 105 and the point group analysis unit 106 to the user. In addition, the UI unit 108 has an input function for inputting a photographing start instruction, a photographing end instruction, and a priority processing position from the user.

[0111] Next, the operation of the terminal device 100 according to the present embodiment will be described. Figure 2 is a timing chart showing information transmission and reception and the like inside the terminal device 100. In Figure 2 , the region shown in light ink indicates that the imaging unit 101 is continuously performing photography. In order to generate a high-quality three-dimensional model, the terminal device 100 analyzes the image, as well as the photographing position and attitude, in real time and gives a photographing instruction to the user. Here, real-time analysis means analyzing while performing photography. Alternatively, real-time analysis means analyzing without generating a three-dimensional model. Specifically, the terminal device 100 estimates the position and attitude of the camera during photography, and based on the estimation result and the photographed image, determines a region that is difficult to restore. The terminal device 100 predicts the photographing position and attitude that can ensure a parallax that is easy to restore this region, and presents the predicted photographing position and attitude using the UI. Here, a sequence during dynamic image photography is presented, but the same process can also be implemented for each still image photography.

[0112] First, the UI unit 108 performs initial processing (S101). As a result, the UI unit 108 sends a photographing start signal to the imaging unit 101. The start processing is performed, for example, by the user clicking the "Photography Start" button on the display of the terminal device 100.

[0113] Next, the UI unit 108 performs in-photography display processing (S102). Specifically, the UI unit 108 presents the image and instructions to the user during photography.

[0114] The imaging unit 101 that has received the photography start signal performs imaging of an image, and transmits the image information of the image obtained by the imaging as an image to the position and orientation estimation unit 103, the three-dimensional reconstruction unit 104, the image analysis unit 105, and the point cloud analysis unit 106. For example, the imaging unit 101 can perform streaming transmission that appropriately transmits the image during imaging, or can transmit the images of that time together at every certain time. That is, the image information is one or more images (frames) included in the image. In the former case, processing can be appropriately performed, so that the waiting time for generating the three-dimensional model can be reduced. In the latter case, a large amount of information obtained by imaging can be utilized, so that high-precision processing can be achieved.

[0115] The position and orientation estimation unit 103 first performs an input waiting process at the start of imaging, and enters a state of waiting for image information from the imaging unit 101. If the position and orientation estimation unit 103 is input with image information from the imaging unit 101, it performs a position and orientation estimation process (S103). That is, the position and orientation estimation process is performed for each frame or for every plurality of frames. When the position and orientation estimation process fails, the position and orientation estimation unit 103 transmits an estimation failure signal to the UI unit 108 in order to notify the user of the failure. When the position and orientation estimation process is successful, the position and orientation estimation unit 103 transmits position and orientation information as an estimation result of the current three-dimensional position and orientation to the UI unit 108. In addition, the position and orientation estimation unit 103 transmits the position and orientation information to the image analysis unit 105 and the three-dimensional reconstruction unit 104.

[0116] The image analysis unit 105 first performs an input waiting process at the start of imaging, and enters a state of waiting for image information from the imaging unit 101 and position and orientation information from the position and orientation estimation unit 103. If the image analysis unit 105 is input with image information and position and orientation information, it performs a photography position candidate determination process (S104). In addition, the photography position candidate determination process can be performed for each frame or at every certain time (a plurality of frames) (for example, every 5 seconds). In addition, the image analysis unit 105 can also determine whether the terminal device 100 is moving toward the photography position candidate generated by the photography position candidate determination process, and if it is moving, does not perform a new photography position candidate determination process. For example, if the current position and orientation are on the straight line connecting the position and orientation of the image at the time when the photography position candidate is to be determined and the calculated candidate position and orientation, the image analysis unit 105 determines that it is moving.

[0117] The 3D reconstruction unit 104 first performs an input waiting process at the start of photography, entering a state of waiting for image information from the imaging unit 101 and position and orientation information from the position and orientation estimation unit 103. If the 3D reconstruction unit 104 is input with image information and position and orientation information, it performs 3D reconstruction processing (S105) to calculate a 3D model. The 3D reconstruction unit 104 sends the point cloud information of the calculated 3D model to the point cloud analysis unit 106.

[0118] The point cloud analysis unit 106 first performs an input waiting process at the start of photography, entering a state of waiting for point cloud information from the 3D reconstruction unit 104. If the point cloud analysis unit 106 is input with point cloud information, it performs a photography position candidate determination process (S106). For example, the point cloud analysis unit 106 determines the density state of the entire point cloud and detects sparse regions. The point cloud analysis unit 106 determines photography position candidates that show more of these sparse regions. In addition, the point cloud analysis unit 106 can also determine photography position candidates using not only point cloud information but also image information or position and orientation information.

[0119] Next, the initial process (S101) will be described. Figure 3 It is a flowchart of the initial process (S101). First, the UI unit 108 displays the current captured image (S201). Next, the UI unit 108 obtains whether there is a position that the user wishes to prioritize for restoration, i.e., a priority position (S202). For example, the UI unit 108 displays a button for specifying the priority mode, and determines that there is a priority position when this button is selected.

[0120] In the case of having a priority position (S203: Yes), the UI unit 108 displays a priority position selection screen (S204) and obtains information on the priority position selected by the user (S205). After step S205, or in the case of not having a priority position (S203: No), next, the UI unit 108 outputs a photography start signal to the imaging unit 101. Thereby, photography starts (S206). For example, photography can be started either by the user pressing a button or automatically after a certain period of time has elapsed.

[0121] Here, for the set priority position, the priority for restoration when instructing the camera to move, etc. is set high. Thus, it is possible to prevent a situation where the camera is instructed to move to restore an area that is difficult to restore but not needed by the user, and it is possible to indicate the area required by the user as the object.

[0122] Figure 4It is a diagram showing an example of an initial display indicating the initial process (S101). A captured image 201, a priority designation button 202 for selecting whether to select a priority designated position, and a photographing start button 203 for starting photographing are displayed. In addition, the captured image 201 can be either a still image or an image being currently captured (moving image).

[0123] Figure 5 It is a diagram showing an example of a method for selecting a priority designated position during priority designation. In this diagram, object recognition is performed on an object such as a window frame, and using a selection bar 204, a desired object is selected from a list of multiple objects included in the image. For example, for an image, labels such as window frame, table, and wall are attached to each pixel by methods such as semantic segmentation, and object pixels are selected together by selecting the label.

[0124] Figure 6 It is a diagram showing an example of another method for selecting a priority designated position during priority designation. In this diagram, the user designates an arbitrary area (an area surrounded by a rectangle in the diagram) by an indicator or a touch operation to select the priority position. In addition, as long as the user can select a specific area, the means is not limited. For example, selection can also be made by designating a color, or a rough designation such as the right area in the image can also be made.

[0125] In addition, input methods other than operating on the screen can also be used. For example, selection operations can also be performed by voice input. In this case, when it is difficult to operate with hands in a state of wearing gloves in a cold region, the input becomes easier.

[0126] In addition, an example of selecting a priority position during the initial process before photographing is shown here, but a priority position can also be appropriately added during photographing. Thereby, a position that is not reflected in the initial state can be selected.

[0127] Next, the position and orientation estimation process (S103) will be described. Figure 7 It is a flowchart of the position and orientation estimation process (S103). First, the position and orientation estimation unit 103 acquires one or more images or videos from the video storage unit 111 (S301). Next, the position and orientation estimation unit 103 calculates or acquires position and orientation information corresponding to the amount of the input image, and this position and orientation information includes the three-dimensional position and orientation (direction) of the camera, as well as camera parameters including lens information, etc. (S302). For example, the position and orientation estimation unit 103 performs image processing such as SLAM or SfM on the image acquired in S301 to calculate the position and orientation information.

[0128] Next, the position and orientation estimation unit 103 stores the position and orientation information obtained in S302 in the camera pose storage unit 112 (S303).

[0129] Alternatively, the image input in S301 may be an image sequence of a certain time amount composed of multiple frames, and the processing after S302 is performed on this image sequence (multiple images). Or, images may be input successively like streaming, and the processing after S302 is repeatedly performed for each input image. In the former case, the accuracy can be improved by using the information of multiple time amounts. In the latter case, since input can be performed successively, a fixed-length input delay can be ensured, and the waiting time required for generating the three-dimensional model can be reduced.

[0130] Next, the photographing position candidate determination process (S104) performed by the image analysis unit 105 will be described. Figure 8 This is a flowchart of the photographing position candidate determination process (S104). First, the image analysis unit 105 acquires multiple images or videos from the video storage unit 111 (S401). In addition, one of the acquired multiple images is set as a key image. Here, the key image is an image used as a reference when performing subsequent three-dimensional reconstruction. For example, the depth at each pixel of the key image is estimated using the information of other images except the key image, and three-dimensional reconstruction is performed using the estimated depth. Next, the image analysis unit 105 acquires the position and orientation information of each of the multiple images (S402).

[0131] Next, the image analysis unit 105 calculates the epipolar lines between the images (between the cameras) using the position and orientation information of each image (S403). Next, the image analysis unit 105 detects the edges in each image (S404). For example, the image analysis unit 105 detects the edges through filtering processing such as a Sobel filter.

[0132] Next, the image analysis unit 105 calculates the angle between the epipolar line and the edge in each image (S405). Next, the image analysis unit 105 calculates the restoration difficulty of each pixel of the key image based on the angle obtained in S405 (S406). Specifically, the more parallel the epipolar line and the edge are, the more difficult the three-dimensional reconstruction is, so the smaller the angle, the higher the restoration difficulty is set. In addition, the restoration difficulty can be set in multiple stages or in 2 stages of high / low. For example, it may be that when the angle is smaller than a predetermined value (for example, 5 degrees), the restoration difficulty is set high, and when the angle is larger than the predetermined value, the restoration difficulty is set low.

[0133] Next, based on the restoration difficulty calculated in S406, the image analysis unit 105 estimates from which position photographing would facilitate the restoration of areas with high restoration difficulty, and determines the estimated position as a candidate for the photographing position (S407). Specifically, an area with high restoration difficulty is synonymous with being on a plane in the same direction as the movement direction between the camera and the object. By moving the camera in a direction perpendicular to this plane, the epipolar line and the edge no longer become horizontal. Therefore, if the camera is moving forward, by moving the camera in the up-down direction or the left-right direction, the restoration difficulty can be reduced in areas with high restoration difficulty.

[0134] In addition, the restoration difficulty is obtained through the processing of S401 to S406, but as long as the restoration difficulty can be calculated, it is not limited to this. For example, the degradation of image quality caused by lens distortion is greater at the image edges than at the image center. Therefore, the image analysis unit 105 can also determine an object that appears only at the edges of the screen in each image, and set a high restoration difficulty for the area of this object. For example, the image analysis unit 105 can also determine the field of view of the camera in the three-dimensional space based on the position and orientation information, and determine the area that appears only at the edges of the screen according to the overlap of the fields of view of each camera.

[0135] Hereinafter, it is explained that the difficulty of restoration varies according to the angle formed by the epipolar line and the edge. Figure 9 It is a diagram showing the situation of observing the camera and the object from above. Figure 10 It is shown in Figure 9 A diagram showing an example of an image obtained by each camera in the shown situation.

[0136] In this case, when searching for corresponding points in the image of camera B with respect to a point in the image of camera A, the search is performed on the epipolar line that can be calculated based on camera geometry. When the epipolar line is parallel to the edge on the image of the object like point A, it is difficult to perform matching using pixel information such as NCC (Normalized Cross Correlation), and it is difficult to determine the correct corresponding point.

[0137] On the other hand, when the epipolar line is perpendicular to the edge on the image of the object like point B, it is easy to perform matching and the correct corresponding point can be determined. That is, by calculating the angle between the epipolar line and the edge on the image, the difficulty of determining the corresponding point can be judged. Correctly obtaining this corresponding point is related to the accuracy of three-dimensional reconstruction. Thus, the angle between the epipolar line and the edge on the image can be used as the difficulty of three-dimensional reconstruction. In addition, as long as the information has the same meaning as the angle, for example, the epipolar line and the edge can be regarded as vectors, and the inner product value of the epipolar line and the edge can also be used.

[0138] In addition, the epipolar line can be calculated using the fundamental matrix between camera A and camera B. This fundamental matrix can be calculated based on the position and attitude information of camera A and camera B. When the internal matrices of camera A and camera B are set as KA and KB, the relative rotation matrix of camera B observed from camera A is set as R, and the relative movement vector is set as T, the epipolar line can be obtained as follows.

[0139] For a pixel (x, y) on camera A, (a, b, c) is calculated by the following formula, and the epipolar line on camera B can be expressed as a straight line satisfying ax + by + c = 0.

[0140] [Equation 1]

[0141] F = K B -T *([t X *R)*K A -1

[0142]

[0143] Hereinafter, an example of determining the candidate photography position will be described. Figures 11 to 13 is a schematic diagram for explaining an example of determining the candidate photography position. Edges with high restoration difficulty, as Figure 11 shown, are mostly straight lines or line segments on a three-dimensional plane passing through the straight line connecting the three-dimensional positions of camera A and camera B. Conversely, the more perpendicular a straight line is to the plane passing through this straight line, the larger the angle between the epipolar line and the edge in the matching between camera A and camera B, and the lower the restoration difficulty.

[0144] That is, as Figure 12 shown, when determining the candidate photography position, for the target edge, if photographing is performed from the position of camera C in the direction of perpendicular movement relative to the plane including the edge under the above conditions, the restoration difficulty at the target edge can be reduced when matching with camera A or camera B. Therefore, the image analysis unit 105 determines camera C as a candidate photography position.

[0145] In addition, when there is no target edge, the image analysis unit 105 can also use the edge with the highest restoration difficulty as a candidate and determine the candidate photography position using the above method. Or, the image analysis unit 105 can also randomly select an edge from the top 10 edges with the lowest restoration difficulty as a candidate.

[0146] In addition, here only the information of a pair of cameras (camera A and camera B) is used to calculate camera C, but when there are multiple edges with high restoration difficulty, it is also possible to, as Figure 13As shown, the image analysis unit 105 determines the candidate photography positions (camera C) for the first edge and the candidate photography positions (camera C) for the second edge, and outputs the path connecting camera C and camera D.

[0147] In addition, the method for determining the candidate photography positions is not limited to this. The image analysis unit 105 can also consider that the closer to the center of the screen, the less the influence of distortion and other factors on image quality degradation. When targeting the edge reflected at the end of the screen by camera A, it determines the position where the edge is reflected at the center of the screen as the candidate photography position.

[0148] Next, the 3D reconstruction process (S105) will be described. Figure 14 It is a flowchart of the 3D reconstruction process (S105). First, the 3D reconstruction unit 104 obtains multiple images or videos from the image storage unit 111 (S501). Next, the 3D reconstruction unit 104 obtains the position and attitude information (camera parameters) of each of the multiple images from the camera pose storage unit 112 (S502).

[0149] Next, the 3D reconstruction unit 104 performs 3D reconstruction using the obtained multiple images and multiple position and attitude information to generate a 3D model (S503). For example, the 3D reconstruction unit 104 performs 3D reconstruction using the visual volume intersection method or SfM. Finally, the 3D reconstruction unit 104 stores the generated 3D model in the 3D model storage unit 113 (S504).

[0150] In addition, the process of S503 may not be performed by the terminal device 100. For example, the terminal device 100 sends images and camera parameters to a cloud server or the like. The cloud server performs 3D reconstruction to generate a 3D model. The terminal device 100 receives the 3D model from the cloud server. Thus, the terminal device 100 can utilize a high-quality 3D model regardless of the performance of the terminal device 100.

[0151] Next, the in-photography display process (S102) will be described. Figure 15 It is a flowchart of the in-photography display process (S102). First, the UI unit 108 displays the UI screen (S601). Next, the UI unit 108 obtains and displays the in-photography image, that is, the captured image (S602). Next, the UI unit 108 determines whether a candidate photography position or an estimation failure signal has been received (S603). Here, the estimation failure signal is a signal sent from the position and attitude estimation unit 103 when the position and attitude estimation in the position and attitude estimation unit 103 fails. In addition, the candidate photography position is sent from the 3D reconstruction unit 104 or the point cloud analysis unit 106.

[0152] When the UI unit 108 receives a candidate for the photographing position (S603: Yes), it indicates that there is a candidate for the photographing position (S604) and presents the candidate for the photographing position (S605). For example, the UI unit 108 can visually present the candidate for the photographing position via the UI, or can present it using the sound emitted by a sound output mechanism such as a speaker. Specifically, it can also be that when the terminal device 100 is moved upward, "Please pick up by 20 cm" is indicated by sound, and when photographing the right side, "Please turn right by 45°" is indicated by sound. Thereby, the user can perform photographing without looking at the screen of the terminal device 100 during photographing accompanied by movement, and thus can perform photographing safely.

[0153] In addition, when the terminal device 100 is equipped with an oscillator such as a vibrator, it can also be presented by vibration. For example, a rule can be determined in advance such that there are two short vibrations when moving upward and one long vibration when turning right, and the presentation can be made in accordance with this rule. In this case, it is not necessary to look at the screen either, and thus safe photographing can be achieved.

[0154] In addition, when the UI unit 108 receives an estimation failure signal, in S604, it indicates that the estimation has failed.

[0155] After S605, or when a candidate for the photographing position is not received (S603: No), next, the UI unit 108 determines whether there is an instruction to end photographing (S606). The instruction to end photographing can be given, for example, by an operation on the UI screen or by a sound instruction. Or, it can also be given by a gesture input such as shaking the terminal device 100 twice.

[0156] When there is an instruction to end photographing (S606: Yes), the UI unit 108 sends a photographing end signal for conveying the end of photographing to the imaging unit 101 (S607). In addition, when there is no instruction to end photographing (S606: No), the UI unit 108 performs the processing after S601 again.

[0157] Figure 16This is a diagram showing an example of a method for visually prompting a candidate shooting position. Here, an example is shown where the candidate shooting position is a position higher than the current position, and it is specified to shoot from a position higher than the current position. In this example, an upward arrow 211 is prompted on the screen. In addition, the UI unit 108 can also change the display mode (color, size, etc.) of the arrow according to the distance from the current position to the candidate shooting position. For example, the UI unit 108 can display a red and large arrow when the current position is far from the candidate shooting position, and the arrow becomes greener and smaller as the current position gets closer to the candidate shooting position. Additionally, when there is no candidate shooting position (that is, when there is no area with high restoration difficulty in the current shooting), the UI unit 108 does not display an arrow, or it can prompt an ○ mark to indicate that the current shooting is good.

[0158] Figure 17 This is a diagram showing another example of a method for visually prompting a candidate shooting position. Here, the UI unit 108 displays a dotted line frame 212 closer to the center of the screen when the camera has moved towards the candidate shooting position, and prompts the user to move the camera to make the dotted line frame 212 closer to the center of the screen. In addition, the distance from the current position to the candidate shooting position can be represented by the color or thickness of the frame.

[0159] In addition, when the user does not follow the instruction to make the camera position closer to the prompted candidate shooting position, the UI unit 108 can also display a message 213 such as an alarm as shown Figure 18 In addition, the UI unit 108 can also switch the indication method according to the situation. For example, the UI unit 108 can display a smaller size at the beginning of the indication and make the display larger after a certain time. Or, the UI unit 108 can issue an alarm based on the time since the start of the indication. For example, the UI unit 108 can issue an alarm when the instruction has not been followed 1 minute after the start of the indication.

[0160] In addition, when the current position is separated from the candidate shooting position, the UI unit 108 can prompt general information and display it within a frame when it becomes a short distance that can be represented within the screen.

[0161] In addition, the display method of the candidate shooting position can also be other than the exemplified method. For example, the UI unit 108 can display the candidate shooting position on map information (two-dimensional or three-dimensional). Thereby, the user can intuitively grasp the moving direction.

[0162] In addition, examples of providing instructions to a user are described herein, but instructions may also be provided to a moving body equipped with a camera, such as a robot or a drone. In this case, the functions of the terminal device 100 may also be included in the moving body. That is, the moving body may also move to a candidate for a determined photographing position and perform photographing. As a result, a highly accurate three-dimensional model can be stably generated even in an automatically controlled device.

[0163] In addition, when performing three-dimensional reconstruction within the terminal device 100 or on a server, the information of pixels determined to have a high difficulty level of restoration may also be utilized. For example, the terminal device 100 or the server may determine the three-dimensional points reconstructed by using the pixels determined to have a high difficulty level of restoration as areas or points with low accuracy. In addition, metadata indicating such areas or points with low accuracy may be assigned to the three-dimensional model or the three-dimensional points. As a result, it is possible to determine whether the generated three-dimensional points are of high accuracy or low accuracy in post-processing. For example, the correction degree of the three-dimensional points in the filtering process can be switched according to the accuracy.

[0164] As described above, the photographing instruction device according to the present embodiment performs the processing as Figure 19 shown. The photographing instruction device (for example, the terminal device 100) detects a region (second region) in which it is difficult to generate a three-dimensional model of the subject using a plurality of images, based on the photographing position and posture of each of the plurality of images obtained by photographing the subject and the plurality of images (S701). Next, the photographing instruction device instructs at least one of the photographing position and the posture to photograph an image in which it is easy to generate a three-dimensional model of the detected region (S702). As a result, the accuracy of the three-dimensional model can be improved.

[0165] Here, an image in which it is easy to generate a three-dimensional model includes at least one of (1) an image of an unphotographed region when there is an unphotographed region among a part of the photographing viewpoints out of the plurality of photographing viewpoints; (2) an image of a region with little jitter; (3) an image of a region with a higher contrast than other regions and thus having many feature points; (4) an image of a region closer to the photographing viewpoint than other regions and in which the error between the calculated three-dimensional position and the actual position is estimated to be small when calculating the three-dimensional position; and (5) an image of a region with a smaller influence of lens distortion than other regions.

[0166] For example, the photographing instruction device further accepts the designation of a priority region (first region), and in the instruction (S702), instructs at least one of the photographing position and the posture to photograph an image in which it is easy to generate a three-dimensional model of the designated priority region. As a result, the accuracy of the three-dimensional model of the region required by the user can be preferentially improved.

[0167] For example, in order to generate a three-dimensional model of a subject based on the photographing positions and postures of a plurality of images obtained by photographing the subject and the plurality of images, the photographing instruction device according to the present embodiment accepts the designation of a first region (e.g., a priority region) ( Figure 3 in S205), and instructs at least one of the photographing position and the posture to photograph an image for generating the three-dimensional model of the designated first region. Thereby, it is possible to preferentially improve the accuracy of the three-dimensional model of the region required by the user, and thus it is possible to improve the accuracy of the three-dimensional model.

[0168] For example, based on the respective photographing positions and postures and the plurality of images, a second region in which it is difficult to generate the three-dimensional model is detected (S701), and at least one of the photographing position and the posture is instructed to generate an image that is easy to generate the three-dimensional model of the second region (S702). In the instruction (S702) corresponding to the first region, at least one of the photographing position and the posture is instructed to photograph an image that is easy to generate the three-dimensional model of the first region.

[0169] For example, as Figure 5 etc. show, the photographing instruction device displays an image of the subject after the identification of the attributes has been performed, and in the designation of the first region, the designation of the attributes is accepted.

[0170] For example, in the detection of the second region, (i) on a two-dimensional image, an edge whose angle difference from the epipolar line based on the photographing position and posture is smaller than a predetermined value is obtained, and (ii) a three-dimensional region corresponding to the obtained edge is detected as the second region. The photographing instruction device instructs at least one of the photographing position and the posture to photograph an image in which the angle difference is larger than the value in the instruction (S702) corresponding to the second region.

[0171] For example, the plurality of images are a plurality of frames included in a moving image being currently photographed and displayed, and the instruction (S702) corresponding to the second region is performed in real time. Thereby, by performing the photographing instruction in real time, the convenience for the user can be improved.

[0172] For example, in the instruction (S702) corresponding to the second region, for example, as Figure 16 shows, the photographing direction is instructed. Thereby, the user can easily perform appropriate photographing according to the instruction. For example, the direction in which the next photographing position exists with respect to the current position is prompted.

[0173] For example, in the instruction (S702) corresponding to the second region, for example, as Figure 17 shows, the photographing region is instructed. Thereby, the user can easily perform appropriate photographing according to the instruction.

[0174] For example, the photographing instruction device includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0175] (Embodiment 2)

[0176] By using a plurality of images photographed by a camera to generate a three-dimensional model, it is possible to generate a three-dimensional map or the like more simply than the method of generating a three-dimensional model using laser measurement. Here, the three-dimensional model is a model that represents the measured object photographed on a computer. The three-dimensional model has, for example, position information of each three-dimensional position on the measured object.

[0177] However, when photographing an image for generating a three-dimensional model of the space (hereinafter referred to as the object space) of the measurement object, it is not easy for the user (photographer) to judge the required image. As a result, the three-dimensional model cannot be reconstructed because appropriate images are not obtained, or the accuracy of the three-dimensional model may be reduced. Here, the accuracy is the error between the position information of the three-dimensional model and the actual position. In the present embodiment, information for assisting photographing is presented to the user during photographing. Thus, the user can effectively photograph appropriate images. In addition, the accuracy of the generated three-dimensional model can be improved.

[0178] Specifically, in the present embodiment, during photographing of the object space, an area where photographing has not been performed is detected and presented to the user (photographer). Here, the area where photographing has not been performed may include: an area where photographing has not been performed at that moment when photographing the object space (for example, an area blocked by other objects), and an area where, although photographing has been performed, three-dimensional points could not be obtained. In addition, an area where three-dimensional reconstruction (generation of a three-dimensional model) is difficult is detected and presented to the user. In addition, thereby, the efficiency of photographing can be improved, and the failure of three-dimensional model reconstruction or the reduction of accuracy can be suppressed.

[0179] First, a configuration example of the three-dimensional reconstruction system according to the present embodiment will be described. Figure 20 It is a diagram showing the configuration of the three-dimensional reconstruction system according to the present embodiment. As Figure 20 shown, the three-dimensional reconstruction system includes a photographing device 301 and a reconstruction device 302.

[0180] The imaging device 301 is a terminal device used by a user, such as a tablet terminal, a smart phone, or a notebook personal computer, etc., which is a portable terminal. The imaging device 301 has functions such as a photographing function, a function of estimating the position and orientation (hereinafter referred to as the position and orientation) of the camera, and a function of displaying a photographed completion area. In addition, during and after photographing, the imaging device 301 transmits the photographed image and the position and orientation to the reconstruction device 302. Here, the image is, for example, a moving image. In addition, the image may also be a plurality of still images. In addition, the imaging device 301 estimates the position and orientation during photographing, determines the photographed completion area using at least one of the position and orientation and the three-dimensional point cloud, and presents the photographed completion area to the user.

[0181] The reconstruction device 302 is, for example, a server connected to the imaging device 301 via a network or the like. The reconstruction device 302 acquires the image photographed by the imaging device 301 and generates a three-dimensional model using the acquired image. For example, the reconstruction device 302 can use the camera position and orientation estimated by the imaging device 301, or can estimate the camera position based on the acquired image.

[0182] In addition, the transfer of data between the imaging device 301 and the reconstruction device 302 can be performed via an HDD (hard disk drive) or the like based on an offline method, or can be performed constantly via a network.

[0183] In addition, the three-dimensional model generated by the reconstruction device 302 can be a three-dimensional point cloud (point cloud) that densely restores the three-dimensional space, or can be a set of three-dimensional meshes. In addition, the three-dimensional point cloud generated by the imaging device 301 is a set of three-dimensional points obtained by sparsely three-dimensionally restoring characteristic points such as the corners of objects in the space. That is, the three-dimensional model (three-dimensional point cloud) generated by the imaging device 301 is a model with a lower spatial resolution than the three-dimensional model generated by the reconstruction device 302. In other words, the three-dimensional model (three-dimensional point cloud) generated by the imaging device 301 is a model simpler than the three-dimensional model generated by the reconstruction device 302. In addition, a simplified model is, for example, a model with less information, a model that is easy to generate, or a model with low accuracy. For example, the three-dimensional model generated by the imaging device 301 is a sparser three-dimensional point cloud than the three-dimensional model generated by the reconstruction device 302.

[0184] Next, the configuration of the imaging device 301 will be described. Figure 21 is a block diagram of the imaging device 301. The imaging device 301 includes an imaging unit 311, a position and orientation estimation unit 312, a position and orientation integration unit 313, a region detection unit 314, a region detection unit 314, a UI unit 315, a control unit 316, an image storage unit 317, a position and orientation storage unit 318, and a region information storage unit 319.

[0185] The imaging unit 311 is an imaging device such as a camera, which acquires an image (a moving image). In addition, although an example using a moving image will be mainly described below, a plurality of still images may be used instead of the moving image. The imaging unit 311 stores the acquired image in the image storage unit 317. In addition, the imaging unit 311 can capture a visible light image or a non-visible light image (e.g., an infrared image). In the case of using an infrared image, imaging can be performed even in a dark environment such as at night. In addition, the imaging unit 311 can be a monocular camera or can have multiple cameras like a stereo camera. By using a calibrated stereo camera, the accuracy of the three-dimensional position and attitude can be improved. In addition, the imaging unit 311 can also be a device such as an RGB-D sensor that can also capture a depth image. In this case, a depth image as three-dimensional information can be acquired, so that the accuracy of estimating the camera position and attitude can be improved. In addition, the depth image can be used as alignment information when synthesizing the three-dimensional attitude described later.

[0186] The control unit 316 controls the overall imaging process of the imaging device 301. The position and attitude estimation unit 312 estimates the three-dimensional position and attitude (position and attitude) of the camera that has captured the image by using the image stored in the image storage unit 317. In addition, the position and attitude estimation unit 312 stores the estimated position and attitude in the position and attitude storage unit 318. For example, the position and attitude estimation unit 312 uses image processing such as SLAM (Simultaneous Localization and Mapping) to estimate the position and attitude. Alternatively, the position and attitude estimation unit 312 can also calculate the position and attitude of the camera by using information obtained from various sensors (GPS or acceleration sensor) provided in the imaging device 301. In the former case, the position and attitude can be estimated based on the information from the imaging unit 311. In the latter case, image processing can be implemented with less processing.

[0187] When imaging is performed multiple times in one environment, the position and attitude synthesis unit 313 synthesizes the position and attitude of the camera estimated in each imaging and calculates the position and attitude that can be processed in the same space. Specifically, the position and attitude synthesis unit 313 uses the three-dimensional coordinate axes of the position and attitude obtained in the first imaging as the reference coordinate axes. Then, the position and attitude synthesis unit 313 transforms the coordinates of the position and attitude obtained in the second and subsequent imagings into the coordinates in the space of the reference coordinate axes.

[0188] The area detection unit 314 uses the images stored in the image storage unit 317 and the position and orientation stored in the position and orientation storage unit 318 to detect areas in the object space where three-dimensional reconstruction cannot be performed or areas where the accuracy of three-dimensional reconstruction is low. Areas in the object space where three-dimensional reconstruction cannot be performed are, for example, areas where the image has not been photographed. Areas where the accuracy of three-dimensional reconstruction is low are, for example, areas where the number of images taken of the area is small (less than a predetermined number). In addition, an area with low accuracy is an area where the error between the three-dimensional position information generated when the three-dimensional position information is generated and the implemented position is large. In addition, the area detection unit 314 stores the information of the detected area in the area information storage unit 319. In addition, the area detection unit 314 may also detect areas where three-dimensional reconstruction can be performed and determine areas other than those areas as areas where three-dimensional reconstruction cannot be performed.

[0189] In addition, the information stored in the area information storage unit 319 can be either two-dimensional information overlaid on the image or three-dimensional information such as three-dimensional coordinate information.

[0190] The UI unit 315 presents the captured images and the area information detected by the area detection unit 314 to the user. In addition, the UI unit 315 has an input function for the user to input a photography start instruction and a photography end instruction. For example, the UI unit 315 is a display with a touch panel.

[0191] Next, the operation of the imaging device 301 according to the present embodiment will be described. Figure 22 This is a flowchart showing the operation of the imaging device 301. In the imaging device 301, the start and stop of photography are performed in response to the user's instruction. Specifically, photography is started by pressing the photography start button on the UI. When a photography start instruction is input (S801: Yes), the imaging unit 311 starts photography. The captured image is stored in the image storage unit 317.

[0192] Next, the position and orientation estimation unit 312 calculates the position and orientation each time an image is added (S802). The calculated position and orientation are stored in the position and orientation storage unit 318. At this time, when there is also, for example, a three-dimensional point cloud generated by SLAM in addition to the position and orientation, the generated three-dimensional point cloud is also stored.

[0193] Next, the position and orientation synthesis unit 313 synthesizes the position and orientation (S803). Specifically, the position and orientation synthesis unit 313 determines whether the three-dimensional coordinate spaces of the position and orientation of the previously captured image and the position and orientation of the newly captured image can be synthesized by using the estimation results of the position and orientation and the image, and synthesizes them if they can be synthesized. That is, the position and orientation synthesis unit 313 transforms the coordinates of the position and orientation of the newly captured image into the coordinate system of the previous position and orientation. Thus, multiple positions and orientations are represented in one three-dimensional coordinate space. As a result, the data obtained in multiple captures can be used in common, and the accuracy of the position and orientation estimation can be improved.

[0194] Next, the region detection unit 314 detects the captured area, etc. (S804). Specifically, the region detection unit 314 uses the estimation results of the position and orientation and the image to generate three-dimensional position information (such as a three-dimensional point cloud, a three-dimensional model, or a depth image), and uses the generated three-dimensional position information to detect a region where three-dimensional reconstruction cannot be performed or a region where the accuracy of three-dimensional reconstruction is low. In addition, the region detection unit 314 saves the information of the detected region to the region information storage unit 319. Next, the UI unit 315 displays the information of the region obtained through the above processing (S805).

[0195] In addition, this series of processes is repeated until the photography ends (S806). For example, these processes are repeated each time one or more frames of images are acquired.

[0196] Figure 23 It is a flowchart of the position and orientation estimation process (S802). First, the position and orientation estimation unit 312 acquires an image from the image storage unit 317 (S811). Next, the position and orientation estimation unit 312 calculates the position and orientation of the camera in each image by using the acquired image (S812). For example, the position and orientation estimation unit 312 calculates the position and orientation by using image processing such as SLAM or SfM (Structure from Motion). In addition, when the imaging device 301 has sensors such as an IMU (inertial measurement unit), the position and orientation estimation unit 312 can also utilize the information obtained from these sensors to estimate the position and orientation.

[0197] In addition, the position and orientation estimation unit 312 can utilize the results obtained by prior calibration as camera parameters such as the focal length of the lens. Or, the position and orientation estimation unit 312 can calculate the camera parameters simultaneously with the estimation of the position and orientation.

[0198] Next, the position and orientation estimation unit 312 stores the calculated position and orientation information in the position and orientation storage unit 318 (S813). In addition, in the case where the calculation of the position and orientation information fails, information indicating the failure may also be stored in the position and orientation storage unit 318. Thereby, it is possible to know the location and time of the failure, as well as in what image the failure occurred, and these information can be utilized during re - photography and the like.

[0199] Figure 24 This is a flowchart of the position and orientation comprehensive processing (S803). First, the position and orientation comprehensive unit 313 acquires an image from the image storage unit 317 (S821). Next, the position and orientation comprehensive unit 313 acquires the current position and orientation (S822). Next, the position and orientation comprehensive unit 313 acquires at least one of the images and position and orientation of the photographed - completed paths other than the current photographing path (S823). In addition, the photographed - completed path can be generated, for example, based on the sequence information of the position and orientation of the camera obtained by SLAM. Further, the information of the photographed - completed path is stored, for example, in the position and orientation storage unit 318. Specifically, the result of SLAM is stored each time a photography attempt is made. In the N - th photography (current path), the three - dimensional coordinate axes of the N - th photography are integrated with the three - dimensional coordinate axes of the results of the 1st to N - 1th times (past paths). In addition, instead of the result of SLAM, the position information obtained by GPS or Bluetooth (registered trademark) may be used.

[0200] Next, the position and orientation comprehensive unit 313 determines whether integration is possible (S824). Specifically, the position and orientation comprehensive unit 313 determines whether the position and orientation and images of the photographed - completed path obtained are similar to the current position and orientation and images. If they are similar, it is determined that integration is possible; if not, it is determined that integration is not possible. More specifically, the position and orientation comprehensive unit 313 calculates a feature quantity representing the overall characteristics of the image based on each image, and determines whether the images are of similar viewpoints by comparing them. In addition, when the photographing device 301 has GPS and knows the absolute position of the photographing device 301, the position and orientation comprehensive unit 313 may also use this information to determine an image photographed at the same or a similar position as the current image.

[0201] In the case where integration is possible (S824: Yes), the position and orientation comprehensive unit 313 performs path integration processing (S825). Specifically, the position and orientation comprehensive unit 313 calculates the three - dimensional relative position between the current image and the reference image that photographed a region similar to this image. The position and orientation comprehensive unit 313 adds the calculated three - dimensional relative position to the coordinates of the reference image, thereby calculating the coordinates of the current image.

[0202] Hereinafter, a specific example of the path integration processing will be described. Figure 25It is a plan view showing a case of photography in an object space. Path C is the path of camera A where photography and position and attitude estimation are completed. In addition, in this figure, the case where camera A exists at a prescribed position on path C is shown. Further, path D is the path of camera B that is currently performing photography. At this time, the current camera B photographs an image with the same field of view as camera A. In addition, an example in which two images are obtained by different cameras is shown here, but two images photographed by the same camera at different times may also be used. Figure 26 It is a figure showing an example of an image and a comparison process in this case.

[0203] As Figure 26 shown, for each image, the position and attitude synthesis unit 313 extracts feature amounts such as ORB (Oriented FAST and Rotated BRIEF) feature amounts, and extracts the feature amount of the entire image based on their distribution or quantity, etc. For example, the position and attitude synthesis unit 313 clusters the feature amounts appearing in the image like a bag of words, and uses the histogram of each class as the feature amount.

[0204] The position and attitude synthesis unit 313 compares the feature amounts of the entire image between images, and in the case of images determined to show the same position, performs feature point matching between the images to calculate the relative three-dimensional position between each camera. In other words, the position and attitude synthesis unit 313 explores images with similar feature amounts of the entire image from multiple images on the path where photography is completed. The position and attitude synthesis unit 313 transforms the three-dimensional position of path D to the coordinate system of path C based on this relative position relationship. Thereby, multiple paths can be represented in one coordinate system. In this way, the position and attitude in multiple paths can be synthesized.

[0205] In addition, when the photographing device 301 has a sensor such as GPS that can detect an absolute position, the position and attitude synthesis unit 313 may also use the detection result for synthesis processing. For example, the position and attitude synthesis unit 313 may perform processing using the detection result of the sensor instead of performing the above-described processing using the image, or may use not only the image but also the detection result of the sensor. For example, the position and attitude synthesis unit 313 may utilize the information of GPS to screen the images to be compared. Specifically, the position and attitude synthesis unit 313 may set, as comparison objects, images with position and attitude within a range where the latitude and longitude are each within ±0.001 degrees with respect to the current camera position through GPS. Thereby, the processing amount can be reduced.

[0206] Figure 27It is a flowchart of the area detection process (S804). First, the area detection unit 314 acquires an image from the image storage unit 317 (S831). Next, the area detection unit 314 acquires the position and orientation from the position and orientation storage unit 318 (S832). In addition, the area detection unit 314 acquires from the position and orientation storage unit 318 a three-dimensional point cloud representing the three-dimensional positions of feature points generated by SLAM or the like.

[0207] Next, the area detection unit 314 uses the acquired image, position and orientation, and three-dimensional point cloud to detect unphotographed areas (S833). Specifically, the area detection unit 314 projects the three-dimensional point cloud onto the image and determines the periphery of the pixels onto which the three-dimensional points are projected (within a predetermined distance from the pixels) as recoverable areas (photography completed areas). In addition, the farther the projected three-dimensional points are from the photography position, the larger the above-mentioned predetermined distance can be made. In addition, in the case of using a stereo camera, the area detection unit 314 can also estimate the recoverable area based on the disparity image. In addition, in the case of using an RGB-D camera, the obtained depth value can also be used for determination. In addition, the details of the determination process for the photography completed area will be described later. In addition, the area detection unit 314 can also determine not only whether a three-dimensional model can be generated, but also the accuracy in the case of three-dimensional reconstruction.

[0208] Finally, the area detection unit 314 outputs the obtained area information (S834). This area information can be, for example, an image obtained by overlapping the information of each area on the photographed image, or information obtained by arranging the information of each area in a three-dimensional space such as a three-dimensional map.

[0209] Figure 28 It is a flowchart of the display process (S805). First, the UI unit 315 confirms whether there is information to be displayed (S841). Specifically, the UI unit 315 confirms whether there is a newly added image in the image storage unit 317, and if so, determines that there is display information. In addition, the UI unit 315 confirms whether there is newly added information in the area information storage unit 319, and if so, determines that there is display information.

[0210] In the case of having display information (S841: Yes), the UI unit 315 acquires display information such as an image and area information (S842). Next, the UI unit 315 displays the acquired display information (S843).

[0211] Figure 29 It is a diagram showing an example of the UI screen displayed by the UI unit 315. This UI screen includes a photography start / stop button 321, a photographed image 322, area information 323, and a character display area 324.

[0212] The photographing start / stop button 321 is an operation unit for the user to indicate the start and stop of photographing. The in-photograph image 322 is the image being currently photographed. The area information 323 shows the photographed area, the low-precision area, etc. Here, the low-precision area means the area where the precision of the three-dimensional model generated when using the photographed image is low. In the character display area 324, information on the unphotographed area or the low-precision area is represented by characters. In addition, sound or the like can be used instead of characters.

[0213] Figure 30 FIG. is an example showing the area information 323. As Figure 30 shown, for example, an image obtained by overlapping and representing information on each area on the image being photographed is used. In addition, as the photographing viewpoint moves, the display of the area information 323 is changed in real time. Thereby, the user can easily grasp the photographed area, etc. while referring to the image at the current camera viewpoint. For example, the photographed area and the low-precision area are displayed in different colors. In addition, the area photographed in the current path and the area photographed in the photographing of other paths are displayed in different colors. In addition, as the information representing each area is Figure 30 shown, it can be either the information overlapped on the image or represented by characters or marks. That is, this information only needs to be information that the user can visually judge each area. Thereby, the user knows that it is only necessary to photograph the area without added color and the low-precision area, so that it is possible to avoid missing photographing, etc. In addition, in order to perform path integration processing, the area to be photographed can be presented to the user in such a way that the photographed area in the current path is continuous with the photographed areas of other paths.

[0214] In addition, for areas such as the low-precision area that the user is desired to pay attention to, a display that is easy to attract the user's attention, such as blinking, can also be used. In addition, the image to which the area information is overlapped can also be a past image.

[0215] In addition, here, the in-photograph image 322 and the area information 323 are displayed individually, but the area information can also be overlapped on the in-photograph image 322. In this case, the area required for display can be reduced, so that it is easy to visually recognize even on a small terminal such as a smart phone. Thereby, these display methods can also be switched according to the type of the terminal. For example, on a terminal with a small screen size such as a smart phone, the area information can be overlapped on the in-photograph image 322, and on a terminal with a large screen size such as a tablet terminal, the in-photograph image 322 and the area information 323 can be displayed individually.

[0216] In addition, the UI unit 315 uses characters or sounds to prompt that a low-precision area has been generated several meters in front of the current position, etc. This distance can be calculated based on the result of position estimation. By using character information, information can be correctly notified to the user. In addition, when using sound, there is no need to look away during photography, etc., so notification can be safely performed.

[0217] In addition, instead of two-dimensionally displaying on an image, information on the overlapping area can be presented in the real space through AR (Augmented Reality) glasses or HUD (Head-Up Display). Thereby, the affinity with the image observed from the real viewpoint is improved, and the position to be photographed can be intuitively presented to the user.

[0218] In addition, when the estimation of the position and orientation fails, the photographing device 301 can also notify the user of the failure of the estimation of the position and orientation. Figure 31 This is a diagram showing a display example in this case.

[0219] For example, the image at the time of failure of position and orientation estimation can also be displayed in the area information 323. In addition, characters or sounds for prompting to restart photography from this location can also be presented. Thereby, the user can quickly restart at the time of failure.

[0220] In addition, the photographing device 301 can also predict the failure position based on the elapsed time and moving speed since the failure, and use sound or characters to prompt information such as "Please return 5 meters back". Alternatively, the photographing device 301 can also display a two-dimensional map or a three-dimensional point map, and prompt the photographing failure position on this map.

[0221] In addition, the photographing device 301 can also detect that the user has returned to the failure position, and prompt this situation to the user using characters, images, sounds, or vibrations, etc. For example, it is possible to detect that the user has returned to the failure position using the feature amount of the entire image, etc.

[0222] In addition, when the photographing device 301 detects a low-precision area, it can also instruct the user to re-photograph and the photographing method. Here, the photographing method is, for example, to make the area appear larger. For example, the photographing device 301 can also perform this instruction using characters, images, sounds, or vibrations, etc.

[0223] Thereby, the quality of the acquired data can be improved, and thus the accuracy of the generated three-dimensional model can be improved.

[0224] Figure 32 This is a diagram showing a display example in the case of detecting a low-precision area. For example, as Figure 32As shown, characters are used to indicate to the user. In this example, in the character display area 324, it is shown that a low-precision area has been detected. In addition, instructions for the user such as "Please return 5 m ahead" for photographing the low-precision area are displayed. In addition, the photographing device 301 can also stop these displays when it detects that the user has moved to the indicated location or has photographed the indicated area.

[0225] Figure 33 is a diagram showing another example of an instruction to the user. For example, it is also possible to Figure 33 show the direction and distance of movement to the user by the arrow 325 shown. Figure 34 is a diagram showing an example of this arrow. For example, as Figure 34 shown, the direction of movement is indicated by the angle of the arrow, and the distance to the movement destination is shown by the size of the arrow. In addition, the display mode of the arrow can also be changed according to the distance. For example, the display mode can be either color, or the presence or absence or size of an effect. The effect is, for example, blinking, moving, or expanding and contracting, etc. In addition, the shade of the arrow can also be changed. For example, the closer the distance, the greater the effect can be made, or the darker the color of the arrow can be made. In addition, a combination of the above can also be used.

[0226] In addition, not limited to the arrow, as long as it is a display that can identify the direction, any display such as a triangle or a finger icon can also be used.

[0227] In addition, as the area information 323, instead of overlapping it on the photographed image, the information of the area can also be overlapped on an image based on a third-party viewpoint such as a floor plan or an oblique view. For example, if it is an environment where a 3D map can be prepared, the photographed completion area, etc. can also be overlapped and displayed on the 3D map.

[0228] Figure 35 is a diagram showing an example of the area information 323 in this case. Figure 35 is a diagram showing the 3D map as a floor plan, showing the photographed completion area and the low-precision area. In addition, it shows the current camera position 331 (the position and direction of the photographing device 301), the current path 332, and the past path 333.

[0229] In addition, in the case where there is a map such as CAD of the object space at a construction site, etc., or in the case where 3D modeling has been performed at the same location before and 3D map information has already been generated, the photographing device 301 can also make use of it. In addition, when the photographing device 301 can use GPS, etc., it can also create map information based on the latitude and longitude information obtained by GPS.

[0230] In this way, by overlapping the estimated results of the position and orientation of the camera and the photographed completed area on the 3D map and displaying it as an aerial view, the user can easily grasp which area has been photographed and which path has been taken for photography, etc. For example, in Figure 35 In the example shown, it is possible to make the user notice omissions in photography such as the area on the back side of the pillar that has not been photographed, which is difficult to understand in the prompt of the photographed image, so that effective photography can be achieved.

[0231] In addition, in an environment without CAD or the like, the display from a third-party viewpoint can be used in the same way without using a map. In this case, although the visual recognition ability decreases because there is no reference object, the user can grasp the positional relationship between the current photography position and the low-precision area. In addition, the user can confirm the situation where the area where an object is expected cannot be photographed. Thus, the efficiency of photography can also be improved in this case.

[0232] In addition, an example of a plan view is shown here, but map information observed from other viewpoints can also be used. In addition, the photographing device 301 may also have a function of changing the viewpoint of the 3D map. For example, a UI for the user to perform a viewpoint change operation can also be used.

[0233] In addition, the photographing device 301 may display both the area information 323 shown in the above Figure 29 etc. and the area information shown in Figure 35 or may have a function of switching these displays.

[0234] Next, an example of a method for determining and presenting the photographed completed area will be described. When SLAM is used in the position estimation of the camera, not only the position and orientation information of the camera is generated, but also a 3D point cloud related to feature points such as the corners of objects in the image is generated. Regarding the area where the 3D points are generated, it can be determined that 3D modeling can be performed. Therefore, by projecting the 3D points onto the image, the photographed completed area can be presented on the image.

[0235] Figure 36 is a plan view showing the photographing status of the object area. The black dots in the figure represent the generated 3D points (feature points). Figure 37 is a diagram showing an example of projecting 3D points into an image and determining the area around the 3D points as the photographed completed area.

[0236] In addition, the photographing device 301 may also connect the 3D points to generate a mesh and use the mesh to determine the photographed completed area. Figure 38 is a diagram showing an example of the photographed completed area in this case. As Figure 38 shown, the photographing device 301 may also determine the area where the mesh is generated as the photographed completed area.

[0237] In addition, when the imaging device 301 uses a stereo camera or an RGB-D sensor, it can also use the parallax image or depth image obtained from them to determine the captured completion area. Thereby, dense three-dimensional information can be obtained with less processing, and thus the captured completion area can be determined more accurately.

[0238] In addition, in this example, SLAM is used for self-position estimation (position and attitude estimation), but it is not limited to this as long as the position and attitude of the camera can be estimated during imaging.

[0239] In addition, the imaging device 301 can not only indicate the captured completion area, but also anticipate the accuracy when being restored and display the anticipated accuracy. Specifically, the three-dimensional points are calculated based on the feature points in the image. When projecting these three-dimensional points onto the image, the projected three-dimensional points sometimes deviate from the reference feature points. This deviation is called the reprojection error, and the accuracy can be evaluated using the reprojection error. Specifically, the larger the reprojection error, the lower the accuracy can be judged.

[0240] For example, the imaging device 301 can also represent the accuracy by the color of the captured completion area. For example, high-accuracy areas are represented by blue, and low-accuracy areas are represented by red. In addition, the accuracy can also be represented stepwise by the color difference or shade. Thereby, the user can easily grasp the accuracy of each area.

[0241] In addition, the imaging device 301 can also use the depth image obtained from an RGB-D sensor or the like to perform the determination of whether restoration is possible and the accuracy evaluation. Figure 39 This is a diagram showing an example of a depth image. In this diagram, the darker the color (the denser the shadow), the farther the distance means.

[0242] Here, there is a tendency that the closer the distance to the camera, the higher the accuracy of the generated three-dimensional model, and the farther the distance, the lower the accuracy. Thus, for example, the imaging device 301 can also determine the pixels up to a certain range of depth (for example, up to 5 m) as the captured range. Figure 40 This is a diagram showing an example when the area up to a certain range of depth is determined as the captured completion area. In this diagram, the shaded area is determined as the captured completion area.

[0243] In addition, the imaging device 301 can also determine the accuracy of the area according to the distance to the area. That is, the imaging device 301 can determine that the closer the distance, the higher the accuracy. For example, the relationship between the distance and the accuracy can be defined linearly, or other definitions can be used.

[0244] In addition, when the imaging device 301 uses a stereo camera, it can also generate a depth image from a disparity image. Therefore, the method for determining the captured area and accuracy can be the same as that in the case of using a depth image. In this case, the imaging device 301 can also determine that a region for which a depth value cannot be calculated from the disparity image cannot be restored (not captured). Alternatively, for a region for which a depth value cannot be calculated, the imaging device 301 can estimate the depth value based on surrounding pixels. For example, the imaging device 301 calculates the average value of 5×5 pixels centered on the target pixel.

[0245] In addition, information indicating the three-dimensional positions of the captured areas determined as described above and the accuracy of each three-dimensional position is stored. As described above, the coordinates of the position and orientation of the camera are integrated, so the coordinates of the obtained area information can also be integrated. Thereby, information indicating the three-dimensional positions of the captured areas in the target space and the accuracy of each three-dimensional position is generated.

[0246] As described above, the imaging instruction device according to the present embodiment performs Figure 41 the processing shown. The imaging device 301 captures a plurality of first images of the target space (S851), generates first three-dimensional position information of the target space (e.g., a sparse three-dimensional point cloud or a depth image) based on the plurality of first images and the first capture position and orientation of each of the plurality of first images (S852), and for a second region in which it is difficult to generate second three-dimensional position information (e.g., a dense three-dimensional point cloud) of the target space that is more detailed than the first three-dimensional position information by using the plurality of first images, the second three-dimensional position information is not generated and the first three-dimensional position information is used for determination (S853).

[0247] Here, the region in which it is difficult to generate three-dimensional position information includes at least one of a region where a three-dimensional position cannot be calculated and a region where the error between the three-dimensional position and the actual position is larger than a specified threshold. Specifically, the second region includes (1) a region that is not captured from some of the plurality of imaging viewpoints; (2) a region that is captured but has a large shake; (3) a region that is captured but has a lower contrast than other regions and thus has fewer feature points; (4) a region that is captured but is farther from the imaging viewpoint than other regions, and even if a three-dimensional position is calculated, the error between the calculated three-dimensional position and the actual position is large; and (5) a region where the influence of lens distortion is larger than that of other regions. In addition, shake can be detected, for example, by obtaining the temporal position change of feature points.

[0248] In addition, the detailed three-dimensional position information is, for example, three-dimensional position information with high spatial resolution. The spatial resolution of the three-dimensional position information represents the distance between two adjacent three-dimensional positions when the two adjacent three-dimensional positions can be discriminated as different three-dimensional positions. In addition, a high spatial resolution means that the distance between two adjacent three-dimensional positions is small. That is, the three-dimensional position information with high spatial resolution has more three-dimensional position information in a space of a specified size. In addition, the three-dimensional position information with high spatial resolution can also be referred to as dense three-dimensional position information, while the three-dimensional position information with low spatial resolution can be referred to as sparse three-dimensional position information.

[0249] In addition, the detailed three-dimensional position information can also be three-dimensional position information with a large amount of information. For example, the first three-dimensional position information can be distance information from one viewpoint like a depth image, and the second three-dimensional position information can be a three-dimensional model such as a three-dimensional point group capable of obtaining distance information from an arbitrary viewpoint.

[0250] In addition, for example, the object space and the subject are the same concept, and commonly mean the photographed area.

[0251] Thus, the imaging device 301 can use the first three-dimensional position information to determine the second area where it is difficult to generate the second three-dimensional position information without generating the second three-dimensional position information, and thus can improve the efficiency of photographing multiple images for generating the second three-dimensional position information.

[0252] For example, the second area is at least one of an area where imaging of the image has not been performed and an area where the accuracy of the second three-dimensional position information is estimated to be lower than a predetermined reference. Here, the reference is, for example, a threshold of the distance between two different three-dimensional positions. That is, the second area is an area where the difference between the generated three-dimensional position information and the implemented position is larger than a predetermined threshold when the three-dimensional position information is generated.

[0253] For example, the first three-dimensional position information includes a first three-dimensional point group, and the second three-dimensional position information includes a second three-dimensional point group that is denser than the first three-dimensional point group.

[0254] For example, in the determination (S853), the imaging device 301 determines the third area of the object space corresponding to the area around the first three-dimensional point group (the area within a specified distance from the first three-dimensional point group), and determines the area other than the third area as the second area (for example Figure 37 ).

[0255] For example, in the determination (S853), the imaging device 301 generates a mesh using the first three-dimensional point group, and determines the area other than the third area of the object space corresponding to the area where the mesh is generated as the second area (for example Figure 38 ).

[0256] For example, in the determination (S853), the imaging device 301 determines the second region based on the reprojection error of the first three-dimensional point group.

[0257] For example, the first three-dimensional position information includes a depth image. The imaging device 301, for example, as Figure 40 shown, based on the depth image, determines the region within a predetermined distance from the imaging viewpoint as the third region, and determines the region outside the third region as the second region.

[0258] For example, the imaging device 301 further uses the plurality of second images that have been captured, the second capture positions and postures of the plurality of second images, the plurality of first images, and the plurality of first capture positions and postures, and makes the coordinate systems of the plurality of first capture positions and postures correspond to the coordinate systems of the plurality of second capture positions and postures. Thereby, the imaging device 301 can use the information obtained in multiple captures to determine the second region.

[0259] For example, the imaging device 301 further displays the second region or the third region (e.g., the captured completion region) outside the second region during the capture in the object space (e.g., Figure 30 ). Thereby, the imaging device 301 can prompt the second region to the user.

[0260] For example, the imaging device 301 overlays and displays the information indicating the second region or the third region on one of the plurality of images (e.g., Figure 30 ). Thereby, the imaging device 301 can prompt the position of the second region within the image to the user, so that the user can easily grasp the position of the second region.

[0261] For example, the imaging device 301 overlays and displays the information indicating the second region or the third region on the map of the object space (e.g., Figure 35 ). Thereby, the imaging device 301 can prompt the position of the second region in the surrounding environment to the user, so that the user can easily grasp the position of the second region.

[0262] For example, the imaging device 301 displays the second region and the restoration accuracy (the accuracy of three-dimensional reconstruction) of each region included in the second region. Thereby, the user can not only grasp the second region but also the restoration accuracy of each region, and thus can perform appropriate imaging accordingly.

[0263] For example, the imaging device 301 further prompts the user with an instruction for the user to image the second region (e.g., Figure 32 and Figure 33 ). Thereby, the user can effectively perform appropriate imaging.

[0264] For example, the indication includes at least one of the direction and distance from the current position to the second area (e.g., Figure 33 and Figure 34 ). Thereby, the user can effectively perform appropriate photography.

[0265] For example, the photographing device includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0266] As described above, the photographing instruction device and the photographing device according to the embodiment of the present disclosure have been described, but the present disclosure is not limited to this embodiment.

[0267] For example, the photographing instruction device according to Embodiment 1 may be combined with the photographing device according to Embodiment 2. For example, the above photographing instruction device may also have at least a part of the processing units included in the above photographing device. In addition, the above photographing device may also have at least a part of the processing units included in the above photographing instruction device. In addition, at least a part of the processing units included in the above photographing instruction device may be combined with at least a part of the processing units included in the above photographing device. That is to say, the photographing instruction method according to Embodiment 1 may also include at least a part of the processing included in the photographing method according to Embodiment 2. In addition, the photographing method according to Embodiment 2 may also include at least a part of the processing included in the photographing instruction method according to Embodiment 1. In addition, at least a part of the processing included in the photographing instruction method according to Embodiment 1 may be combined with at least a part of the processing included in the photographing method according to Embodiment 2.

[0268] In addition, each processing unit included in the photographing instruction device and the photographing device according to the above embodiment is typically implemented as an LSI which is an integrated circuit. These may be individually a single chip, or may be a single chip in a manner including a part or all of them.

[0269] In addition, the formation of the integrated circuit is not limited to LSI, and may also be implemented by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after manufacturing the LSI, or a reconfigurable processor that can reconfigure the connection and setting of circuit units inside the reconfigurable LSI may also be used.

[0270] In addition, in each of the above embodiments, each component may also be constituted by dedicated hardware, or may be implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.

[0271] In addition, the present disclosure can also be implemented as a photographing instruction method or the like executed by a photographing instruction device or the like.

[0272] In addition, the division of the functional modules in the block diagram is an example. Multiple functional modules can also be implemented as one functional module, or one functional module can be divided into multiple ones, or a part of the function can be transferred to other functional modules. In addition, the functions of multiple functional modules with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.

[0273] In addition, the order in which each step in the flowchart is executed is an order illustrated for specifically explaining the present disclosure, and it can also be an order other than the above. In addition, a part of the above steps can also be executed simultaneously (in parallel) with other steps.

[0274] The photographing instruction device and the like related to one or more aspects have been described above based on the embodiments, but the present disclosure is not limited to these embodiments. As long as it does not deviate from the gist of the present disclosure, the embodiments obtained by applying various modifications conceivable by those skilled in the art to this embodiment, and the embodiments constructed by combining the constituent elements in different embodiments can also be included within the scope of one or more aspects.

[0275] Industrial Applicability

[0276] The present disclosure can be applied to a photographing instruction device.

[0277] Description of Reference Numerals:

[0278] 100 Terminal device

[0279] 101 Imaging unit

[0280] 102 Control unit

[0281] 103 Position and attitude estimation unit

[0282] 104 3D reconstruction unit

[0283] 105 Image analysis unit

[0284] 106 Point cloud analysis unit

[0285] 107 Communication unit

[0286] 108 UI unit

[0287] 111 Image storage unit

[0288] 112 Camera attitude storage unit

[0289] 113 3D model storage unit

[0290] 301 Photographing device

[0291] 302 Reconstruction Device

[0292] 311 Camera Unit

[0293] 312 Position and Orientation Estimation Unit

[0294] 313 Position and Orientation Integration Unit

[0295] 314 Region Detection Unit

[0296] 315 UI Unit

[0297] 316 Control Unit

[0298] 317 Image Storage Unit

[0299] 318 Position and Orientation Storage Unit

[0300] 319 Region Information Storage Unit

[0301] 321 Photograph Start / Stop Button

[0302] 322 In-Progress Photograph Image

[0303] 323 Region Information

[0304] 324 Character Display Area

[0305] 325 Arrow

[0306] 331 Camera Position

[0307] 332 Current Path

[0308] 333 Past Path

Claims

1. A photography method, performed by a photography device, a plurality of first images of a photographed object space, generate first three-dimensional position information of the object space based on the plurality of first images, and the first photographing positions and postures of the plurality of first images respectively, for a second area in which it is difficult to generate second three-dimensional position information of the object space that is more detailed than the first three-dimensional position information, instead of generating the second three-dimensional position information, use the first three-dimensional position information for determination, the first three-dimensional position information includes a first three-dimensional point cloud, the second three-dimensional position information includes a second three-dimensional point cloud that is denser than the first three-dimensional point cloud, in the determination, use the first three-dimensional point cloud to generate a mesh, and determine an area outside the third area of the object space corresponding to the area where the mesh is generated as the second area, The second region includes: at least one of an area where the second three-dimensional point cloud cannot be calculated and an area where the error between the position of the second three-dimensional point cloud and the actual position is greater than a specified threshold.

2. The photography method according to claim 1, the photography method further uses a plurality of second images that have been photographed, the second photographing positions and postures of the plurality of second images respectively, the plurality of first images, and the plurality of first photographing positions and postures to make the coordinate systems of the plurality of first photographing positions and postures correspond to the coordinate systems of the plurality of second photographing positions and postures.

3. The photography method according to claim 1, the photography method further displays the second area or the third area during the photography of the object space.

4. The photography method according to claim 3, display information indicating the second area or the third area by overlapping it on one of the plurality of first images.

5. The photography method according to claim 3, on a map of the object space, overlap and display information indicating the second area or the third area.

6. The photography method according to any one of claims 3 to 5, display the second area and the restoration accuracy of each area included in the second area.

7. A photography device, comprising: a processor; and a memory, the processor uses the memory, a plurality of first images of a photographed object space, generate first three-dimensional position information of the object space based on the plurality of first images, and the first photographing positions and postures of the plurality of first images respectively, for a second area in which it is difficult to generate second three-dimensional position information of the object space that is more detailed than the first three-dimensional position information, instead of generating the second three-dimensional position information, use the first three-dimensional position information for determination, the first three-dimensional position information includes a first three-dimensional point cloud, the second three-dimensional position information includes a second three-dimensional point cloud that is denser than the first three-dimensional point cloud, in the determination, use the first three-dimensional point cloud to generate a mesh, and determine an area outside the third area of the object space corresponding to the area where the mesh is generated as the second area, The second region includes: It is not possible to calculate at least one of the region of the second three-dimensional point group and the region where the error between the position of the second three-dimensional point group and the actual position is greater than a specified threshold value.

Citation Information

Patent Citations

  • Image management apparatus, image management method and program

    JP2017130146A

  • Contour line measurement apparatus and robot system

    CN105444691A

  • Smart guide to capture digital images that align with a target image model

    US20190253614A1