Photographing method and photographing device

The photography instruction method and device improve three-dimensional model accuracy by designating specific areas for imaging and providing real-time shooting guidance, addressing challenges in optimal positioning and posture determination.

JP7745167B2Active Publication Date: 2025-09-29PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022512007
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-27
Filing Date
2021-03-24
Publication Date
2025-09-29
Estimated Expiration
2041-03-24

AI Technical Summary

Technical Problem

Existing methods for generating three-dimensional models from multiple images face challenges in achieving accurate reconstruction, particularly in determining optimal shooting positions and postures to improve model accuracy.

Method used

A photography instruction method and device that designates specific areas for three-dimensional model generation, providing shooting position and posture instructions to enhance accuracy, especially in regions difficult to model, and displays real-time guidance for improved image capture.

Benefits of technology

Enhances the accuracy of three-dimensional models by prioritizing critical areas for imaging, reducing the need for re-shooting and improving work efficiency through precise shooting instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007745167000002
    Figure 0007745167000002
  • Figure 0007745167000003
    Figure 0007745167000003
  • Figure 0007745167000004
    Figure 0007745167000004
Patent Text Reader

Abstract

This imaging instruction method is executed by an imaging instruction device, wherein in order to generate a three-dimensional model of a subject on the basis of imaging positions and orientations of a plurality of images in which the subject has been imaged, and the plurality of images, the designation of a first region is received (S205) and at least one among an imaging position and orientation is instructed so as to capture an image to be used in the generation of the three-dimensional model of the designated first region. For example, this imaging instruction method may detect, on the basis of the respective imaging positions and orientations and the plurality of images, a second region for which it is difficult to generate a three-dimensional model, may instruct at least one among the imaging position and orientation so that an image is generated with which it is easy to generate a three-dimensional model of the second region, and, in an instruction corresponding to the first region, may instruct at least one among the imaging position and orientation so that an image is captured with which it is easy to generate a three-dimensional model of the first region.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a photography instruction method, a photography method, a photography instruction device, and a photography device. [Background technology]

[0002] Patent Document 1 discloses a technique for generating a three-dimensional model of a subject using a plurality of images obtained by photographing the subject from a plurality of viewpoints. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-130146 Summary of the Invention [Problem to be solved by the invention]

[0004] In a process for generating a three-dimensional model, it is desirable to be able to improve the accuracy of the three-dimensional model. An object of the present disclosure is to provide a photography instruction method or a photography instruction device that can improve the accuracy of the three-dimensional model. [Means for solving the problem]

[0005] An imaging method according to one aspect of the present disclosure is an imaging method executed by an imaging device, the imaging method comprising: capturing a plurality of first images of a target space; generating first three-dimensional position information of the target space based on the plurality of first images and a first imaging position and orientation of each of the plurality of first images; and generating second three-dimensional position information of the target space that is more detailed than the first three-dimensional position information. 2nd area using the first three-dimensional position information without generating the second three-dimensional position information, the first three-dimensional position information including a first three-dimensional point cloud, and the second three-dimensional position information including a second three-dimensional point cloud that is denser than the first three-dimensional point cloud, and in the determination, a mesh is generated using the first three-dimensional point cloud, and a region of the object space corresponding to the region in which the mesh is generated is determined. Third areaThe area other than the above 2nd area It is determined that:

[0006] According to one aspect of the present disclosure How to give instructions for shooting teeth, A photography instruction method executed by a photography instruction device, which accepts designation of a first area for generating a three-dimensional model of a subject based on the shooting position and posture of each of a plurality of images of the subject and the plurality of images, instructs at least one of the shooting position and posture to capture an image to be used in generating the three-dimensional model of the specified first area, displays an image of the subject for which attribute recognition has been performed, and accepts designation of the attributes when designating the first area. [Effects of the Invention]

[0007] The present disclosure can provide a photography instruction method or photography instruction device that can improve the accuracy of a three-dimensional model. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram of a terminal device according to the first embodiment. [Figure 2] FIG. 2 is a sequence diagram of the terminal device according to the first embodiment. [Figure 3] FIG. 3 is a flowchart of the initial processing according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing an example of an initial display according to the first embodiment. [Figure 5] FIG. 5 is a diagram showing an example of a method for selecting a priority designated portion according to the first embodiment. [Figure 6] FIG. 6 is a diagram showing an example of a method for selecting a priority designated portion according to the first embodiment. [Figure 7] FIG. 7 is a flowchart of the position and orientation estimation process according to the first embodiment. [Figure 8] FIG. 8 is a flowchart of the shooting position candidate determination process according to the first embodiment. [Figure 9] FIG. 9 is a diagram showing a situation in which the camera and the object according to the first embodiment are viewed from above. [Figure 10] FIG. 10 is a diagram showing an example of an image obtained by each camera according to the first embodiment. [Figure 11] FIG. 11 is a schematic diagram for explaining an example of determining candidate shooting positions according to the first embodiment. [Figure 12] FIG. 12 is a schematic diagram for explaining an example of determining candidate shooting positions according to the first embodiment. [Figure 13] FIG. 13 is a schematic diagram for explaining an example of determining candidate shooting positions according to the first embodiment. [Figure 14] FIG. 14 is a flowchart of the three-dimensional reconstruction process according to the first embodiment. [Figure 15] FIG. 15 is a flowchart of the display process during shooting according to the first embodiment. [Figure 16] FIG. 16 is a diagram showing an example of a method for visually presenting shooting position candidates according to the first embodiment. [Figure 17] FIG. 17 is a diagram showing an example of a method for visually presenting shooting position candidates according to the first embodiment. [Figure 18] FIG. 18 is a diagram showing an example of an alert display according to the first embodiment. [Figure 19] FIG. 19 is a flowchart of the shooting instruction process according to the first embodiment. [Figure 20] FIG. 20 is a diagram illustrating a configuration of a three-dimensional reconstruction system according to the second embodiment. [Figure 21] FIG. 21 is a block diagram of an imaging device according to the second embodiment. [Figure 22] FIG. 22 is a flowchart showing the operation of the imaging device according to the second embodiment. [Figure 23] FIG. 23 is a flowchart of a position and orientation estimation process according to the second embodiment. [Figure 24] FIG. 24 is a flowchart of the position and orientation integration process according to the second embodiment. [Figure 25] FIG. 25 is a plan view showing how an image is captured in a target space according to the second embodiment. [Figure 26] FIG. 26 shows example images and an example of comparison processing according to the second embodiment. [Figure 27] FIG. 27 is a flowchart of the area detection process according to the second embodiment. [Figure 28] FIG. 28 is a flowchart of the display process according to the second embodiment. [Figure 29]FIG. 29 is a diagram showing a display example of a UI screen according to the second embodiment. [Figure 30] FIG. 30 is a diagram illustrating an example of region information according to the second embodiment. [Figure 31] FIG. 31 is a diagram showing a display example when position and orientation estimation according to the second embodiment fails. [Figure 32] FIG. 32 is a diagram showing a display example when a low-accuracy area according to the second embodiment is detected. [Figure 33] FIG. 33 is a diagram showing an example of an instruction to a user according to the second embodiment. [Figure 34] FIG. 34 is a diagram showing an example of an instruction (arrow) according to the second embodiment. [Figure 35] FIG. 35 is a diagram illustrating an example of region information according to the second embodiment. [Figure 36] FIG. 36 is a plan view showing an imaging situation of a target area according to the second embodiment. [Figure 37] FIG. 37 is a diagram showing an example of an imaged area when three-dimensional points according to the second embodiment are used. [Figure 38] FIG. 38 is a diagram showing an example of an imaged area when a mesh according to the second embodiment is used. [Figure 39] FIG. 39 is a diagram illustrating an example of a depth image according to the second embodiment. [Figure 40] FIG. 40 is a diagram showing an example of an imaged region when a depth image according to the second embodiment is used. [Figure 41] FIG. 41 is a flowchart of the imaging method according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] A photographing instruction method according to one aspect of the present disclosure is a photographing instruction method executed by a photographing instruction device, which accepts designation of a first area for generating a three-dimensional model of a subject based on the photographing position and posture of each of a plurality of images of the subject and the plurality of images, and instructs at least one of the photographing position and posture to photograph an image to be used in generating the three-dimensional model of the specified first area.

[0010] This allows the accuracy of the three-dimensional model to be improved preferentially in the area required by the user, thereby improving the accuracy of the three-dimensional model.

[0011] For example, based on each of the shooting positions and postures and the multiple images, a second region for which it is difficult to generate a three-dimensional model is detected, and at least one of the shooting position and posture is instructed so as to generate an image that makes it easier to generate a three-dimensional model of the second region, and in the instruction corresponding to the first region, at least one of the shooting position and posture is instructed so as to capture an image that makes it easier to generate the three-dimensional model of the first region.

[0012] For example, an image of the subject for which attribute recognition has been executed is displayed, and the attribute is designated in the first area.

[0013] For example, in detecting the second region, (i) an edge is found on the two-dimensional image whose angular difference with the epipolar line based on the shooting position and posture is smaller than a predetermined value, and (ii) a three-dimensional region corresponding to the found edge is detected as the second region, and the instruction corresponding to the second region may instruct at least one of the shooting position and posture to capture an image in which the angular difference is larger than the value.

[0014] For example, the plurality of images may be a plurality of frames included in a moving image currently being captured and displayed, and the instruction corresponding to the second area may be given in real time.

[0015] This allows for real-time shooting instructions, improving user convenience.

[0016] For example, the instruction corresponding to the second area may instruct a shooting direction.

[0017] This allows the user to easily take a suitable photograph by following the instructions.

[0018] For example, the instruction corresponding to the second area may instruct a shooting area.

[0019] This allows the user to easily take a suitable photograph by following the instructions.

[0020] In addition, a photography instruction device according to one aspect of the present disclosure includes a processor and a memory, and accepts designation of a first area for generating a three-dimensional model of a subject based on the shooting position and posture of each of a plurality of images of the subject and the plurality of images, and instructs at least one of the shooting position and posture to capture an image to be used in generating the three-dimensional model of the specified first area.

[0021] This allows the accuracy of the three-dimensional model to be improved preferentially in the area required by the user, thereby improving the accuracy of the three-dimensional model.

[0022] A photography instruction method according to one aspect of the present disclosure detects an area where it is difficult to generate a three-dimensional model of the subject using multiple images based on the shooting position and posture of each of multiple images of the subject and the multiple images, and instructs at least one of the shooting position and posture so as to capture an image that makes it easier to generate a three-dimensional model of the detected area.

[0023] This allows for improved accuracy of the three-dimensional model.

[0024] For example, the photographing instruction method may further receive a specification of a priority area, and the instruction may specify at least one of the photographing position and posture so as to photograph an image that makes it easier to generate a three-dimensional model of the specified priority area.

[0025] This allows the accuracy of the three-dimensional model of the area required by the user to be improved with priority.

[0026] A photographing method according to one aspect of the present disclosure is a photographing method executed by a photographing device, which photographs a plurality of first images of a target space, generates first three-dimensional position information of the target space based on the plurality of first images and a first photographing position and orientation of each of the plurality of first images, and determines a second region of the target space for which it is difficult to generate second three-dimensional position information that is more detailed than the first three-dimensional position information, using the first three-dimensional position information without generating the second three-dimensional position information.

[0027] According to this, the photographing method can determine the second area for which it is difficult to generate second three-dimensional position information using the first three-dimensional position information without generating second three-dimensional position information, thereby improving the efficiency of photographing multiple images for generating second three-dimensional position information.

[0028] For example, the second area may be at least one of an area where no image has been captured and an area where the accuracy of the second three-dimensional position information is estimated to be lower than a predetermined standard.

[0029] For example, the first three-dimensional position information may include a first three-dimensional point cloud, and the second three-dimensional position information may include a second three-dimensional point cloud that is denser than the first three-dimensional point cloud.

[0030] For example, the determination may include determining a third region of the target space corresponding to a region surrounding the first three-dimensional point cloud, and determining a region other than the third region as the second region.

[0031] For example, in the determination, a mesh may be generated using the first three-dimensional point group, and an area other than a third area in the target space corresponding to an area in which the mesh is generated may be determined to be the second area.

[0032] For example, the determination may determine the second region based on a reprojection error of the first three-dimensional point cloud.

[0033] For example, the first three-dimensional position information may include a depth image, and based on the depth image, an area within a predetermined distance from the shooting viewpoint may be determined as a third area, and an area other than the third area may be determined as the second area.

[0034] For example, the photographing method may further include aligning a coordinate system of the plurality of first photographing positions and orientations with a coordinate system of the plurality of second photographing positions and orientations using a plurality of second images that have already been photographed, second photographing positions and orientations of each of the plurality of second images, the plurality of first images, and the plurality of first photographing positions and orientations.

[0035] This allows the second region to be determined using information obtained by capturing images multiple times.

[0036] For example, the imaging method may further include displaying the second area or a third area other than the second area while imaging the target space.

[0037] This allows the second area to be presented to the user.

[0038] For example, information indicating the second region or the third region may be displayed superimposed on one of the plurality of images.

[0039] This allows the position of the second area in the image to be presented to the user, allowing the user to easily grasp the position of the second area.

[0040] For example, information indicating the second area or the third area may be displayed superimposed on a map of the target space.

[0041] This allows the position of the second area in the surrounding environment to be presented to the user, allowing the user to easily grasp the position of the second area.

[0042] For example, the second region and the restoration accuracy of each region included in the second region may be displayed.

[0043] This allows the user to grasp the restoration accuracy of each area in addition to the second area, and based on this, can perform appropriate photography.

[0044] For example, the photographing method may further include presenting to the user an instruction to cause the user to photograph the second area.

[0045] This allows the user to efficiently take appropriate photographs.

[0046] For example, the instruction may include at least one of a direction and a distance from the current location to the second area.

[0047] This allows the user to efficiently take appropriate photographs.

[0048] In addition, an imaging device according to one aspect of the present disclosure includes a processor and a memory, and the processor uses the memory to capture a plurality of first images of a target space, generate first three-dimensional position information of the target space based on the plurality of first images and a first imaging position and orientation of each of the plurality of first images, and determine a second area of ​​the target space for which it is difficult to generate second three-dimensional position information of the target space that is more detailed than the first three-dimensional position information using the plurality of first images, using the first three-dimensional position information without generating the second three-dimensional position information.

[0049] This allows the photographing device to use the first three-dimensional position information to determine the second area for which it is difficult to generate second three-dimensional position information without generating second three-dimensional position information, thereby improving the efficiency of photographing multiple images for generating second three-dimensional position information.

[0050] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0051] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components not recited in independent claims will be described as optional components.

[0052] (Embodiment 1) Generating a three-dimensional model using multiple images captured by a camera makes it easier to generate three-dimensional maps and the like than using laser measurement to generate three-dimensional models. For this reason, the method of generating three-dimensional models using images is used when measuring distances in construction management at building sites, etc. Here, a three-dimensional model is a computer representation of an imaged measurement target. The three-dimensional model contains, for example, positional information for each three-dimensional location on the measurement target.

[0053] However, when generating a 3D model using images, multiple images with parallax that capture the same location are required. Furthermore, the more images there are, the more densely the 3D model can be generated. Furthermore, because parallax affects the accuracy of the restoration, when measuring distances, it is best to take photos while moving appropriately relative to the subject. However, it is difficult to determine the appropriate shooting position while taking photos while responding to the subject's situation.

[0054] In this embodiment, a UI (user interface) or the like is described that detects areas where 3D reconstruction (generation of a 3D model) is difficult and instructs the shooting position or shooting posture based on the detection results. This makes it possible to suppress failure or deterioration in accuracy of reconstruction of a 3D model. In addition, since the need for re-shooting can be reduced, work efficiency can be improved.

[0055] First, the configuration of a terminal device 100, which is an example of a photography instruction device according to this embodiment, will be described. FIG. 1 is a block diagram of the terminal device 100 according to this embodiment. This terminal device 100 has an imaging function, a function of estimating the three-dimensional position and orientation during photography, a function of determining next photography position candidates from the photographed image, and a function of presenting the estimated photography position candidates to a user. Note that the terminal device 100 may also have a function of performing three-dimensional reconstruction using the estimated three-dimensional position and orientation to generate a three-dimensional model that is a three-dimensional point cloud of the photography environment, a function of determining next photography position candidates using the three-dimensional model, a function of presenting the estimated photography position candidates to a user, and a function of transmitting and receiving at least one of the photographed video, the three-dimensional position and orientation, and the three-dimensional model to and from another terminal device, a management server, etc.

[0056] The terminal device 100 includes an imaging unit 101, a control unit 102, a position and orientation estimation unit 103, a three-dimensional reconstruction unit 104, an image analysis unit 105, a point cloud analysis unit 106, a communication unit 107, a UI unit 108, a video storage unit 111, a camera orientation storage unit 112, and a three-dimensional model storage unit 113.

[0057] The imaging unit 101 is an imaging device such as a camera, and acquires video (moving images). Note that, although an example in which video is mainly used will be described below, multiple still images may be used instead of video. The imaging unit 101 stores the acquired video in the video storage unit 111. Note that the imaging unit 101 may capture visible light images or infrared images. When infrared images are used, imaging is possible even in dark environments such as at night. Note that the imaging unit 101 may be a monocular camera, or may have multiple cameras like a stereo camera. The accuracy of three-dimensional position and orientation can be improved by using a calibrated stereo camera. Even if the stereo camera is not calibrated, parallax images with parallax can be acquired.

[0058] The control unit 102 performs overall control of the image capturing process and the like of the terminal device 100. The position and orientation estimation unit 103 estimates the three-dimensional position and orientation of the camera that captured the image using the image stored in the image storage unit 111. The position and orientation estimation unit 103 also stores the estimated three-dimensional position and orientation in the camera orientation storage unit 112. For example, the position and orientation estimation unit 103 uses image processing such as SLAM (Simultaneous Localization and Mapping) to estimate the position and orientation of the camera. Alternatively, the position and orientation estimation unit 103 may calculate the position and orientation of the camera using information obtained by various sensors (GPS or acceleration sensor) included in the terminal device 100. In the former case, the position and orientation can be estimated from information from the image capturing unit 101. In the latter case, image processing can be performed with low processing power.

[0059] The 3D reconstruction unit 104 generates a 3D model by performing 3D reconstruction using the images stored in the image storage unit 111 and the 3D positions and orientations stored in the camera orientation storage unit 112. The 3D reconstruction unit 104 also stores the generated 3D model in the 3D model storage unit 113. For example, the 3D reconstruction unit 104 performs 3D reconstruction using image processing such as SfM (Structure from Motion). Alternatively, the 3D reconstruction unit 104 may utilize stereo parallax when using images obtained by a calibrated camera such as a stereo camera. The former method makes it possible to generate a highly accurate 3D model by using many images. The latter method makes it possible to generate a 3D model quickly with light processing.

[0060] When SLAM is used in the position and orientation estimation unit 103, a three-dimensional model of the surrounding environment is generated at the same time as the three-dimensional position and orientation are estimated. Therefore, this three-dimensional model may be utilized.

[0061] The image analysis unit 105 analyzes the position from which imaging should be performed in order to perform accurate 3D reconstruction, using the images stored in the image storage unit 111 and the 3D positions and orientations stored in the camera orientation storage unit 112, and determines candidate imaging positions based on the analysis results. Information indicating the determined candidate imaging positions is output to the UI unit 108 and presented to the user.

[0062] The point cloud analysis unit 106 determines the density of the point cloud included in the three-dimensional model using the video stored in the video storage unit 111, the three-dimensional position and orientation stored in the camera orientation storage unit 112, and the three-dimensional model in the three-dimensional model storage unit 113. The point cloud analysis unit 106 determines candidate shooting positions that capture sparse areas. The point cloud analysis unit 106 also detects areas of the point cloud generated using peripheral areas of the image where lens distortion and the like are likely to occur, and determines candidate shooting positions that capture those areas at the center of the camera. The determined candidate shooting positions are output to the UI unit 108 and presented to the user.

[0063] The communication unit 107 transmits and receives the captured video, the calculated three-dimensional posture, and the three-dimensional model to and from a cloud server or another terminal device via communication.

[0064] The UI unit 108 presents the captured video and the candidate shooting positions determined by the image analysis unit 105 and the point cloud analysis unit 106 to the user. The UI unit 108 also has an input function for the user to input instructions to start shooting, instructions to end shooting, and priority processing locations.

[0065] Next, the operation of the terminal device 100 according to this embodiment will be described. FIG. 2 is a sequence diagram illustrating the exchange of information within the terminal device 100. In FIG. 2, areas indicated by light gray indicate areas where the imaging unit 101 is continuously capturing images. To generate a high-quality three-dimensional model, the terminal device 100 analyzes the video and the capturing position and orientation in real time and instructs the user to capture the image. Here, real-time analysis refers to performing analysis while capturing images. Alternatively, real-time analysis refers to performing analysis without generating a three-dimensional model. Specifically, the terminal device 100 estimates the position and orientation of the camera during capture and, based on the estimation result and the captured video, determines areas that are difficult to restore. The terminal device 100 predicts the capturing position and orientation that will ensure a parallax that makes it easy to restore the area and presents the predicted capturing position and orientation on a UI. While a sequence for capturing video is presented here, similar processing may be performed for capturing individual still images.

[0066] First, the UI unit 108 performs an initial process (S101). As a result, the UI unit 108 sends a shooting start signal to the imaging unit 101. The start process is performed, for example, when the user clicks a "shooting start" button on the display of the terminal device 100.

[0067] Next, the UI unit 108 performs display processing during shooting (S102). Specifically, the UI unit 108 presents the video being shot and instructions to the user.

[0068] Upon receiving the shooting start signal, the image capturing unit 101 captures video and transmits the captured video, which is image information, to the position and orientation estimation unit 103, the 3D reconstruction unit 104, the image analysis unit 105, and the point cloud analysis unit 106. For example, the image capturing unit 101 may perform streaming transmission, in which the video is transmitted as needed at the same time as shooting, or may transmit video at regular intervals all at once. In other words, the image information is one or more images (frames) contained in the video. In the former case, processing can be performed as needed, thereby reducing the waiting time for generating a 3D model. In the latter case, a large amount of captured information can be utilized, thereby achieving high-precision processing.

[0069] The position and orientation estimation unit 103 first performs input standby processing at the start of image capture, and enters a state of waiting for image information from the imaging unit 101. When image information is input from the imaging unit 101, the position and orientation estimation unit 103 performs position and orientation estimation processing (S103). That is, the position and orientation estimation processing is performed for each frame or for each set of frames. If the position and orientation estimation processing fails, the position and orientation estimation unit 103 transmits an estimation failure signal to the UI unit 108 to notify the user of the failure. If the position and orientation estimation processing is successful, the position and orientation estimation unit 103 transmits position and orientation information, which is the estimation result of the three-dimensional position and orientation, to the UI unit 108 to output the current three-dimensional position and orientation. Furthermore, the position and orientation estimation unit 103 transmits the position and orientation information to the image analysis unit 105 and the three-dimensional reconstruction unit 104.

[0070] The image analysis unit 105 first performs an input standby process at the start of shooting, and waits for image information from the imaging unit 101 and position and orientation information from the position and orientation estimation unit 103. When the image information and position and orientation information are input, the image analysis unit 105 performs a shooting position candidate determination process (S104). Note that the shooting position candidate determination process may be performed for each frame, or may be performed every certain time (a plurality of frames) (e.g., every 5 seconds). The image analysis unit 105 also determines whether the terminal device 100 is moving to a shooting position candidate generated by the shooting position candidate determination process, and if it is moving, it does not need to perform a new shooting position candidate determination process. For example, the image analysis unit 105 determines that it is moving if the current position and orientation are on a line connecting the position and orientation of the image when the shooting position candidate was determined and the position and orientation of the calculated candidate.

[0071] The three-dimensional reconstruction unit 104 first performs input standby processing at the start of imaging, and enters a state of waiting for image information from the imaging unit 101 and position and orientation information from the position and orientation estimation unit 103. When the image information and position and orientation information are input, the three-dimensional reconstruction unit 104 performs three-dimensional reconstruction processing (S105) to calculate a three-dimensional model. The three-dimensional reconstruction unit 104 transmits point cloud information, which is the calculated three-dimensional model, to the point cloud analysis unit 106.

[0072] The point cloud analysis unit 106 first performs an input standby process at the start of imaging, and waits for point cloud information from the three-dimensional reconstruction unit 104. When the point cloud information is input, the point cloud analysis unit 106 performs a photographing position candidate determination process (S106). For example, the point cloud analysis unit 106 determines the density state of the entire point cloud and detects sparse areas. The point cloud analysis unit 106 determines a photographing position candidate that captures as many of these sparse areas as possible. Note that the point cloud analysis unit 106 may determine the photographing position candidate using image information or position and orientation information in addition to the point cloud information.

[0073] Next, the initial processing (S101) will be described. FIG. 3 is a flowchart of the initial processing (S101). First, the UI unit 108 displays the currently captured image (S201). Next, the UI unit 108 acquires whether or not there is a priority portion that the user wants to restore preferentially (S202). For example, the UI unit 108 displays a button for specifying a priority mode, and if the button is selected, it determines that there is a priority portion.

[0074] If there is a priority location (Yes in S203), the UI unit 108 displays a priority location selection screen (S204) and acquires information about the priority location selected by the user (S205). After step S205, or if there is no priority location (No in S203), the UI unit 108 then outputs a shooting start signal to the imaging unit 101. This starts shooting (S206). For example, shooting may start when the user presses a button, or shooting may start automatically after a certain period of time has passed.

[0075] Here, the set priority areas are given a high priority for restoration when issuing instructions to move the camera, etc. This prevents instructions to move the camera to restore areas that are difficult to restore but are not necessary for the user, and allows instructions to be given to target areas that the user needs.

[0076] 4 is a diagram showing an example of the initial display of the initial process (S101). A captured image 201, a priority designation button 202 for selecting whether to select a priority designation location, and a start shooting button 203 for starting shooting are displayed. Note that the captured image 201 may be a still image or may be a video (moving image) currently being shot.

[0077] 5 is a diagram showing an example of a method for selecting a priority designation location when designating a priority. In this diagram, attribute recognition of an object such as a window frame is performed, and a desired object is selected from a list of multiple objects contained in the image using a selection box 204. For example, labels such as window frame, desk, and wall are added to each pixel of the image using a method such as semantic segmentation, and target pixels are selected all at once by selecting the label.

[0078] 6 is a diagram showing an example of another method for selecting a priority designation area when designating a priority. In this figure, the priority area is selected by the user designating an arbitrary area (a rectangular area in the figure) using a pointer or touch operation. Note that any method may be used as long as the user can select a specific area. For example, the selection may be made by designating a color, or a rough designation such as the right-hand area in the image may be made.

[0079] Furthermore, input methods other than on-screen operation may be used. For example, selection operations may be performed by voice input. In this case, inputting is easier when it is difficult to operate by hand, such as when wearing gloves in cold climates.

[0080] Although the example shown here shows the selection of priority locations in the initial processing before shooting, priority locations may be added as appropriate during shooting, which makes it possible to select locations that were not initially captured.

[0081] Next, the position and orientation estimation process (S103) will be described. Fig. 7 is a flowchart of the position and orientation estimation process (S103). First, the position and orientation estimation unit 103 acquires one or more images or videos from the video storage unit 111 (S301). Next, the position and orientation estimation unit 103 calculates or acquires position and orientation information for each input image, including camera parameters such as the three-dimensional position and orientation (direction) of the camera and lens information (S302). For example, the position and orientation estimation unit 103 calculates the position and orientation information by performing image processing such as SLAM or SfM on the images acquired in S301.

[0082] Next, the position and orientation estimation unit 103 stores the position and orientation information acquired in S302 in the camera orientation storage unit 112 (S303).

[0083] The images input in S301 may be an image sequence consisting of multiple frames spanning a fixed period of time, and the processing from S302 onwards may be performed on this image sequence (multiple images). Alternatively, images may be input sequentially, as in streaming, and the processing from S302 onwards may be repeated for each input image. In the former case, accuracy can be improved by using information from multiple time points. In the latter case, sequential input is possible, ensuring a fixed length of input delay and reducing the waiting time required for generating a 3D model.

[0084] Next, the imaging position candidate determination process (S104) by the image analysis unit 105 will be described. FIG. 8 is a flowchart of the imaging position candidate determination process (S104). First, the image analysis unit 105 acquires a plurality of images or videos from the video storage unit 111 (S401). One of the acquired plurality of images is set as a key image. Here, the key image is an image that serves as a reference when performing three-dimensional reconstruction in the subsequent stage. For example, the depth of each pixel of the key image is estimated using information of images other than the key image, and three-dimensional reconstruction is performed using the estimated depth. Next, the image analysis unit 105 acquires position and orientation information of each of the plurality of images (S402).

[0085] Next, the image analysis unit 105 calculates an epipolar line between the images (between the cameras) using the position and orientation information of each image (S403). Next, the image analysis unit 105 detects edges in each image (S404). For example, the image analysis unit 105 detects edges by filtering using a Sobel filter or the like.

[0086] Next, the image analysis unit 105 calculates the angle between the epipolar line and the edge in each image (S405). Next, the image analysis unit 105 calculates the degree of difficulty of restoration for each pixel of the key image based on the angle obtained in S405 (S406). Specifically, the more parallel the epipolar line and the edge are, the more difficult it is to perform three-dimensional reconstruction, so the smaller the angle, the higher the degree of difficulty of restoration is set. Note that the degree of difficulty of restoration may be set in multiple stages, or in two stages: high and low. For example, if the angle is smaller than a predetermined value (e.g., 5 degrees), the degree of difficulty of restoration may be set to high, and if the angle is larger than the predetermined value, the degree of difficulty of restoration may be set to low.

[0087] Next, based on the restoration difficulty calculated in S406, the image analysis unit 105 estimates from which position the image should be captured to make it easier to restore the area with a high restoration difficulty, and determines the estimated position as a candidate image capturing position (S407). Specifically, an area with a high restoration difficulty is synonymous with being on the same plane as the movement direction between the cameras, and moving the camera in a direction perpendicular to that plane causes the epipolar line and the edge to become non-horizontal. Therefore, if the camera is moving forward, the restoration difficulty in the area with a high restoration difficulty can be reduced by moving the camera in the vertical or horizontal direction.

[0088] Although the restoration difficulty is calculated through the processes of S401 to S406, this is not necessarily true as long as the restoration difficulty can be calculated. For example, the degradation of image quality due to lens distortion is greater at the edges of the image than at the center of the image. Therefore, the image analysis unit 105 may determine an object that appears only at the edge of the screen in each image and set a high restoration difficulty for the area of ​​that object. For example, the image analysis unit 105 may determine the field of view of the camera in three-dimensional space from the position and orientation information, and determine the area that appears only at the edge of the screen from the overlap of the field of view of each camera.

[0089] Below, we will explain how the difficulty of restoration varies depending on the angle between the epipolar line and the edge. Fig. 9 is a diagram showing a situation in which the camera and the object are viewed from above. Fig. 10 is a diagram showing examples of images obtained by each camera in the situation shown in Fig. 9.

[0090] In this case, when searching for a corresponding point in the image from camera B for a point in the image from camera A, the epipolar line, which can be calculated from the camera geometry, is searched. When the epipolar line and the edge in the image of the object are parallel, as in the case of point A, matching using pixel information such as NCC (Normalized Cross Correlation) is difficult, and it is difficult to determine the correct corresponding point.

[0091] On the other hand, when the epipolar line and the edge on the image of the object are perpendicular, as in point B, matching becomes easier and the correct corresponding point can be determined. In other words, the difficulty of determining the corresponding point can be determined by calculating the angle between the epipolar line and the edge on the image. Correctly determining this corresponding point affects the accuracy of 3D reconstruction. Therefore, the angle between this epipolar line and the edge on the image can be used as an indication of the difficulty of 3D reconstruction. Note that any information with the same meaning as the angle will suffice; for example, the epipolar line and the edge can be considered as vectors and the dot product of the epipolar line and the edge can be used.

[0092] The epipolar line can be calculated using the fundamental matrix between camera A and camera B. This fundamental matrix can be calculated from the position and orientation information of camera A and camera B. If the internal matrices of camera A and camera B are KA and KB, and the relative rotation matrix of camera B as seen from camera A is R and the relative movement vector is T, the epipolar line can be calculated as follows:

[0093] For pixel (x, y) on camera A, (a, b, c) are calculated using the following formula, and the epipolar line on camera B can be expressed as a straight line that satisfies ax+by+c=0.

[0094]

number

[0095] An example of determining candidate shooting positions will be described below. Fig. 11 to Fig. 13 are schematic diagrams for explaining an example of determining candidate shooting positions. An edge with a high degree of difficulty in restoration is often a line or line segment on a three-dimensional plane that passes through a line connecting the three-dimensional positions of camera A and camera B, as shown in Fig. 11. Conversely, the more perpendicular the line is to the plane that passes through this line, the greater the angle between the epipolar line and the edge in matching between camera A and camera B, and the lower the degree of difficulty in restoration.

[0096] 12, when determining a candidate shooting position, if an image is taken from the position of camera C in a direction perpendicular to the plane containing the edge and meeting the above conditions with respect to the target edge, it becomes possible to lower the difficulty of restoration at the target edge when matching with camera A or camera B. Therefore, the image analysis unit 105 determines camera C as the candidate shooting position.

[0097] Furthermore, if there is no target edge, the image analysis unit 105 may determine the shooting position candidate by the above method, using the edge with the highest restoration difficulty as a candidate, or may randomly select an edge as a candidate from the top 10 edges in terms of restoration difficulty.

[0098] In addition, here, camera C is calculated using only information from a pair of cameras (camera A and camera B). However, if there are multiple edges with high restoration difficulty, as shown in FIG. 13, the image analysis unit 105 calculates a candidate shooting position (camera C) for the first edge and a candidate shooting position (camera B) for the second edge. D ) and output a route connecting camera C and camera D.

[0099] However, the method for determining candidate shooting positions is not limited to this. Taking into consideration that the center of the screen is less susceptible to image quality degradation due to distortion and the like, if an edge captured at the edge of the screen by camera A is targeted, the image analysis unit 105 may determine a position where the edge is captured in the center of the screen as the candidate shooting position.

[0100] Next, the three-dimensional reconstruction process (S105) will be described. Fig. 14 is a flowchart of the three-dimensional reconstruction process (S105). First, the three-dimensional reconstruction unit 104 acquires a plurality of images or videos from the video storage unit 111 (S501). Next, the three-dimensional reconstruction unit 104 acquires position and orientation information (camera parameters) of each of the plurality of images from the camera orientation storage unit 112 (S502).

[0101] Next, the 3D reconstruction unit 104 generates a 3D model by performing 3D reconstruction using the acquired multiple images and multiple pieces of position and orientation information (S503). For example, the 3D reconstruction unit 104 performs 3D reconstruction using volume intersection or SfM. Finally, the 3D reconstruction unit 104 stores the generated 3D model in the 3D model storage unit 113 (S504).

[0102] Note that the processing of S503 does not have to be performed by the terminal device 100. For example, the terminal device 100 transmits the image and camera parameters to a cloud server or the like. The cloud server generates a three-dimensional model by performing three-dimensional reconstruction. The terminal device 100 receives the three-dimensional model from the cloud server. This allows the terminal device 100 to use a high-quality three-dimensional model regardless of the performance of the terminal device 100.

[0103] Next, the display processing during shooting (S102) will be described. FIG. 15 is a flowchart of the display processing during shooting (S102). First, the UI unit 108 displays a UI screen (S601). Next, the UI unit 108 acquires and displays a captured image, which is an image being shot, (S602). Next, the UI unit 108 determines whether a shooting position candidate or an estimation failure signal has been received (S603). Here, the estimation failure signal is a signal transmitted from the position and orientation estimation unit 103 when position and orientation estimation in the position and orientation estimation unit 103 fails. Furthermore, the shooting position candidate is transmitted from the three-dimensional reconstruction unit 104 or the point cloud analysis unit 106.

[0104] When the UI unit 108 receives the shooting position candidates (Yes in S603), it displays a message indicating that there are shooting position candidates (S604) and presents the shooting position candidates (S605). For example, the UI unit 108 may visually display the shooting position candidates via the UI, or may present the shooting position candidates using audio from a mechanism that outputs audio, such as a speaker. Specifically, when moving the terminal device 100 upward, audio instructions may be given such as "Lift the terminal device 100 by 20 cm," or when photographing the right side, "Turn 45 degrees to the right." This allows the user to shoot safely because they do not need to stare at the screen of the terminal device 100 while shooting while moving.

[0105] Furthermore, if the terminal device 100 is equipped with a vibrator or other vibrator, the notification may be via vibration. For example, rules may be set in advance, such as two short vibrations when moving upwards and one long vibration when turning right, and the notification may be made in accordance with these rules. In this case, too, it is not necessary to keep your eyes on the screen, making it possible to achieve safe shooting.

[0106] When the UI unit 108 receives the estimation failure signal, it displays in S604 that the estimation has failed.

[0107] After S605, or if no shooting position candidates have been received (No in S603), the UI unit 108 then determines whether or not an instruction to end shooting has been given (S606). The instruction to end shooting may be given, for example, by operating the UI screen, or by voice instruction. Alternatively, the instruction may be given by gesture input, such as shaking the terminal device 100 twice.

[0108] If an instruction to end image capture has been given (Yes in S606), the UI unit 108 transmits an image capture end signal to the imaging unit 101 to notify the end of image capture (S607). If an instruction to end image capture has not been given (No in S606), the UI unit 108 performs the processes from S601 onwards again.

[0109] FIG. 16 is a diagram showing an example of a method for visually presenting candidate shooting positions. Here, an example is shown in which the candidate shooting position is located above the current position, and a designation is made to shoot from a position above the current position. In this example, an up arrow 211 is presented on the screen. The UI unit 108 may change the display format (color, size, etc.) of the arrow depending on the distance from the current position to the candidate shooting position. For example, the UI unit 108 may display a large red arrow when the current position is far from the candidate shooting position, and a smaller green arrow as the current position gets closer to the candidate shooting position. Furthermore, if there are no candidate shooting positions (i.e., there are no areas in the current shooting that are difficult to restore), the UI unit 108 may not display an arrow, or may present a circle symbol to indicate that the current shooting is good.

[0110] 17 is a diagram showing another example of a method for visually presenting candidate shooting positions. Here, the UI unit 108 displays a dotted frame 212 that will come to the center of the screen when the camera moves to the candidate shooting position, and prompts the user to move the camera so that the dotted frame 212 approaches the center of the screen. Note that the color or thickness of the frame may indicate the distance from the current position to the candidate shooting position.

[0111] Furthermore, if the user does not follow instructions to move the camera position closer to the presented shooting position candidate, the UI unit 108 may display a message 213 such as an alert as shown in FIG. 18 . The UI unit 108 may also switch the way of giving instructions depending on the situation. For example, the UI unit 108 may display a smaller image shortly after the start of the instructions and then enlarge the image after a certain period of time has passed. Alternatively, the UI unit 108 may issue an alert based on the time elapsed since the start of the instructions. For example, the UI unit 108 may issue an alert if the instructions are not followed one minute after the start of the instructions.

[0112] Furthermore, when the current position is far from the candidate shooting position, the UI unit 108 may present general information, and when the distance becomes close enough to be displayed on the screen, display the information in a frame.

[0113] Note that the method of displaying the shooting position candidates may be other than the exemplified method. For example, the UI unit 108 may display the shooting position candidates on map information (two-dimensional or three-dimensional). This allows the user to intuitively understand the direction in which to move.

[0114] Although the example of giving instructions to a user has been described above, instructions may also be given to a mobile object equipped with a camera, such as a robot or a drone. In this case, the functions of the terminal device 100 may be included in the mobile object. In other words, the mobile object may move to the determined candidate shooting position and take a photograph. This allows a highly accurate three-dimensional model to be generated stably even in automatically controlled equipment.

[0115] Furthermore, information about pixels determined to have a high degree of restoration difficulty may be utilized when performing 3D reconstruction within the terminal device 100 or the server. For example, the terminal device 100 or the server may determine that 3D points reconstructed using pixels determined to have a high degree of restoration difficulty are areas or points with low accuracy. Furthermore, metadata indicating these areas or points with low accuracy may be added to the 3D model or the 3D points. This allows post-processing to determine whether the generated 3D points are high or low accuracy. For example, the degree of correction in the filtering process of the 3D points can be switched depending on the accuracy.

[0116] As described above, the photography instruction device according to this embodiment performs the process shown in Fig. 19. The photography instruction device (for example, terminal device 100) detects an area (second area) where it is difficult to generate a three-dimensional model of the subject using multiple images, based on the photography position and orientation of each of multiple images of the subject and the multiple images (S701). Next, the photography instruction device instructs at least one of the photography position and orientation so as to capture images that make it easy to generate a three-dimensional model of the detected area (S702). This can improve the accuracy of the three-dimensional model.

[0117] Here, images that facilitate the generation of a three-dimensional model include at least one of: (1) an image of an area that is not photographed from some of the multiple photographing viewpoints; (2) an image of an area with little blur; (3) an image of an area that has higher contrast than other areas and therefore has more feature points; (4) an image of an area that is closer to the photographing viewpoint than other areas and is estimated to have a smaller error between the calculated three-dimensional position and the actual position when the three-dimensional position is calculated; and (5) an image of an area that is less affected by lens distortion than other areas.

[0118] For example, the photography instruction device further receives a designation of a priority area (first area), and in the instruction (S702), instructs at least one of a photographing position and a posture so as to photograph an image that makes it easy to generate a three-dimensional model of the designated priority area. This makes it possible to preferentially improve the accuracy of the three-dimensional model of the area required by the user.

[0119] For example, the photography instruction device according to this embodiment accepts designation of a first region (e.g., a priority region) for generating a three-dimensional model of a subject based on the photographing position and posture of each of a plurality of images of the subject and the plurality of images (S205 in FIG. 3), and instructs at least one of the photographing position and posture so as to photograph images to be used in generating a three-dimensional model of the specified first region. This allows the accuracy of the three-dimensional model of the region required by the user to be improved preferentially, thereby improving the accuracy of the three-dimensional model.

[0120] For example, based on each of the shooting positions and postures and the multiple images, a second region for which it is difficult to generate a three-dimensional model is detected (S701), and at least one of the shooting position and posture is instructed so as to generate an image that makes it easier to generate a three-dimensional model of the second region (S702), and in the instruction (S702) corresponding to the first region, at least one of the shooting position and posture is instructed so as to capture an image that makes it easier to generate the three-dimensional model of the first region.

[0121] For example, as shown in FIG. 5, the photography instruction device displays an image of the subject for which attribute recognition has been performed, and accepts designation of the attribute in designating the first area.

[0122] For example, in detecting the second region, (i) an edge on the two-dimensional image is found whose angular difference with respect to the epipolar line based on the photographing position and posture is smaller than a predetermined value, and (ii) a three-dimensional region corresponding to the found edge is detected as the second region. In the instruction (S702) corresponding to the second region, the photographing instruction device instructs at least one of the photographing position and posture so as to photograph an image in which the angular difference is larger than the value.

[0123] For example, the plurality of images are a plurality of frames included in a video currently being captured and displayed, and the instruction (S702) corresponding to the second area is given in real time. This improves user convenience by giving a capturing instruction in real time.

[0124] For example, the instruction (S702) corresponding to the second area instructs the shooting direction, for example, as shown in Fig. 16. This allows the user to easily take a suitable shot by following the instruction. For example, the direction in which the next shooting position exists relative to the current position is presented.

[0125] For example, the instruction (S702) corresponding to the second area instructs an imaging area as shown in Fig. 17. This allows the user to easily perform suitable imaging by following the instruction.

[0126] For example, the photography instruction device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0127] (Embodiment 2) By generating a three-dimensional model using multiple images captured by a camera, it is possible to generate a three-dimensional map or the like more easily than by using laser measurement to generate a three-dimensional model. Here, a three-dimensional model is a computer representation of the imaged measurement target. The three-dimensional model contains, for example, positional information for each three-dimensional location on the measurement target.

[0128] However, when capturing images for generating a three-dimensional model of a space to be measured (hereinafter referred to as the target space), it is not easy for the user (photographer) to determine the necessary images. As a result, the three-dimensional model may not be reconstructed because an appropriate image is not obtained, or the accuracy of the three-dimensional model may be reduced. Here, accuracy refers to the error between the position information of the three-dimensional model and the actual position. In this embodiment, information to assist the user in capturing images is presented to the user during capturing. This allows the user to efficiently capture appropriate images. Furthermore, the accuracy of the generated three-dimensional model can be improved.

[0129] Specifically, in this embodiment, when photographing the target space, an area where no photographing has been performed is detected, and the detected area is presented to the user (photographer). Here, the area where no photographing has been performed may include an area where no photographing has been performed at the time of photographing the target space (for example, an area hidden by another object) and an area where photographing has been performed but no 3D points have been obtained. In addition, an area where 3D reconstruction (generation of a 3D model) is difficult is detected, and the detected area is presented to the user. This also improves the efficiency of photographing and prevents failure or deterioration in the accuracy of reconstructing the 3D model.

[0130] First, a configuration example of a three-dimensional reconstruction system according to this embodiment will be described. Fig. 20 is a diagram showing the configuration of a three-dimensional reconstruction system according to this embodiment. As shown in Fig. 20, the three-dimensional reconstruction system includes an imaging device 301 and a reconstruction device 302.

[0131] The image capturing device 301 is a terminal device used by a user, such as a mobile terminal such as a tablet terminal, smartphone, or laptop personal computer. The image capturing device 301 has a photographing function, a function for estimating the position and orientation of the camera (hereinafter referred to as the position and orientation), and a function for displaying the photographed area. The image capturing device 301 also sends the photographed images and their position and orientation to the reconstruction device 302 during and after photographing. Here, the image is, for example, a moving image. Note that the image may also be a plurality of still images. The image capturing device 301 also estimates the position and orientation during photographing, determines the photographed area using at least one of the position and orientation and the three-dimensional point cloud, and presents the photographed area to the user.

[0132] The reconstruction device 302 is, for example, a server connected to the image capturing device 301 via a network or the like. The reconstruction device 302 acquires images captured by the image capturing device 301 and generates a three-dimensional model using the acquired images. For example, the reconstruction device 302 may use the camera position and orientation estimated by the image capturing device 301, or may estimate the camera position from the acquired images.

[0133] Furthermore, data may be exchanged between the imaging device 301 and the reconstruction device 302 offline via a hard disk drive (HDD) or constantly via a network.

[0134] The three-dimensional model generated by the reconstruction device 302 may be a three-dimensional point cloud (point cloud) that densely reconstructs a three-dimensional space, or may be a collection of three-dimensional meshes. The three-dimensional point cloud generated by the image capture device 301 is a collection of three-dimensional points that sparsely reconstruct characteristic points, such as corners of an object in space. In other words, the three-dimensional model (three-dimensional point cloud) generated by the image capture device 301 has a lower spatial resolution than the three-dimensional model generated by the reconstruction device 302. In other words, the three-dimensional model (three-dimensional point cloud) generated by the image capture device 301 is a simpler model than the three-dimensional model generated by the reconstruction device 302. A simpler model is, for example, a model with less information, a model that is easy to generate, or a model with lower accuracy. For example, the three-dimensional model generated by the image capture device 301 is a sparser three-dimensional point cloud than the three-dimensional model generated by the reconstruction device 302.

[0135] Next, we will explain the configuration of the image capturing device 301. Fig. 21 is a block diagram of the image capturing device 301. The image capturing device 301 includes an image capturing unit 311, a position and orientation estimation unit 312, a position and orientation integration unit 313, a region detection unit 314, a UI unit 315, a control unit 316, an image storage unit 317, a position and orientation storage unit 318, and a region information storage unit 319.

[0136] The imaging unit 311 is an imaging device such as a camera, and acquires images (moving images). While the following description primarily focuses on examples in which moving images are used, multiple still images may be used instead of moving images. The imaging unit 311 stores the acquired images in the image storage unit 317. The imaging unit 311 may capture visible light images or non-visible light images (e.g., infrared images). Using infrared images enables imaging even in dark environments, such as at night. The imaging unit 311 may be a monocular camera or may have multiple cameras, such as a stereo camera. Using a calibrated stereo camera can improve the accuracy of the three-dimensional position and orientation. The imaging unit 311 may also be a device capable of capturing depth images, such as an RGB-D sensor. In this case, depth images, which serve as three-dimensional information, can be acquired, improving the accuracy of camera position and orientation estimation. Furthermore, the depth images can be used as information for alignment during three-dimensional orientation integration, as described below.

[0137] The control unit 316 performs overall control of the image capture process and the like of the image capture device 301. The position and orientation estimation unit 312 estimates the three-dimensional position and orientation (position and orientation) of the camera that captured the image using images stored in the image storage unit 317. The position and orientation estimation unit 312 also stores the estimated position and orientation in the position and orientation storage unit 318. For example, the position and orientation estimation unit 312 uses image processing such as SLAM (Simultaneous Localization and Mapping) to estimate the position and orientation. Alternatively, the position and orientation estimation unit 312 may calculate the position and orientation of the camera using information obtained by various sensors (GPS or acceleration sensor) included in the image capture device 301. In the former case, the position and orientation can be estimated from information from the image capture unit 311. In the latter case, image processing can be performed with low processing power.

[0138] When multiple images are captured in one environment, the position and orientation integration unit 313 integrates the camera positions and orientations estimated in each capture to calculate a position and orientation that can be handled in the same space. Specifically, the position and orientation integration unit 313 uses the three-dimensional coordinate axes of the position and orientation obtained in the first capture as the reference coordinate axes. Then, the position and orientation integration unit 313 converts the coordinates of the position and orientation obtained in the second and subsequent captures into coordinates in the space of the reference coordinate axes.

[0139] The region detection unit 314 detects a region in the target space where three-dimensional reconstruction is not possible or a region where the accuracy of three-dimensional reconstruction is low, using the images stored in the image storage unit 317 and the positions and orientations stored in the position and orientation storage unit 318. A region in the target space where three-dimensional reconstruction is not possible is, for example, a region where no images have been captured. A region where the accuracy of three-dimensional reconstruction is low is, for example, a region where the number of images captured of the region is small (fewer than a predetermined number). A region with low accuracy is a region where, when three-dimensional position information is generated, there is a large error between the generated three-dimensional position information and the actual position. The region detection unit 314 also stores information about the detected region in the region information storage unit 319. The region detection unit 314 may detect a region where three-dimensional reconstruction is possible and determine that regions other than the detected region are regions where three-dimensional reconstruction is not possible.

[0140] Furthermore, the information stored in the region information storage unit 319 may be two-dimensional information superimposed on an image, or may be three-dimensional information such as three-dimensional coordinate information.

[0141] The UI unit 315 presents the user with the captured image and area information detected by the area detection unit 314. The UI unit 315 also has an input function that enables the user to input an instruction to start shooting and an instruction to stop shooting. For example, the UI unit 315 is a display with a touch panel.

[0142] Next, the operation of the image capturing device 301 according to this embodiment will be described. FIG. 22 is a flowchart showing the operation of the image capturing device 301. In the image capturing device 301, image capturing is started and stopped in response to a user instruction. Specifically, image capturing is started by pressing a start image capturing button on the UI. When an instruction to start image capturing is input (Yes in S801), the image capturing unit 311 starts image capturing. The captured image is stored in the image storage unit 317.

[0143] Next, the position and orientation estimation unit 312 calculates the position and orientation each time an image is added (S802). The calculated position and orientation are stored in the position and orientation storage unit 318. At this time, if a three-dimensional point cloud generated by, for example, SLAM exists in addition to the position and orientation, the generated three-dimensional point cloud is also stored.

[0144] Next, the position and orientation integration unit 313 integrates the positions and orientations (S803). Specifically, the position and orientation integration unit 313 uses the position and orientation estimation results and the images to determine whether the three-dimensional coordinate spaces of the positions and orientations of the images captured up to now and the position and orientation of the newly captured image can be integrated, and if so, integrates them. In other words, the position and orientation integration unit 313 converts the coordinates of the position and orientation of the newly captured image into the coordinate system of the previous positions and orientations. This allows multiple positions and orientations to be expressed in a single three-dimensional coordinate space. This makes it possible to share data obtained by multiple captures, thereby improving the accuracy of position and orientation estimation.

[0145] Next, the region detection unit 314 detects an imaged region, etc. (S804). Specifically, the region detection unit 314 generates three-dimensional position information (a three-dimensional point cloud, a three-dimensional model, a depth image, etc.) using the position and orientation estimation result and the image, and detects a region where three-dimensional reconstruction is not possible or where the accuracy of three-dimensional reconstruction is low, using the generated three-dimensional position information. Furthermore, the region detection unit 314 stores information about the detected region in the region information storage unit 319. Next, the UI unit 315 displays information about the region obtained by the above process (S805).

[0146] This series of processes is repeated until the end of imaging (S806). For example, this series of processes is repeated every time one or more frames of images are acquired.

[0147] 23 is a flowchart of the position and orientation estimation process (S802). First, the position and orientation estimation unit 312 acquires images from the image storage unit 317 (S811). Next, the position and orientation estimation unit 312 calculates the position and orientation of the camera in each image using the acquired images (S812). For example, the position and orientation estimation unit 312 calculates the position and orientation using image processing such as SLAM or SfM (Structure from Motion). Note that if the image capture device 301 has sensors such as an IMU (Inertial Measurement Unit), the position and orientation estimation unit 312 may estimate the position and orientation using information obtained by these sensors.

[0148] Furthermore, the position and orientation estimation unit 312 may use the results of a calibration performed in advance as camera parameters such as the focal length of the lens, or may calculate the camera parameters simultaneously with the estimation of the position and orientation.

[0149] Next, the position and orientation estimation unit 312 stores the calculated position and orientation information in the position and orientation storage unit 318 (S813). If the calculation of the position and orientation information fails, information indicating the failure may be stored in the position and orientation storage unit 318. This makes it possible to know the point and time of the failure and the type of image for which the failure occurred, and this information can be used when re-capturing images, etc.

[0150] FIG. 24 is a flowchart of the position and orientation integration process (S803). First, the position and orientation integration unit 313 acquires an image from the image storage unit 317 (S821). Next, the position and orientation integration unit 313 acquires the current position and orientation (S822). Next, the position and orientation integration unit 313 acquires at least one image and position and orientation of a captured route other than the current capture route (S823). Note that the captured route can be generated, for example, from time-series information of the camera position and orientation obtained by SLAM. Also, information on the captured route is stored, for example, in the position and orientation storage unit 318. Specifically, the SLAM results are stored for each capture attempt, and in the Nth capture (current route), the three-dimensional coordinate axes of the Nth capture are integrated with the three-dimensional coordinate axes of the 1st to N-1th results (past routes). Note that instead of the SLAM results, position information obtained by GPS or Bluetooth (registered trademark) may be used.

[0151] Next, the position and orientation integration unit 313 determines whether integration is possible (S824). Specifically, the position and orientation integration unit 313 determines whether the position and orientation of the acquired photographed path and the image are similar to the current position and orientation and image, and determines that integration is possible if they are similar, and determines that integration is not possible if they are not similar. More specifically, the position and orientation integration unit 313 calculates feature amounts that express the characteristics of the entire image from each image, and determines whether the images have similar viewpoints by comparing them. Furthermore, if the image capture device 301 has a GPS and the absolute position of the image capture device 301 is known, the position and orientation integration unit 313 may use that information to determine whether the image was captured at the same or similar position as the current image.

[0152] If integration is possible (Yes in S824), the position and orientation integration unit 313 performs path integration processing (S825). Specifically, the position and orientation integration unit 313 calculates the three-dimensional relative position between the current image and a reference image that captures an area similar to the current image. The position and orientation integration unit 313 calculates the coordinates of the current image by adding the calculated three-dimensional relative position to the coordinates of the reference image.

[0153] A specific example of the path integration process will be described below. Fig. 25 is a plan view showing how images are captured in the target space. Path C is the path of camera A, which has already captured images and whose position and orientation have been estimated. Note that this figure shows a case where camera A is located at a predetermined position on path C. Path D is the path of camera B, which is currently capturing images. At this time, the current cameras B and A are capturing images with similar fields of view. Note that although an example is shown here in which two images are obtained by different cameras, two images captured at different times by the same camera may also be used. Fig. 26 is a diagram showing example images and an example of comparison processing in this case.

[0154] 26, the position and orientation integration unit 313 extracts features such as ORB (Oriented FAST and Rotated BRIEF) features for each image, and extracts features for the entire image based on the distribution or number of the features. For example, the position and orientation integration unit 313 clusters features that appear in the image, such as a bag of words, and uses a histogram for each class as a feature.

[0155] The position and orientation integration unit 313 compares the overall image feature amounts between the images, and if it is determined that the images show the same location, it calculates the relative three-dimensional positions between the cameras by performing feature point matching between the images. In other words, the position and orientation integration unit 313 searches for an image with similar overall screen feature amounts from multiple images of the captured route. Based on this relative positional relationship, the position and orientation integration unit 313 converts the three-dimensional position of route D into the coordinate system of route C. This makes it possible to represent multiple routes in a single coordinate system. In this way, the positions and orientations of multiple routes can be integrated.

[0156] If the image capturing device 301 has a sensor capable of detecting an absolute position, such as a GPS, the position and orientation integration unit 313 may perform integration processing using the detection results. For example, the position and orientation integration unit 313 may perform processing using the detection results of the sensor without performing the above-described processing using the image, or may use the detection results of the sensor in addition to the image. For example, the position and orientation integration unit 313 may use GPS information to narrow down the images to be compared. Specifically, the position and orientation integration unit 313 may set, as comparison targets, images whose positions and orientations are within a range of ±0.001 degrees or less from the latitude and longitude of the current camera position as determined by GPS. This reduces the amount of processing.

[0157] 27 is a flowchart of the area detection process (S804). First, the area detection unit 314 acquires an image from the image storage unit 317 (S831). Next, the area detection unit 314 acquires the position and orientation from the position and orientation storage unit 318 (S832). Furthermore, the area detection unit 314 acquires from the position and orientation storage unit 318 a three-dimensional point cloud indicating the three-dimensional positions of feature points generated by SLAM or the like.

[0158] Next, the region detection unit 314 detects an uncaptured region using the acquired image, position and orientation, and 3D point cloud (S833). Specifically, the region detection unit 314 projects the 3D point cloud onto the image and determines the periphery of the pixel onto which the 3D point is projected (within a predetermined distance from the pixel) as a restorable region (captured region). Note that the farther the projected 3D point is from the capture position, the greater the predetermined distance may be. Also, when a stereo camera is used, the region detection unit 314 may estimate the restorable region from a parallax image. Also, when an RGB-D camera is used, the determination may be made using the obtained depth value. Note that the process of determining the captured region will be described in detail later. Also, the region detection unit 314 may determine not only whether a 3D model can be generated, but also the accuracy of the 3D reconstruction.

[0159] Finally, the region detection unit 314 outputs the obtained region information (S834). This region information may be, for example, an image in which information about each region is superimposed on the captured image, or information in which information about each region is arranged in a three-dimensional space such as a three-dimensional map.

[0160] 28 is a flowchart of the display process (S805). First, the UI unit 315 checks whether there is information to display (S841). Specifically, the UI unit 315 checks whether there is any newly added image in the image storage unit 317, and if there is, determines that there is display information. The UI unit 315 also checks whether there is any newly added information in the area information storage unit 319, and if there is, determines that there is display information.

[0161] If there is display information (Yes in S841), the UI unit 315 acquires the display information such as image and area information (S842). Next, the UI unit 315 displays the acquired display information (S843).

[0162] 29 is a diagram showing an example of a UI screen displayed by the UI unit 315. The UI screen includes a start / stop shooting button 321, an image being shot 322, area information 323, and a character display area 324.

[0163] The start / stop shooting button 321 is an operation unit that allows the user to instruct the start and stop of shooting. The image being shot 322 is the image currently being shot. The area information 323 displays areas that have been shot and low-accuracy areas. Here, a low-accuracy area is an area where the accuracy of the generated three-dimensional model is low when a three-dimensional model is generated using a shot image. The text display area 324 displays information about areas that have not been shot or low-accuracy areas in text. Note that audio or the like may be used instead of text.

[0164] FIG. 30 is a diagram illustrating an example of the region information 323. As illustrated in FIG. 30, for example, an image in which information indicating each region is superimposed on the image being captured is used. Furthermore, as the capturing viewpoint moves, the display of the region information 323 changes in real time. This allows the user to easily understand captured regions, etc., while referring to the image from the current camera viewpoint. For example, captured regions and low-accuracy regions are displayed in different colors. Furthermore, captured regions on the current route and captured regions on a different route are displayed in different colors. Note that the information indicating each region may be superimposed on the image, as illustrated in FIG. 30, or may be represented by letters or symbols. In other words, the information may be information that allows the user to visually identify each region. This allows the user to know that they should capture uncolored regions and low-accuracy regions, thereby avoiding missing captures. Furthermore, in order to perform the route integration process, the user can be presented with regions to be captured so that the captured regions are continuous with captured regions on a different route.

[0165] Note that for areas such as low-precision areas that should be brought to the user's attention, a display that makes it easy for the user to notice may be used, such as blinking. Note that the image on which the area information is superimposed may be a past image.

[0166] Note that, although the image being captured 322 and the area information 323 are displayed separately here, the area information may be superimposed on the image being captured 322. In this case, it is possible to reduce the area required for display, making it easier to view even on small devices such as smartphones. Therefore, these display methods may be switched depending on the type of device. For example, on devices with small screen sizes such as smartphones, the area information may be superimposed on the image being captured 322, and on devices with large screen sizes such as tablet devices, the image being captured 322 and the area information 323 may be displayed separately.

[0167] The UI unit 315 also presents, by text or voice, information such as how many meters ahead of the current position a low-accuracy area has occurred. This distance can be calculated from the results of position estimation. By using text information, the user can be notified of the information correctly. Furthermore, when voice is used, the notification can be given safely because there is no need to take one's eyes off the camera while shooting.

[0168] Instead of displaying the area in two dimensions on an image, area information may be superimposed on real space using AR (Augmented Reality) glasses or HUD (Head-Up Display). This increases the affinity with the image as seen from a real viewpoint, allowing the user to intuitively identify areas to be photographed.

[0169] If the image capturing device 301 fails to estimate the position and orientation, the image capturing device 301 may notify the user of the failure. Fig. 31 shows an example of a display in this case.

[0170] For example, an image when position and orientation estimation fails may be displayed in the area information 323. Also, text or audio may be presented to prompt the user to resume shooting from that point. This allows the user to quickly try again when the error occurs.

[0171] The camera device 301 may also predict the location of the failure based on the time elapsed since the failure and the moving speed, and provide information such as "Please go back 5 meters" by voice or text. Alternatively, the camera device 301 may display a two-dimensional map or a three-dimensional dot map and display the location of the shooting failure on the map.

[0172] Furthermore, the photographing device 301 may detect that the user has returned to the position of failure and notify the user of this fact by using text, an image, sound, vibration, etc. For example, it is possible to detect that the user has returned to the position of failure by using the feature amount of the entire image.

[0173] Furthermore, when the image capturing device 301 detects a low-accuracy area, it may instruct the user to capture the image again and instruct the user on how to capture the image. Here, the capturing method may be, for example, capturing a larger image of the area. For example, the image capturing device 301 may issue this instruction using text, an image, sound, vibration, or the like.

[0174] This improves the quality of the acquired data, thereby improving the accuracy of the generated three-dimensional model.

[0175] FIG. 32 is a diagram showing a display example when a low-accuracy area is detected. For example, as shown in FIG. 32, instructions to the user are given using text. In this example, a message indicating that a low-accuracy area has been detected is displayed in text display area 324. Also, instructions to the user for capturing an image of the low-accuracy area, such as "Please go back 5 m," are displayed. Furthermore, the image capturing device 301 may erase these displays when it detects that the user has moved to the instructed location or captured an image of the instructed area.

[0176] FIG. 33 is a diagram showing another example of instructions to the user. For example, the direction and distance to be moved may be displayed to the user using an arrow 325 shown in FIG. 33. FIG. 34 is a diagram showing an example of this arrow. For example, as shown in FIG. 34, the direction of movement is indicated by the angle of the arrow, and the distance to the destination is indicated by the size of the arrow. The display form of the arrow may be changed depending on the distance. For example, the display form may be color, or the presence or absence of an effect or the size of the effect. Examples of effects include blinking, movement, zooming, etc. The darkness of the arrow may also be changed. For example, the closer the distance, the greater the effect may be, or the darker the color of the arrow may be. A combination of these may also be used.

[0177] Furthermore, any display that allows the user to recognize the direction, such as a triangle or a finger icon, may be used instead of an arrow.

[0178] Furthermore, instead of superimposing on the image being photographed, the area information 323 may be superimposed on an image from a third-person viewpoint such as a plan view or a perspective view. For example, in an environment where a three-dimensional map is available, the photographed area, etc. may be superimposed on the three-dimensional map.

[0179] Fig. 35 is a diagram showing an example of area information 323 in this case. Fig. 35 is a plan view of a three-dimensional map, showing photographed areas and low-accuracy areas. Also shown are a current camera position 331 (the position and direction of the photographing device 301), a current route 332, and a past route 333.

[0180] If a CAD map or other such map of the target space exists at a construction site or the like, or if 3D modeling has previously been performed at the same location and 3D map information has already been generated, the image capture device 301 may utilize that information. Also, if a GPS or the like is available, the image capture device 301 may create map information based on latitude and longitude information obtained by the GPS.

[0181] In this way, by superimposing the estimated position and orientation of the camera and the captured area on a 3D map and displaying it as a bird's-eye view, the user can easily understand which area has been captured and which route was taken to capture it, etc. For example, in the example shown in Fig. 35, the user can be made aware of missed shots, such as the back side of a pillar, which is difficult to see when presented in the image being captured, thereby realizing efficient photography.

[0182] Note that even in environments without CAD or other devices, a third-person view can be used without a map. In this case, visibility is reduced because there is no reference object, but the user can understand the positional relationship between the current shooting position and low-accuracy areas. The user can also confirm situations where areas where objects are expected to be present have not been captured. Therefore, even in this case, shooting efficiency can be improved.

[0183] Although a plan view is shown as an example here, map information viewed from other viewpoints may also be used. Furthermore, the image capture device 301 may have a function for changing the viewpoint of the 3D map. For example, a UI may be used to allow the user to change the viewpoint.

[0184] Furthermore, the image capturing device 301 may display both the area information 323 shown in FIG. 29 and the area information shown in FIG. 35, or may have a function to switch between these displays.

[0185] Next, we will explain examples of methods for determining and presenting photographed areas. When SLAM is used to estimate the camera position, a 3D point cloud is generated for feature points such as the corners of objects in the image in addition to the camera's position and orientation information. Since it can be determined that 3D modeling is possible for the area for which these 3D points have been generated, the photographed area can be presented on the image by projecting the 3D points onto the image.

[0186] Fig. 36 is a plan view showing the shooting situation of the target area. The black circles in the figure indicate the generated 3D points (feature points). Fig. 37 is a diagram showing an example in which the 3D points are projected onto an image and the area around the 3D points is determined to be the shot area.

[0187] Alternatively, the camera device 301 may generate a mesh by connecting three-dimensional points and determine the captured area using the mesh. Fig. 38 is a diagram showing an example of the captured area in this case. As shown in Fig. 38, the camera device 301 may determine the area in which the mesh is generated as the captured area.

[0188] If the image capturing device 301 uses a stereo camera or an RGB-D sensor, the captured area may be determined using a parallax image or depth image obtained from the stereo camera or RGB-D sensor. This allows for obtaining dense three-dimensional information with light processing, thereby enabling more accurate determination of the captured area.

[0189] Also, in this example, self-position estimation (position and orientation estimation) is performed using SLAM, but this does not necessarily apply if the position and orientation of the camera can be estimated during shooting.

[0190] In addition to presenting the captured area, the image capturing device 301 may also predict the accuracy of the restored image and display the predicted accuracy. Specifically, 3D points are calculated from feature points in the image. When these 3D points are projected onto the image, the projected 3D points may deviate from the reference feature points. This deviation is called a reprojection error, and the accuracy can be evaluated using the reprojection error. Specifically, the larger the reprojection error, the lower the accuracy can be determined to be.

[0191] For example, the image capturing device 301 may indicate the accuracy by the color of the captured area. For example, high accuracy areas may be displayed in blue, and low accuracy areas may be displayed in red. The accuracy may be expressed in stages by different colors or shades of color. This allows the user to easily understand the accuracy of each area.

[0192] The image capturing device 301 may also use a depth image obtained by an RGB-D sensor or the like to determine whether restoration is possible and evaluate the accuracy. Fig. 39 is a diagram showing an example of a depth image. In this figure, the darker the color (the denser the hatching), the farther the distance.

[0193] Here, the closer the distance from the camera, the higher the accuracy of the generated three-dimensional model tends to be, and the farther the distance, the lower the accuracy tends to be. Therefore, for example, the image capturing device 301 may determine that pixels up to a certain depth range (for example, up to 5 m) are the captured range. Figure 40 is a diagram showing an example in which an area up to a certain depth range is determined to be an imaged area. In this figure, the hatched area is determined to be an imaged area.

[0194] The image capturing device 301 may also determine the accuracy of an area based on the distance to the area. That is, the image capturing device 301 may determine that the closer the distance, the higher the accuracy. For example, the relationship between distance and accuracy may be defined linearly, or another definition may be used.

[0195] Furthermore, when using a stereo camera, the image capturing device 301 can generate a depth image from a parallax image, so the imaged area and accuracy may be determined in the same manner as when using a depth image. In this case, the image capturing device 301 may determine that an area for which a depth value cannot be calculated from the parallax image is unrecoverable (uncaptured). Alternatively, the image capturing device 301 may estimate a depth value from surrounding pixels for an area for which a depth value cannot be calculated. For example, the image capturing device 301 calculates an average value of 5 × 5 pixels centered on the target pixel.

[0196] Furthermore, information indicating the three-dimensional positions of the photographed areas determined above and the accuracy of each three-dimensional position is stored. As described above, the coordinates of the camera position and orientation are integrated, so the coordinates of the obtained area information can also be integrated. This generates information indicating the three-dimensional positions of the photographed areas in the target space and the accuracy of each three-dimensional position.

[0197] As described above, the photography instruction device according to this embodiment performs the processing shown in Fig. 41. The photography device 301 photographs a plurality of first images of the target space (S851), generates first three-dimensional position information of the target space (for example, a sparse three-dimensional point cloud or a depth image) based on the plurality of first images and the first photographing positions and orientations of each of the plurality of first images (S852), and determines a second region for which it is difficult to generate second three-dimensional position information of the target space (for example, a dense three-dimensional point cloud) that is more detailed than the first three-dimensional position information using the plurality of first images, by using the first three-dimensional position information without generating the second three-dimensional position information (S853).

[0198] Here, the region where it is difficult to generate three-dimensional position information includes at least one of a region where it is impossible to calculate the three-dimensional position and a region where the error between the three-dimensional position and the actual position is larger than a predetermined value. Specifically, the second region includes at least one of (1) a region that is not captured from some of the multiple capture viewpoints, (2) a region that is captured but has significant blur, (3) a region that is captured but has lower contrast than other regions and therefore fewer feature points, (4) a region that is captured but is farther from the capture viewpoint than other regions and is therefore estimated to have a large error between the calculated three-dimensional position and the actual position even if the three-dimensional position is calculated, and (5) a region that is more affected by lens distortion than other regions. Note that blur can be detected, for example, by determining the positional change of feature points over time.

[0199] Furthermore, detailed three-dimensional position information is, for example, three-dimensional position information with high spatial resolution. The spatial resolution of three-dimensional position information refers to the distance between two adjacent three-dimensional positions when the two adjacent three-dimensional positions can be distinguished as different three-dimensional positions. Furthermore, high spatial resolution refers to a small distance between two adjacent three-dimensional positions. In other words, three-dimensional position information with high spatial resolution has information on more three-dimensional positions in a space of a given size. Furthermore, three-dimensional position information with high spatial resolution may be referred to as dense three-dimensional position information, and three-dimensional position information with low spatial resolution may be referred to as sparse three-dimensional position information.

[0200] The detailed three-dimensional position information may be three-dimensional position information with a large amount of information. For example, the first three-dimensional position information may be distance information from one viewpoint such as a depth image, and the second three-dimensional position information may be a three-dimensional model such as a three-dimensional point cloud from which distance information from an arbitrary viewpoint can be obtained.

[0201] Furthermore, for example, the target space and the subject are the same concept and refer to a common area that is photographed.

[0202] This allows the photographing device 301 to use the first three-dimensional position information to determine the second area for which it is difficult to generate second three-dimensional position information without generating second three-dimensional position information, thereby improving the efficiency of photographing multiple images for generating second three-dimensional position information.

[0203] For example, the second region is at least one of a region where no image has been captured and a region where the accuracy of the second three-dimensional position information is estimated to be lower than a predetermined standard. Here, the standard is, for example, a threshold value for the distance between two different three-dimensional positions. In other words, the second region is a region where, when three-dimensional position information is generated, the difference between the generated three-dimensional position information and the actual position is greater than a predetermined threshold.

[0204] For example, the first three-dimensional position information includes a first three-dimensional point cloud, and the second three-dimensional position information includes a second three-dimensional point cloud that is denser than the first three-dimensional point cloud.

[0205] For example, in the determination (S853), the imaging device 301 determines a third area of ​​the target space corresponding to the area surrounding the first three-dimensional point cloud (an area within a predetermined distance from the first three-dimensional point cloud), and determines an area other than the third area as a second area (e.g., Figure 37).

[0206] For example, in the determination (S853), the photographing device 301 generates a mesh using the first three-dimensional point cloud, and determines that the area other than the third area in the target space corresponding to the area in which the mesh is generated is the second area (e.g., Figure 38).

[0207] For example, in the determination (S853), the image capturing apparatus 301 determines the second region based on the reprojection error of the first 3D point cloud.

[0208] For example, the first three-dimensional position information includes a depth image. Based on the depth image, the image capturing device 301 determines an area within a predetermined distance from the image capturing viewpoint as a third area, and determines an area other than the third area as a second area, as shown in FIG.

[0209] For example, the image capturing device 301 further aligns the coordinate system of the multiple first image capturing positions and orientations with the coordinate system of the multiple second image capturing positions and orientations using multiple second images that have already been captured, the second image capturing positions and orientations of each of the multiple second images, the multiple first images, and the multiple first image capturing positions and orientations. This allows the image capturing device 301 to determine the second area using information obtained from multiple captures.

[0210] For example, the image capturing device 301 further displays the second area or a third area other than the second area (e.g., an area that has been photographed) while capturing an image of the target space (e.g., FIG. 30). In this way, the image capturing device 301 can present the second area to the user.

[0211] For example, the image capturing device 301 displays information indicating the second area or the third area by superimposing it on one of the multiple images (for example, FIG. 30). This allows the image capturing device 301 to present the position of the second area in the image to the user, so that the user can easily grasp the position of the second area.

[0212] For example, the image capturing device 301 displays information indicating the second area or the third area superimposed on a map of the target space (e.g., FIG. 35). This allows the image capturing device 301 to present the position of the second area in the surrounding environment to the user, allowing the user to easily grasp the position of the second area.

[0213] For example, the imaging device 301 displays the second region and the restoration accuracy (accuracy of three-dimensional reconstruction) of each region included in the second region. This allows the user to understand the restoration accuracy of each region in addition to the second region, and allows appropriate imaging to be performed based on this.

[0214] For example, the image capturing device 301 further presents the user with instructions to cause the user to capture an image of the second area (for example, FIGS. 32 and 33). This allows the user to efficiently capture an appropriate image.

[0215] For example, the instruction includes at least one of the direction and the distance from the current position to the second area (for example, FIGS. 33 and 34). This allows the user to efficiently take appropriate photographs.

[0216] For example, the imaging device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0217] Although the photography instruction device, photography device, etc. according to the embodiment of the present disclosure have been described above, the present disclosure is not limited to this embodiment.

[0218] For example, the photographing instruction device according to Embodiment 1 may be combined with the photographing device according to Embodiment 2. For example, the photographing instruction device may have at least a part of the processing units included in the photographing device. Furthermore, the photographing device may have at least a part of the processing units included in the photographing instruction device. Furthermore, at least a part of the processing units included in the photographing instruction device may be combined with at least a part of the processing units included in the photographing device. In other words, the photographing instruction method according to Embodiment 1 may include at least a part of the processing included in the photographing method according to Embodiment 2. Furthermore, the photographing method according to Embodiment 2 may include at least a part of the processing included in the photographing instruction method according to Embodiment 1. Furthermore, at least a part of the processing included in the photographing method according to Embodiment 1 may be combined with at least a part of the processing included in the photographing method according to Embodiment 2.

[0219] Furthermore, each processing unit included in the photography instruction device and the photography device according to the above-described embodiments is typically realized as an LSI, which is an integrated circuit. These may be individually implemented as single chips, or some or all of them may be integrated into a single chip.

[0220] Furthermore, the integration is not limited to LSI, but may be realized by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays), which can be programmed after LSI fabrication, or reconfigurable processors, which allow the connections and settings of circuit cells within LSIs to be reconfigured, may also be used.

[0221] In each of the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0222] The present disclosure may also be realized as a photography instruction method or the like executed by a photography instruction device or the like.

[0223] The division of functional blocks in the block diagram is an example, and multiple functional blocks may be realized as a single functional block, one functional block may be divided into multiple blocks, or some functions may be moved to another functional block.Furthermore, the functions of multiple functional blocks having similar functions may be processed in parallel or in time-sharing by a single piece of hardware or software.

[0224] The order in which the steps in the flowchart are executed is merely an example for specifically explaining the present disclosure, and other orders may be used. Some of the steps may be executed simultaneously (in parallel) with other steps.

[0225] While the photography instruction device and the like according to one or more aspects have been described based on the embodiments, the present disclosure is not limited to these embodiments. As long as they do not deviate from the spirit of the present disclosure, various modifications conceivable by those skilled in the art to the present embodiments and configurations constructed by combining components of different embodiments may also be included within the scope of one or more aspects. [Industrial Applicability]

[0226] The present disclosure can be applied to a photography instruction device. [Explanation of symbols]

[0227] 100 Terminal Device 101 Imaging unit 102 Control section 103 Position and orientation estimation unit 104 Three-dimensional reconstruction unit 105 Image Analysis Unit 106 Point cloud analysis section 107 Communications Department 108 UI section 111 Video Storage Department 112 Camera posture storage unit 113 3D Model Storage Unit 301 Imaging equipment 302 Reconfiguration device 311 Imaging unit 312 Position and orientation estimation unit 313 Position and orientation integration unit 314 Area detection unit 315 UI section 316 Control Unit 317 Image Storage Unit 318 Position and orientation storage section 319 Area information storage section 321 Start / Stop Recording Button 322 Images being taken 323 Area information 324 character display area 325 Arrow 331 Camera Position 332 Current Route 333 Past Routes

Claims

1. An imaging method performed by an imaging device, comprising: capturing a plurality of first images of the target space; generating first three-dimensional position information of the target space based on the plurality of first images and a first shooting position and orientation of each of the plurality of first images; determining a second region for which it is difficult to generate second three-dimensional position information of the target space that is more detailed than the first three-dimensional position information, using the first three-dimensional position information without generating the second three-dimensional position information; the first three-dimensional position information includes a first three-dimensional point cloud; the second three-dimensional position information includes a second three-dimensional point cloud that is denser than the first three-dimensional point cloud; In the determination, a mesh is generated using the first three-dimensional point cloud, and a region other than a third region in the target space corresponding to the region in which the mesh is generated is determined to be the second region. Shooting method.

2. The photographing method further comprises: Using a plurality of second images already taken, second image capturing positions and orientations of each of the plurality of second images, the plurality of first images, and the plurality of first image capturing positions and orientations, a coordinate system of the plurality of first image capturing positions and orientations is aligned with a coordinate system of the plurality of second image capturing positions and orientations. The photographing method according to claim 1.

3. The photographing method further comprises: The second area or the third area is displayed while the target space is being photographed. The photographing method according to claim 1.

4. information indicating the second region or the third region is displayed superimposed on any one of the plurality of first images; The photographing method according to claim 3.

5. and displaying information indicating the second area or the third area on a map of the target space in a superimposed manner. The photographing method according to claim 3.

6. Displaying the second region and the restoration accuracy of each region included in the second region. The photographing method according to any one of claims 3 to 5.

7. a processor; a memory; The processor uses the memory to: capturing a plurality of first images of the target space; generating first three-dimensional position information of the target space based on the plurality of first images and a first shooting position and orientation of each of the plurality of first images; determining a second region for which it is difficult to generate second three-dimensional position information of the target space that is more detailed than the first three-dimensional position information, using the first three-dimensional position information without generating the second three-dimensional position information; the first three-dimensional position information includes a first three-dimensional point cloud; the second three-dimensional position information includes a second three-dimensional point cloud that is denser than the first three-dimensional point cloud; In the determination, a mesh is generated using the first three-dimensional point cloud, and a region other than a third region in the target space corresponding to the region in which the mesh is generated is determined to be the second region. Filming equipment.

Citation Information

Patent Citations

  • Three-dimensional information acquisition system, three-dimensional information acquisition method, and program

    JP2008089410A

  • Image capturing device for three-dimensional measurement and method therefor

    JP2010256253A

  • Contour line measurement device and robot system

    JP2016061687A

  • Image management apparatus, image management method and program

    JP2017130146A