Information processing system, information processing method, and program

WO2025187211A8PCT designated stage Publication Date: 2025-10-02SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/001174
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-04
Filing Date
2025-01-16
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing methods for generating 3D models from multiple captured images are prone to quality degradation due to the influence of pure rotational shooting motion, which is difficult for users to detect and correct.

Method used

An information processing system that estimates the shooting position and orientation for each captured image based on a common field of view with other images, calculates the reliability of these estimates, and provides guidance to users to minimize the impact of pure rotational motion through target position and orientation adjustments.

Benefits of technology

This approach enables the generation of high-quality 3D models by ensuring accurate estimation of shooting positions and orientations, thereby reducing the quality degradation caused by rotational motion and enhancing the overall model precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025001174_02102025_PF_FP_ABST
    Figure JP2025001174_02102025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to one embodiment of the present technology is provided with an estimation unit and a reliability calculation unit. The estimation unit estimates an imaging position attitude for each of a plurality of captured images which have been captured in order to generate a 3D model of a target object in a real space, such estimation performed on the basis of a field of view shared with other captured images. The reliability calculation unit calculates, for the estimated imaging position attitude, the degree of reliability relating to the impact on the estimation accuracy of the imaging position attitude caused by imaging motion when the captured image is imaged. Due to this configuration, it is possible to generate a high-quality 3D model.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method, and program

[0001] The present technology relates to an information processing system, an information processing method, and a program that can be applied to generating a 3D model.

[0002] U.S. Patent No. 6,263,666 discloses an unmanned aerial vehicle used to generate a 3D model, which acquires multiple images of a scanned object and updates the 3D model in real time, allowing for rapid in-situ inspection of the 3D model.

[0003] Japanese Patent Application Laid-Open No. 2003-124222 discloses an imaging system that can easily calculate at least one of a blind spot area and a visible area from a subject.

[0004] JP 2023-72064 A JP 2023-5022 A

[0005] There is a demand for technology that enables the generation of high-quality 3D models from multiple captured images of a real object.

[0006] In view of the above circumstances, an object of the present technology is to provide an information processing system, an information processing method, and a program that enable generation of a high-quality 3D model.

[0007] To achieve the above object, an information processing device according to one embodiment of the present technology includes an estimation unit and a reliability calculation unit. The estimation unit estimates a shooting position and orientation for each of a plurality of captured images taken for generating a 3D model of a target object in real space, based on a common field of view with other captured images. The reliability calculation unit calculates a reliability of the estimated shooting position and orientation regarding an influence on estimation accuracy of the shooting position and orientation caused by shooting motion when capturing the captured images.

[0008] In this information processing system, for each of a plurality of captured images taken to generate a 3D model of a target object in real space, the shooting position and orientation are estimated based on a common field of view with the other captured images. Furthermore, for the estimated shooting position and orientation, a reliability is calculated regarding the influence of shooting motion on the estimation accuracy of the shooting position and orientation. Based on this reliability, it becomes possible to generate a high-quality 3D model.

[0009] The reliability calculation unit may calculate, for each of the plurality of captured images, a degree of consistency of a geometric relationship with the other captured images that have the common field of view, and calculate the reliability based on the calculated degree of consistency of the geometric relationship.

[0010] The shooting motion may be a pure rotational shooting motion.

[0011] The information processing system may further include an acquisition unit that acquires a captured image of the target object in response to capturing of the target object. In this case, the estimation unit may estimate the shooting position and orientation of the new captured image each time a new captured image is acquired based on the common field of view with an existing captured image that has already been acquired and whose shooting position and orientation have been estimated. Furthermore, the reliability calculation unit may calculate the reliability for the shooting position and orientation of the new captured image.

[0012] The estimation unit may estimate the shooting position and orientation of the new captured image by selecting the existing captured image that has corresponding feature points as a corresponding point pair for the new captured image, and calculating a geometric relationship between the new captured image and the existing captured image based on the corresponding point pair.

[0013] The reliability calculation unit may calculate the reliability of the new captured image based on the consistency of the geometric relationship with the existing captured image that has the common field of view, has already had the reliability calculated, and has the reliability higher than a predetermined threshold.

[0014] The reliability calculation unit may calculate, using the corresponding point pair as a reference, a consistency of the geometric relationship with the existing captured image for which the reliability is higher than a predetermined threshold.

[0015] The reliability calculation unit may select, as an anchor image, an existing captured image whose capture timing is closest to that of the new captured image from among the existing captured images whose reliability has already been calculated and whose reliability is higher than a predetermined threshold, and calculate the reliability based on whether or not there is a common field of view with the anchor image and the degree of consistency of the geometric relationship with the anchor image.

[0016] When the common field of view exists, the reliability calculation unit may calculate, as the reliability, a degree of consistency of the geometric relationship with the anchor image.

[0017] The information processing system may further include an information presenting unit that, when the reliability is smaller than a predetermined threshold, calculates a target position and orientation for photographing the target object and presents the calculated target position and orientation to a user.

[0018] The information presenting unit may calculate the target position and orientation based on the shooting position and orientation of the anchor image.

[0019] The target position and orientation may include a target shooting position and a target shooting orientation. In this case, the information presentation unit may calculate the target shooting position and the target shooting orientation based on the shooting position and orientation of the anchor image so as to avoid a pure rotational shooting motion.

[0020] The information presenting unit may calculate, as the target shooting position, a shooting position moved a predetermined distance in one direction from the shooting position of the anchor image included in the shooting position and orientation of the anchor image.

[0021] The one direction may be a direction from the shooting position of the existing captured image taken immediately before the anchor image, which is included in the shooting position and orientation of the existing captured image taken immediately before the anchor image, toward the shooting position of the anchor image.

[0022] The information presenter may calculate the target position so that the distance from the shooting position of the anchor image to the target position is greater than a predetermined minimum movement distance and less than a predetermined maximum movement distance.

[0023] The information processing system may further include a generation unit that generates a 3D model of the target object based on the plurality of captured images, the shooting positions and orientations of each of the plurality of captured images, and the reliability calculated for the shooting positions and orientations.

[0024] The estimation unit may estimate the shooting position and orientation by performing structure from motion (SfM) processing. In this case, the generation unit may generate the 3D model by performing multi-view stereo (MVS) processing.

[0025] According to one aspect of the present technology, there is provided an information processing method executed by a computer system, the information processing method including: estimating a shooting position and orientation for each of a plurality of captured images taken for generating a 3D model of a target object in real space, based on a common field of view with other captured images; and calculating a reliability of the estimated shooting position and orientation regarding an influence on estimation accuracy of the shooting position and orientation caused by a shooting motion when capturing the captured images.

[0026] A program according to an embodiment of the present technology causes a computer system to execute the information processing method.

[0027] 9 is a schematic diagram for explaining an overview of photogrammetry. It is a schematic diagram for explaining an overview of a photogrammetry algorithm. It is a schematic diagram for explaining shooting motion when shooting a target object. It is a schematic diagram for explaining the relationship between SfM processing and the quality of a generated 3D model. It is a schematic diagram showing an example of the configuration of a 3D model generation system (basic configuration). It is a block diagram showing an example of the functional configuration of a smartphone. It is a flowchart showing an example of the basic operation of a smartphone. It is a schematic diagram showing an example of the configuration of a 3D model generation system (specific configuration). It is a flowchart showing an example of the operation of the 3D model generation system shown in FIG. 8. It is a block diagram showing an example of the configuration of a shooting position and orientation estimation unit. It is a block diagram showing an example of the configuration of a reliability calculation unit. It is a schematic diagram for explaining Sampson distance. It is a schematic diagram showing an example of calculation of a target position and orientation. It is a block diagram showing an example of the configuration of a 3D model generation unit. It is a schematic diagram for explaining an example of calculation of 3D coordinates by a point cloud integration unit. It is a schematic diagram showing another example of the configuration of a reliability calculation unit. It is a schematic diagram showing another example of the configuration of a reliability calculation unit. It is a block diagram showing an example of the hardware configuration of a computer that can be used as a smartphone or the like.

[0028] Hereinafter, embodiments of the present technology will be described with reference to the drawings.

[0029] [Outline of Photogrammetry] Figure 1 is a schematic diagram for explaining the outline of photogrammetry. Photogrammetry is a term that refers to the technology of measuring a target object using images. By using this technology, it is possible to generate a 3D model of a target object (real object) in real space and reconstruct the three-dimensional shape of the target object.

[0030] 1 , for example, a user 1 uses an imaging device 2 to capture multiple images of a target object 3. A 3D model of the target object 3 is generated based on the multiple captured images. Photogrammetry makes it possible to generate a 3D model of the target object 3 using the relatively simple method of capturing images of the target object 3 in real space.

[0031] Any device capable of capturing a two-dimensional image can be used as the imaging device 2. For example, any imaging device can be used, such as a digital camera equipped with an image sensor such as a complementary metal-oxide semiconductor (CMOS) sensor or a charge coupled device (CCD) sensor, a single-lens reflex camera, a smartphone, or a smart terminal.

[0032] For example, a user 1 can take a photograph of a target object with a familiar imaging device 2 such as a smartphone, and have a 3D model of the target object 3 displayed on the display of the smartphone, etc. Of course, by using a high-resolution single-lens reflex camera, etc., it is also possible to generate and display a very high-quality 3D model.

[0033] 1, a user 1 using an imaging device 2 captures multiple images of a car as a target object 3. Of course, the specific type of target object 3 is not limited, and a 3D model can be generated by photogrammetry for any object, including natural objects such as trees and utility poles, artificial objects, humans, and other animals.

[0034] 1, the shooting position (position of the shooting device 2) and shooting direction (direction of the shooting device 2) when a photograph is taken by the shooting device 2 are schematically shown by a white triangle 4. The triangle is an isosceles triangle, and the position of the vertex 4a of the isosceles triangle represents the shooting position. Furthermore, the direction of the perpendicular line from the vertex 4a to the base 4b represents the shooting direction.

[0035] Hereinafter, the shooting position and shooting orientation when a target object is photographed will be collectively referred to as the “shooting position and orientation.” For example, the shooting position and orientation at a certain shooting timing refers to the shooting position and shooting orientation at that shooting timing.

[0036] The photographing position and orientation of a photographed image refers to the photographing position and orientation when the photographed image was photographed. The photographing position and orientation can also be expressed as the camera position and orientation or the photographing device position and orientation.

[0037] As shown by the arrow 5 in Figure 1, generating a 3D model by photogrammetry requires captured images of the target object 3 taken at various shooting positions and orientations. For example, images of the target object 3 are captured 360 degrees around the target object 3. Images of the target object 3 are also captured from below and above the target object 3. A 3D model of the target object 3 is generated based on the multiple captured images.

[0038] 2 is a schematic diagram for explaining an overview of the photogrammetry algorithm. The input is a plurality of captured images 7 taken at various shooting positions and orientations of a target object 3. The shooting position and orientation are estimated for each of the input captured images 7.

[0039] The estimation of the photographing position and orientation is typically performed by SfM (Structure from Motion) processing, which calculates the geometric relationship between each of the photographed images 7 based on a common field of view from the plurality of photographed images 7, and makes it possible to estimate the photographing position and orientation of each of the photographed images.

[0040] In the present disclosure, the "common field of view" between images refers to an image area (pixel area) in which the same subject (the same part of the target object 3) is photographed. Furthermore, processing based on the "common field of view" is not limited to processing using the entire image area in which the same subject is photographed. For example, processing based on corresponding point pairs (e.g., pairs of pixels in which the same subject is photographed) that are feature points that correspond to each other and are included in the "common field of view" is also included in processing based on the "common field of view."

[0041] In the SfM process, the geometric relationship between the captured images 7 having a "common field of view" is calculated based on the corresponding point pairs described above.

[0042] 2, the shooting position and orientation of each captured image 7 is represented by a quadrangular pyramid 8. The position of the apex 8a of the quadrangular pyramid 8 represents the shooting position. The direction of the perpendicular line from the apex 8a to the base 8b represents the shooting direction. The base 8b of the quadrangular pyramid 8 corresponds to the angle of view of the captured image 7.

[0043] A 3D model 9 is generated based on the estimated shooting position and orientation of each captured image 7. Typically, MVS (Multi View Stereo) processing is executed to generate the 3D model 9. Through the MVS processing, depth information of the target object 3 is estimated in a triangulation manner based on the shooting position and orientation of each captured image 7. This makes it possible to define three-dimensional coordinates for each pixel at which the target object 3 is captured, and to generate a 3D model 9 made up of a three-dimensional point cloud. Note that the 3D model 9 is not limited to being composed of three-dimensional point cloud data, and the 3D model 9 may be composed of other data structures.

[0044] 3 is a schematic diagram for explaining a photographing motion when photographing a target object 3. In the present disclosure, the photographing motion refers to the movement (motion) of the photographing device 2 when the user 1 photographs the target object 3.

[0045] 3A, after taking a photograph at a predetermined photographing timing, user 1 moves the photographing position in one predetermined direction and takes the next photograph at the new photographing position. This photographing motion in which the photographing position is moved in one predetermined direction is called a translational photographing motion.

[0046] 3B, after taking a photograph at a predetermined photographing timing, user 1 rotates the photographing direction on the spot without changing the photographing position, and takes photographs multiple times. This photographing motion in which the photographing direction is rotated without changing the photographing position is called a pure rotation photographing motion.

[0047] Note that a shooting motion that rotates the shooting direction, regardless of whether the shooting position changes, is referred to as a rotational shooting motion. For example, in the example shown in Figure 3A, if the shooting direction is rotated when the second shooting is performed, the shooting motion will be a shooting motion that includes a translational shooting motion and a rotational shooting motion. Because the shooting motion includes a translational shooting motion, it is not a pure rotational shooting motion.

[0048] Therefore, the pure rotational shooting motion can also be said to be a rotational shooting motion that does not include a translational shooting motion.

[0049] The present inventors have conducted extensive research into how to generate a high-quality 3D model 9. As a result, they have discovered a new relationship between the SfM process, which is a process for estimating the shooting position and orientation shown in Fig. 2, and the shooting motion shown in Fig. 3.

[0050] 4 is a schematic diagram for explaining the relationship between the SfM processing and the quality of the generated 3D model 9. It has been newly discovered that, in terms of the SfM processing algorithm, when the target object 3 is photographed with the pure rotational photographing motion shown in FIG. 3B (or a photographing motion close to the pure rotational photographing motion), it is impossible to distinguish in the photographed images 7 whether the photographing was performed with a translational photographing motion or a pure rotational photographing motion, and this may result in a decrease in the accuracy of calculation of the geometric relationship between the photographed images 7.

[0051] A decrease in the accuracy of calculating the geometric relationships between the captured images 7 also reduces the quality of the 3D model generated by the MVS process. In the 3D model 9 of the square shown in Figure 4, distortion occurs in the circular area in the foreground. Specifically, a problem occurs in the positional relationship between the curved fence 11 and the stairs 12 inside the fence 11 (the far side of the fence 11 in the figure), resulting in a shape that differs from the actual streetscape.

[0052] 1, when outside-in photography is performed in which the target object 3 is located at the center of the photography environment and photography is performed while moving around the target object 3, pure rotational photography motion as shown in Fig. 3B is unlikely to occur. Therefore, degradation in quality of the 3D model 9 due to the influence of pure rotational photography motion is unlikely to occur.

[0053] On the other hand, as illustrated in Figure 4, there may be cases where user 1 is in a plaza or the like and takes pictures of the surrounding streetscape to generate a 3D model 9, or where user 1 is in a room and takes pictures of the surrounding room to generate a 3D model 9. In such cases, inside-out shooting is required, in which user 1 takes pictures while standing in the same position and looking around, which tends to result in a pure rotational shooting motion as shown in Figure 3B. As a result, the quality of the 3D model 9 may be reduced due to the influence of the pure rotational shooting motion.

[0054] It is often difficult for the user 1 photographing the target object 3 to always know whether or not they are performing a pure rotational shooting motion (or a shooting motion close to a pure rotational shooting motion). For example, they may be performing a pure rotational shooting motion (or a shooting motion close to a pure rotational shooting motion) without even realizing it.

[0055] Therefore, it is thought that for users 1 who are not particularly familiar with photogrammetry, taking photographs while being careful not to perform pure rotational photographing motions can often be a significant burden.

[0056] The present inventor has devised a new technique to solve the problem of quality degradation of 3D models caused by the influence of such shooting motion, which will be described below.

[0057] [Basic Configuration and Basic Operation of 3D Model Generation System] In describing the 3D model generation system according to the present technology, the basic configuration and basic operation of the 3D model generation system will be described first, followed by a detailed description of an embodiment for realizing the 3D model generation system.

[0058] 5 is a schematic diagram showing an example configuration of a 3D model generation system 14 according to this embodiment. In this embodiment, the 3D model generation system 14 is realized by an imaging device 2 formed of a single-lens reflex camera and a smartphone 15. The 3D model generation system 14 corresponds to an embodiment of an information processing system according to the present technology.

[0059] The photographing device 2 and the smartphone 15 are connected to each other so that they can communicate with each other. The connection between the two devices is not limited, and any connection form via wire or wireless may be adopted.

[0060] A liquid crystal monitor 16 is provided on the back side of the photographing device 2, and the user 1 can photograph the target object 3 while checking the target object 3 displayed on the liquid crystal monitor 16.

[0061] The smartphone 15 has hardware necessary for a computer, such as a processor such as a CPU, a GPU, or a DSP, a memory such as a ROM or a RAM, and a storage device such as an HDD (see FIG. 18 ). The processor loads a program according to the present technology stored in the storage unit or the memory into the RAM and executes the program, thereby realizing the information processing method according to the present technology (a reliability calculation method and a 3D model generation method).

[0062] The smartphone 15 functions as an embodiment of an information processing device according to the present technology. Of course, any computer other than a smartphone can be used, and any hardware such as an FPGA or an ASIC may be used.

[0063] Furthermore, in this embodiment, real-time generation of the 3D model 9 is realized by the smartphone 15. The real-time generation of the 3D model is a method of generating the 3D model 9 in real time each time a captured image 7 is acquired in response to photographing the target object 3. The generation of the real-time 3D model 9 can also be said to be the generation of a sequential 3D model 9. Of course, the application of the present technology is not limited to the generation of the real-time 3D model 9.

[0064] 5 , the smartphone 15 is installed so that the display (touch panel) 17 faces the user (the rear side of the camera). In this embodiment, a 3D model 9 of the target object 3, which is generated in real time in response to an image captured by the imaging device 2, is displayed on the display 17 of the smartphone 15.

[0065] The user 1 can decide to photograph an incomplete part while checking the 3D model 9. Also, guide information for photographing the target object 3 is displayed on the display 17. This point will be described later.

[0066] Fig. 6 is a block diagram showing an example of the functional configuration of the smartphone 15. Fig. 7 is a flowchart showing an example of the basic operation of the smartphone 15.

[0067] As shown in FIG. 6 , the smartphone 15 includes a captured image acquisition unit 19 , a capturing position and orientation estimation unit 20 , a reliability calculation unit 21 , a 3D model generation unit 22 , and an information presentation unit 23 .

[0068] These functional blocks are configured by executing a predetermined program of the present technology by the processor of the smartphone 15. Of course, dedicated hardware such as an IC (integrated circuit) may be used to realize each functional block.

[0069] The program is installed on the smartphone 15 via, for example, various recording media. Alternatively, the program may be installed via the Internet or the like. The type of recording medium on which the program is recorded is not limited, and any computer-readable recording medium may be used. For example, any computer-readable non-transitory storage medium may be used.

[0070] 7 , a new image is captured by the user 1. Then, the captured image acquisition unit 19 acquires captured images of the target object 3 in response to photographing of the target object 3 (step 101). For example, each time photographing of the target object 3 is performed at various photographing positions and postures, the captured images are acquired as new images by the captured image acquisition unit 19.

[0071] The new image acquired by the photographed image acquisition unit 19 is an embodiment of the "new photographed image" according to the present technology. The existing image that has already been acquired at a timing before the "new photographed image" is an embodiment of the "existing photographed image" according to the present technology.

[0072] The photographing position and orientation estimation unit 20 estimates the photographing position and orientation for each of the plurality of photographed images 7 photographed for generating a 3D model of the target object 3 in real space based on a common field of view with the other photographed images 7 (step 102). In this embodiment, the photographing position and orientation are estimated by executing SfM processing.

[0073] That is, each time a new image is acquired, the photographing position and orientation estimation unit 20 estimates the photographing position and orientation of the new image based on the common field of view with an existing image that has already been acquired and whose photographing position and orientation has been estimated.

[0074] For example, the shooting position and orientation estimation unit 20 selects an existing image that has corresponding feature points as a pair of corresponding points for the new image, and estimates the shooting position and orientation of the new image by calculating the geometric relationship between the new image and the existing image based on the pair of corresponding points.

[0075] It should be noted that it may be possible to estimate the shooting position and orientation of a new image by performing processing other than SfM processing.

[0076] The reliability calculation unit 21 calculates the reliability of the estimated photographing position and orientation regarding the influence on the estimation accuracy of the photographing position and orientation caused by the photographing motion when photographing the photographed image 7 (step 103). Specifically, the reliability calculation unit 21 calculates the reliability of the influence on the estimation accuracy of the photographing position and orientation caused by the pure rotational photographing motion.

[0077] The reliability calculation unit 21 calculates the consistency of the geometric relationship between each of the multiple captured images 7 and other captured images 7 that have a common field of view, and calculates the reliability based on the calculated consistency of the geometric relationship.

[0078] In this embodiment, the reliability calculation unit 21 calculates the reliability of a new image based on the consistency of the geometric relationship with an existing image that has a common field of view, has already had its reliability calculated, and has its reliability higher than a predetermined threshold. For example, the reliability can be calculated based on corresponding point pairs.

[0079] If the consistency of the geometric relationship is low with respect to an existing image that has a common field of view and has already been calculated with a high reliability, the new image is calculated to have a low reliability, assuming that the pure rotational shooting motion has a high impact on the estimation accuracy of the shooting position and orientation compared to the existing image.If the consistency of the geometric relationship is high, the new image is calculated to have a high reliability, assuming that the pure rotational shooting motion has a low impact on the estimation accuracy of the shooting position and orientation.Note that the specific value of the threshold for reliability is not limited and may be set appropriately.

[0080] Calculating the consistency of the geometric relationship between images can also be considered as calculating the inconsistency of the geometric relationship between images. It is possible to calculate the consistency of the geometric relationship between images using any algorithm that calculates not only the "consistency" but also the "inconsistency," and thus to calculate the reliability.

[0081] The 3D model generation unit 22 generates a 3D model 9 based on the reliability calculated for the shooting position and orientation of each of the multiple captured images 7 (existing images and new images) (step 104). In this embodiment, the 3D model 9 is generated by executing MVS processing. Generating the 3D model 9 based on the reliability makes it possible to generate a high-quality 3D model.

[0082] The information presenting unit 23 determines whether the reliability calculated for the new image is low. Specifically, it determines whether the calculated reliability is smaller than a predetermined threshold (step 105). If the reliability is larger than the predetermined threshold (NO in step 105), the reliability calculation flow ends and proceeds to step 108.

[0083] If the reliability is smaller than the predetermined threshold (YES in step 105 ), the information presenting unit 23 calculates the target position and orientation for photographing the target object 3 and presents it to the user 1 .

[0084] First, the information presentation unit 23 calculates a target position and orientation (step 106). The target position and orientation include a target shooting position, which is a target shooting position, and a target shooting orientation. The information presentation unit 23 calculates, as the target position and orientation, a target shooting position and a target shooting orientation that can suppress the influence of a pure rotational shooting motion on the estimation accuracy of the shooting position and orientation.

[0085] The target position and orientation is information about the shooting position and orientation presented to the user 1, and can also be called a presented shooting position and orientation. The target shooting position can also be called a presented shooting position or a presented spot. The target shooting orientation can also be called a presented shooting orientation.

[0086] Furthermore, the specific value of the threshold value relating to the reliability used in the threshold processing in step 105 is not limited, and may be set appropriately.

[0087] The information presenting unit 23 generates visualized information (step 107). The visualized information includes any image information that can be viewed by the user 1, such as various image information such as text, icons, and GUI. In this embodiment, the information presenting unit 23 generates, as visualized information, a GUI or the like that includes both the 3D model 9 generated in step 104 and the guide information including the target position and posture generated in step 106.

[0088] The generated visualization information is displayed on the display 17 of the smartphone 15. Information on the shooting position and shooting direction that can be taken is calculated.

[0089] It is determined whether or not the user 1 has taken additional photographs (step 108). If additional photographs have been taken (YES in step 108), the process returns to step 101. If additional photographs have not been taken (NO in step 108), the generation of the 3D model 9 ends.

[0090] For example, when a real-time 3D model 9 is generated, it is difficult for the user 1 to recognize whether or not a pure rotational shooting motion is being performed. For example, if the quality of the 3D model 9 displayed as visualization information is not very good, the user 1 may become anxious about how to proceed with subsequent shooting.

[0091] In this embodiment, while the user 1 is continuously photographing the target object 3 at various photographing positions and orientations, if the reliability calculated for a new image becomes low, a target position and orientation is presented.

[0092] The user 1 photographs the target object 3 according to the target position and orientation displayed on the display 17. This makes it possible to capture a new image in which the influence on the estimation accuracy of the photographing position and orientation caused by pure rotational photographing motion is suppressed. As a result, a high-quality 3D model 9 is generated and displayed on the display 17, allowing the user 1 to continue photographing with peace of mind.

[0093] [Example of a specific embodiment of a 3D model generation system] An example of a specific embodiment of a 3D model generation system according to the present technology will be described. Fig. 8 is a schematic diagram showing an example of the configuration of a 3D model generation system 25 according to this embodiment. In this embodiment, the 3D model generation system 25 is also configured by a single-lens reflex camera and a smartphone 15 shown in Fig. 14.

[0094] The 3D model generation system 25 includes a single-lens reflex camera (photographing device) 2 and a display 17 of a smartphone 15. The 3D model generation system 25 also includes functional blocks configured within the smartphone 15.

[0095] 8, the smartphone 15 includes a photographing position and orientation estimation unit 20, a reliability calculation unit 21, a 3D model generation unit 22, a target position and orientation calculation unit 26, and a visualization information generation unit 27. The smartphone 15 also includes a geometric information DB 28.

[0096] In the example shown in Fig. 8, the photographing position and orientation estimation unit 20 also functions as the photographed image acquisition unit 19 shown in Fig. 6. In addition, the target position and orientation calculation unit 26 and the visualization information generation unit 27 realize the information presentation unit 23 shown in Fig. 6.

[0097] Fig. 9 is a flowchart showing an example of the operation of the 3D model generation system 25 shown in Fig. 8. When the user 1 captures an image 7 of the target object 3, the image capturing position and orientation estimation unit 20 acquires a new image from the image capturing device 2 (step 201).

[0098] The photographing position and orientation estimation unit 20 estimates the photographing positions and orientations of all images including existing images (step 202). That is, in this embodiment, in response to acquisition of a new image, the photographing positions and orientations of not only the new image but also all images are estimated.

[0099] 10 is a block diagram showing an example of the configuration of the photographing position and orientation estimation unit 20. The photographing position and orientation estimation unit 20 includes a feature point matching unit 30, a relative position estimation unit 31, and an overall optimization unit 32.

[0100] The feature point matching unit 30 detects feature points in the new image and calculates feature amounts around the feature points by performing, for example, SIFT (Scale-Invariant Feature Transform) processing, etc. The feature point matching unit 30 also checks whether there are pairs of feature points having similar feature amounts between the new image and the existing image.

[0101] 10 , feature point information (2D coordinates and feature amounts of feature points) of feature points detected from the existing image is read from the geometric information DB 28 and compared with the feature amounts of feature points detected from the new image. Then, feature points with similar feature amounts are written as a feature point pair to the geometric information DB.

[0102] Specifically, the feature point information (2D coordinates and feature amounts) of the feature points detected from the new image and the feature point information (2D coordinates and feature amounts) of the feature points of the existing image that form the feature point pair are associated with each other and stored in the geometric information DB 28.

[0103] Of the feature points detected from the new image, feature point information of those that do not form feature point pairs with feature points in the existing image is also written to the geometric information DB 28. This is because even if a feature point does not form a feature point pair this time, there is a possibility that it may form a feature point pair with a feature point detected in a new image later.

[0104] The relative position estimation unit 31 calculates the relative position and orientation of the new image with respect to the existing image based on the feature points of the new image and the feature points of the existing image stored as feature point pairs by the feature point matching unit 30. The relative position and orientation between two images corresponds to the relative geometric relationship between the two images.

[0105] Specifically, true feature point pairs representing the same subject are first detected from among the feature point pairs. That is, a noise removal process is performed to treat feature point pairs representing different subjects as noise. For example, by using a RANSAC (Random Sample Consensus) process or the like, feature point pairs that are geometrically consistent based on the provisional relative positions and orientations can be extracted as true feature point pairs.

[0106] Hereinafter, the extracted true feature point pair will be referred to as an inlier pair. The inlier pair is one embodiment of a corresponding point pair, which is feature points that correspond to each other, according to the present technology.

[0107] The 3D coordinates of the feature points of the existing image that forms an inlier pair are read from the geometric information DB 28. The 3D information is calculated by the subsequent global optimization unit 32 and written to the geometric information DB 28 when the existing image is acquired as a new image.

[0108] The relative position estimation unit 31 calculates the relative position and orientation of the new image with respect to the existing image for the inlier pair based on the 3D coordinates of the feature points of the existing image, for example, by using a PnP (Perspective-n-Point) algorithm, etc. The calculated relative position and orientation are stored in the geometric information DB 28.

[0109] The global optimization unit 32 calculates the shooting positions and orientations of all images, including the new image and the existing images, through optimization processing such as bundle adjustment processing. The optimization processing updates the shooting positions and orientations of the existing images. The global optimization unit 32 then uses the shooting positions and orientations of all images to calculate the 3D coordinates of feature points of all images using the triangulation capacity.

[0110] 10 , the relative position and orientation calculated for the new image, and the shooting positions and orientations and 3D coordinates of feature points calculated for the previous frame for the existing images are input to the global optimization processing unit 32. Optimization processing is performed based on this information, and the 3D coordinates of feature points for all images are calculated. The calculated shooting positions and orientations and 3D coordinates of feature points for all images are stored in the geometric information DB 28.

[0111] Furthermore, the photographing positions and orientations of all images calculated by the overall optimization unit 32 are output to the 3D model generation unit 22 .

[0112] When the shooting positions and orientations of all images are estimated in step 202, the reliability of the shooting positions and orientations of the new images is calculated by the reliability calculation unit 21. In this embodiment, the reliability is calculated by determining whether or not there is any inconsistency in the geometric relationship even though the new images have a sufficient common field of view.

[0113] 11 is a block diagram showing an example of the configuration of the reliability calculation unit 21. The reliability calculation unit 21 includes a common-view detection unit 34, a geometric mismatch detection unit 35, and a reliability integration unit 36.

[0114] The common-view detection unit 34 first selects an anchor image. In this embodiment, of existing images whose reliability has already been calculated and whose reliability is higher than a predetermined threshold, the existing image whose capture timing is closest to that of the new image is selected as the anchor image. In other words, the captured image whose capture timing is close in time to that of the new image and whose reliability calculated when the captured image was acquired as the new image is high is selected as the anchor image. Note that the specific value of the threshold value related to the reliability that serves as the criterion for selecting the anchor image is not limited and may be set as appropriate.

[0115] The common field of view detection unit 34 detects whether or not there is a common field of view between the new image and the anchor image (step 203). In this embodiment, the following detection method is used to detect, as a binary value, whether or not there is a sufficient common field of view between the new image and the anchor image.

[0116] The number of inlier pairs between the new image and the anchor image > a predetermined threshold, and the size of the rectangle spanned by the feature points that are inlier pairs in the new image > a predetermined threshold.

[0117] If the condition is met, the detection result of the common field of view is set to true. If the condition is not met, the detection result of the common field of view is set to false. Note that the specific values ​​of the threshold for the number of inlier pairs and the threshold for the size of the rectangle are not limited and may be set as appropriate.

[0118] The geometric mismatch detection unit 35 calculates the degree of match (geometric reliability) of the geometric relationship between the new image and the anchor image (step 204).

[0119] In this embodiment, the Sampson distance is used to calculate the consistency of the geometric relationship. Fig. 12 is a schematic diagram for explaining the Sampson distance.

[0120] The Sampson distance indicates whether a pair of feature points in two images exists on the same epipolar plane, and is an index for determining the degree of consistency of the geometric relationship. The closer the Sampson distance is to zero, the more accurate the geometric relationship between the two images is, and the higher the degree of consistency of the geometric relationship.

[0121] Images A and B shown in Figure 12 have a correct geometric relationship. Feature point P in image A is a 3D coordinate point projected onto image A. Feature point P' in image B is the same 3D coordinate point projected onto image B. Since images A and B have a correct geometric relationship, the feature point pair (P, P') exists on the epipolar plane. That is, in each of images A and B, feature points P and P' exist on the epipolar line. In this case, the Sampson distance is zero.

[0122] The lower the degree of consistency (i.e., the greater the degree of inconsistency) in the geometric relationship between the inlier pair (pair of feature points) of the new image and the anchor image, the larger the Sampson distance.

[0123] In this embodiment, in order to detect geometric inconsistencies due to the influence of pure rotational shooting motion, the variance value of the Sampson distance with respect to rotation is calculated, and the consistency of the geometric relationship (geometric reliability) is calculated by taking the reciprocal of that variance value.

[0124] The variance value for rotation of the Sampson distance can be calculated using the following formula, in accordance with the paper Revisiting Rotation Averaging: Uncertainties and Robust Losses by Ganlin Zhang, et al. (2023 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR2023)).

[0125]

[0126] The variables in equation (1) are as follows (see FIG. 10): p: 2D coordinates of the feature points of image A; K: internal parameters of camera A that captured image A; P': 2D image of the feature points of image B; K': internal parameters of camera B that captured image B; t: relative position of image A with respect to image B; R: relative orientation of image A with respect to image B; D: rotation direction component of the geometric reliability of image A with respect to image B (so3 representation 3 elements)

[0127] The geometric mismatch detection unit 35 calculates the geometric reliability as a scalar value by taking the norm of the geometric reliability calculated from the variance value in the rotation direction, as shown in the following formula: Geometric reliability = |D|

[0128] The calculated geometric reliability is output to the reliability integration unit 36 ​​.

[0129] The reliability integration unit 36 ​​calculates the reliability of the photographing position and orientation of the new image based on the detection results from the common view detection unit 34 and the detection results from the geometric mismatch detection unit 35 (step 205).

[0130] If there is a sufficient common field of view with respect to the anchor image, but there is a rotational geometric mismatch (i.e., low geometric confidence), it can be determined that there is a pure rotational photographic motion that adversely affects the photogrammetry process.

[0131] If the common field of view detection result is true, the reliability integration unit 36 ​​sets the reliability to geometric reliability. On the other hand, if the common field of view detection result is false, an invalid value is output as the reliability. In other words, if there is no common field of view with the anchor image, it is not possible to determine whether or not there is pure rotational shooting motion, and the calculation flow of reliability due to the influence of pure rotational shooting motion ends.

[0132] As described above, in this embodiment, when a common field of view exists, the reliability calculation unit 21 calculates the consistency of the geometric relationship with the anchor image as the reliability.

[0133] The target position and orientation calculation unit 26 determines whether the calculated reliability is smaller than a predetermined threshold (step 206). If the reliability is smaller than the predetermined threshold, the target position and orientation for photographing the target object 3 is calculated (step 207).

[0134] 13 is a schematic diagram showing an example of calculation of a target position and orientation. In the example shown in Fig. 13, a new image is captured at spot C at time t = T + 1. Also, an image captured at spot B immediately before that at time t = T is selected as the anchor image.

[0135] As shown in Figure 13, when user 1 captures a new image following an anchor image, a shooting motion similar to a pure rotation shooting motion is performed. In this embodiment, it is possible to detect the influence of the pure rotation shooting motion on the estimation accuracy of the shooting position and orientation. In other words, the reliability of the shooting position and orientation estimated for the new image captured at spot C is low. As a result, in step 207, the target position and orientation is calculated.

[0136] As the target posture, a shooting position and a shooting orientation that can suppress the influence on the estimation accuracy of the shooting position and orientation caused by a pure rotation shooting motion are calculated. In this embodiment, the target position and orientation are calculated based on the shooting position and orientation of the anchor image. That is, when the shooting position and orientation of the anchor image are used as a reference, a target shooting position and a target shooting orientation that do not result in a pure rotation shooting motion are calculated.

[0137] As shown in Fig. 13, for example, a shooting position moved a predetermined distance in one direction from the shooting position (spot B) of the anchor image included in the shooting position and orientation of the anchor image is calculated as the target shooting position. In this way, by adding a translational shooting motion of a predetermined distance, it is possible to suppress the influence of pure rotational shooting motion. For example, even if the shooting orientation when a new image is shot is maintained as it is, a translational movement is added, so the influence of pure rotational shooting motion is suppressed.

[0138] 13 is presented, it is not a particularly large burden for user 1 because all that is required is to move the shooting position while keeping the shooting orientation the same. Of course, the shooting orientation does not have to be maintained, and the shooting orientation may be changed.

[0139] Regarding one direction of translational movement from the shooting position (spot B) of the anchor image, it is possible to use the direction from spot A, which is the shooting position of the existing image taken at time t = T-1 immediately before the anchor image (the shooting position included in the shooting position and orientation of the existing image taken immediately before), toward the shooting position of the anchor image.

[0140] This makes it possible to easily set a predetermined direction for calculating the target shooting position. Also, for the user 1, this movement is natural compared to the movement while shooting as in the past, so it is stress-free.

[0141] It is also possible to set the target imaging position as a position obtained by adding the amount of movement from spot A to spot B to spot B. This makes it possible to more easily calculate the target imaging position.

[0142] Regarding the calculation of the target photographing position, it is also possible to specify the distance from the photographing position (spot B) of the anchor image to the target photographing position. For example, the target photographing position is calculated so that the distance from the photographing position (spot B) of the anchor image to the target photographing position is greater than a preset minimum movement distance and less than a preset maximum movement distance.

[0143] The maximum movement distance will now be described. In this 3D model generation system 25, a new image captured at spot C shown in Fig. 13 is an image with low reliability. Therefore, in relation to the anchor image with high reliability, the new image has low geometric reliability.

[0144] In this situation, suppose that an image is captured near spots B and C that has a common field of view with both the anchor image and the new image and does not have a pure rotational shooting motion with respect to the anchor image. In this case, the captured image has a high geometric reliability with respect to the anchor image, and the reliability calculation unit 21 calculates a high reliability for the captured image.

[0145] By having a common field of view and acquiring captured images with high geometric reliability relative to the anchor image, for example, a portion captured in a new image with low reliability can be re-captured in a captured image with high reliability. As a result, it is possible to improve the calculation accuracy of the 3D coordinates of feature points that form inlier pairs calculated in the global optimization.

[0146] Considering this point of view, it is possible to define the maximum movement distance within a range in which a common field of view can be detected for the anchor image and the new image. Specifically, depending on the shooting environment and the type and size of the object to be shot, it is thought that a distance of, for example, several tens of centimeters to several meters can be adopted. Of course, the maximum movement distance is not limited to this range.

[0147] Regarding the minimum distance, if the distance from the imaging position (spot B) of the anchor image is too small (for example, 1 cm), it is difficult to suppress the influence of the pure rotational imaging motion. It is considered that a translational movement of at least several tens of cm or more is necessary.

[0148] For example, it is possible to adopt settings such that the minimum movement distance is several tens of centimeters and the maximum movement distance is several tens of centimeters to several meters (a value greater than the initial movement distance).

[0149] For example, it is assumed that the target photographing position is the position obtained by adding the amount of movement from spot A to spot B to spot B as described above.

[0150] In this case, if the distance from the anchor image to the target shooting position is greater than the maximum movement distance, the position that is the maximum movement distance from the anchor image is set as the target shooting position, whereas if the distance from the anchor image to the target shooting position is less than the minimum movement distance, the position that is the minimum movement distance from the anchor image is set as the target shooting position.

[0151] This makes it possible to sufficiently suppress the influence of pure rotational shooting motion on the estimation accuracy of the shooting position and orientation, thereby enabling the generation of high-quality 3D models.

[0152] Of course, the calculation of the target shooting position is not limited to the example shown in Fig. 13. Any target shooting position may be set as long as it is a target shooting position that can suppress the influence of pure rotational shooting motion, that is, a target shooting position that does not result in pure rotational shooting motion when the shooting position and orientation of the anchor image are used as a reference.

[0153] The 3D model generation unit 22 generates a 3D model 9 based on the shooting positions and orientations and the reliability of all images (step 208). In this embodiment, the 3D model generation unit 22 executes MVS processing to calculate depth information, and generates a dense 3D point cloud as a 3D model.

[0154] 14 is a block diagram showing an example of the configuration of the 3D model generation unit 22. The 3D model generation unit 22 includes a depth image generation unit 38 and a point cloud integration unit 39.

[0155] The depth image generating unit 38 calculates the depth for each pixel of each image using the photographing position and orientation in a triangulation manner, and generates a depth image for each image.

[0156] The point cloud integration unit 39 collects 3D points obtained from the depth image and calculates the 3D coordinates of the same points using averaging or median processing.

[0157] FIG. 15 is a schematic diagram illustrating an example of calculation of 3D coordinates by the point cloud integration unit 39. For example, suppose that averaging is performed on 3D points determined in each depth image. In this case, as shown in FIG. 15 , if the center of gravity of the 3D point cloud of the same point is shifted due to the influence of 3D points calculated from a captured image (depth image) with low reliability calculated in step 205, the accuracy of the 3D coordinates will decrease. Therefore, in this embodiment, the 3D coordinates are calculated by weighted averaging or weighted median processing shown in the following equations, using the reliability as a weight.

[0158]

[0159] The variables in equation (2) are as follows: p: 3D coordinate w: reliability i: image index The weighted median process is performed after sorting for each 3D coordinate.

[0160] Note that a threshold value may be set for the reliability, and a pruning process may be performed so that reliability below the threshold is zero. Furthermore, if the detection result of the common field of view with the anchor image is false in step 203, an invalid value is output as the reliability. For captured images for which an invalid value is output as the reliability, a predetermined value may be assigned as the reliability when performing this point cloud integration process. For example, the reliability may be set to 0, or conversely, the 3D coordinates may be calculated with the reliability set to 1. Note that the specific value of the threshold that serves as the basis for the pruning process is not limited and may be set arbitrarily.

[0161] The generation of the 3D model 9 using the confidence measure makes it possible to denoise the 3D point cloud.

[0162] The visualization information generating unit 27 generates visualization information (image information) including both the generated 3D model 9 and guide information including the target position and posture, and outputs it to the display 17 (step 209).

[0163] Fig. 16 is a schematic diagram showing an example of display of guide information. As shown in Fig. 16A, using a 3D view, it is possible to display the target shooting position and target position and posture from a first-person perspective as seen by the user 1 using a 3D model 9 and a view cone 41. Note that although Fig. 16 shows the view cone 41 in a triangular shape, it is also possible to display a 3D figure of a square pyramid as the view cone 41.

[0164] 16A is the target position and orientation. The white view cone 41b is the shooting position and orientation of the image 7 captured in the past (this can also be considered as historical information of the shooting position and orientation). For example, the shooting position and orientation of the image 7 selected as the anchor image in the past may be displayed.

[0165] Furthermore, the view cone 41a of the target position and orientation is not necessarily visible from the first-person viewpoint. In order to present the target position and orientation without omission, an overhead view may be displayed as shown in FIG. 16B. To identify the up-down direction when generating the overhead view, for example, by collecting the shooting directions of all images and performing PCA processing, it is possible to set the axis in the up-down direction (Z direction) with the smallest variance value.

[0166] The display of the first-person perspective view shown in Fig. 16A and the display of the bird's-eye view shown in Fig. 16B may be switchable. Furthermore, the two views may be displayed simultaneously on the same screen. For example, when the first-person perspective view is selected, the bird's-eye view may be displayed in a small size in the upper right corner of the display 17. Furthermore, when the bird's-eye view is selected, the first-person perspective view may be displayed in a small size in the upper right corner of the display 17.

[0167] 5 , the user 1 photographs the target object 3 while checking the LCD monitor 16. Every time the user 1 photographs the target object 3, the photographed image is acquired as a new image, and the 3D model 9 is displayed in real time on the display 17 of the smartphone 15.

[0168] If the reliability of the new image is low, a target position and orientation (view cone 41 a) is displayed on the display 17 of the smartphone 15. The user 1 changes the shooting position and orientation using the target position and orientation as a guide to capture images. This allows for the continuous capture of highly reliable images. As a result, it is possible to suppress the impact of pure rotational shooting motion on the estimation accuracy of the shooting position and orientation, making it possible to generate a high-quality 3D model.

[0169] As described above, in the 3D model generation system according to this embodiment, for each of a plurality of captured images 7 taken to generate a 3D model 9 of a target object 3 in real space, the shooting position and orientation are estimated based on a common field of view with the other captured images 7. Furthermore, for the estimated shooting position and orientation, a reliability is calculated regarding the influence of shooting motion on the estimation accuracy of the shooting position and orientation. Based on this reliability, it becomes possible to generate a high-quality 3D model 9.

[0170] Photogrammetry has been attracting attention as a method for generating digital twins of the real world. As mentioned above, due to the characteristics of the algorithm, photogrammetry is not good at capturing motions that are close to pure rotation, and there is a risk that distortion will occur in the generated 3D model 9. By implementing this technology, it is possible to automatically detect shooting motions that are not good at capturing images, and to present alternative shooting spots to the user 1 using a GUI or the like.

[0171] This reduces the risk of distortion of the generated 3D model 9. Furthermore, since the target position and orientation can be presented, the user 1 can take photos with confidence even if they are not skilled in photogrammetry photography. Furthermore, since the 3D model 9 is generated based on the reliability, the accuracy of the 3D model itself can be improved.

[0172] Other Embodiments The present technology is not limited to the above-described embodiments, and various other embodiments can be realized.

[0173] FIG. 17 is a schematic diagram showing another example configuration of the reliability calculation unit 21. For example, an IMU 43 is often installed in the image capture device 2, a smartphone, or the like. In this case, it is possible to directly detect the position and orientation information of the image capture device 2 based on the sensing results (acceleration and angular velocity) of the IMU 43. Of course, if a GPS sensor is installed, it is also possible to use the position information of the image capture device 2 based on the GPS signal. It is also possible to use the position information of the image capture device obtained by executing SLAM on the captured image information.

[0174] That is, sensors, algorithms, etc. that can directly detect the photographing position, orientation, and photographing motion of the photographing device 2 may be used in combination as appropriate.

[0175] 17, the reliability calculation unit 21 is configured with a pure rotation detection unit 44. If the amount of translational movement is equal to or less than a threshold and the amount of rotation is equal to or greater than a threshold, the pure rotation detection result is set to true. If these conditions are not met, the pure rotation detection result is set to false.

[0176] If there is a sufficient common field of view for the anchor image, but there is a rotational geometric misalignment, it can be determined that there is a pure rotational photographic motion that adversely affects the photogrammetry process.

[0177] To perform this determination, the reliability integration unit 36 ​​sets the reliability to geometric reliability when the common field of view detection result is true or the pure rotation detection result is true. On the other hand, when the common field of view detection result is false and the pure rotation detection result is false, an invalid value is output as the reliability. This makes it possible to determine pure rotation shooting motion with high accuracy.

[0178] In the above embodiment, the variance value related to the rotation of the Sampson distance is calculated, and the consistency of the geometric relationship (geometric reliability) is calculated by taking the inverse of the variance value. However, this is not limiting, and for example, the inverse of the Sampson distance may be calculated as the consistency of the geometric relationship (geometric reliability).

[0179] The guide information including the target position and posture may be presented by audio output or by tactile sensation such as vibration. For example, it is possible to set a predetermined audio or tactile sensation to be output as the target position and posture is approached. Also, an instruction such as "Please move a little further" may be output by image, audio, or the like.

[0180] In the above embodiment, real-time generation of a 3D model has been exemplified. However, the present technology is not limited to this, and can also be applied to batch generation of a 3D model 9 in which a 3D model is generated after all captured images are collected. For example, calculation of reliability and presentation of a target position and orientation according to the present technology may be performed when a user captures all captured images. Furthermore, when the 3D model is finally generated, generation of the 3D model based on the reliability may be performed.

[0181] In each process described in the above embodiment, any machine learning algorithm using, for example, a DNN (Deep Neural Network), an RNN (Recurrent Neural Network), or a CNN (Convolutional Neural Network) may be used. For example, by using AI (artificial intelligence) that performs deep learning, each process can be executed with high accuracy. Of course, the application of a machine learning algorithm may be executed for any process within the present disclosure.

[0182] FIG. 18 is a block diagram showing an example of the hardware configuration of a computer 60 that can be used as the smartphone 15 or the like.

[0183] The computer 60 includes a CPU 61, a ROM 62, a RAM 63, an input / output interface 65, and a bus 64 interconnecting these components. The input / output interface 65 is connected to a display unit 66, an input unit 67, a storage unit 68, a communication unit 69, a drive unit 70, and other components. The display unit 66 is a display device using, for example, an LCD or EL display. The input unit 67 is a keyboard, a pointing device, a touch panel, or other operating device. If the input unit 67 includes a touch panel, the touch panel may be integrated with the display unit 66. The storage unit 68 is a non-volatile storage device such as a HDD, flash memory, or other solid-state memory. The drive unit 70 is a device capable of driving a removable storage medium 71 such as an optical storage medium or magnetic recording tape. The communication unit 69 is a modem, router, or other communication device connectable to a LAN, WAN, or the like for communicating with other devices. The communication unit 69 may communicate via either a wired or wireless connection. The communication unit 69 is often used separately from the computer 60. Information processing by the computer 60 having the above-described hardware configuration is realized by cooperation between software stored in the storage unit 68 or the ROM 62, etc. and the hardware resources of the computer 60. Specifically, the information processing method according to the present technology is realized by loading a program constituting the software stored in the ROM 62, etc., into the RAM 63 and executing it. The program is installed in the computer 60 via, for example, the recording medium 71. Alternatively, the program may be installed in the computer 60 via a global network, etc. Alternatively, any computer-readable, non-transitory storage medium may be used.

[0184] The information processing method (reliability calculation method, 3D model generation method) and program according to the present technology may be executed by cooperation between multiple computers connected to each other via a network or the like, thereby constructing an information processing system or information processing device according to the present technology. In other words, the information processing method and program according to the present technology can be executed not only in a computer system composed of a single computer, but also in a computer system in which multiple computers operate in conjunction with each other. In this disclosure, a "system" refers to a collection of multiple components (devices, modules (parts), etc.), regardless of whether all the components are located in the same housing. Therefore, both multiple devices housed in separate housings and connected via a network and a single device in which multiple modules are housed in a single housing are systems.

[0185] The execution of the information processing method and program according to the present technology by a computer system includes both cases where, for example, estimation of the image capture position and orientation, calculation of reliability, calculation of the consistency of geometric relationships, calculation of the target position and orientation, detection of a common field of view, and generation of a 3D model are performed by a single computer, and cases where each process is performed by a different computer. Furthermore, the execution of each process by a specific computer also includes having another computer execute part or all of the process and obtaining the results. In other words, the information processing method and program according to the present technology can also be applied to a cloud computing configuration in which a single function is shared and processed collaboratively by multiple devices via a network.

[0186] The 3D model generation system, the imaging device, the smartphone, the functional blocks, the configurations of the GUI for guide information, the estimation of the imaging position and orientation, the calculation of reliability, the calculation of the consistency of geometric relationships, the calculation of the target position and orientation, the detection of a common field of view, the generation of a 3D model, and other processing flows described with reference to the drawings are merely one embodiment and can be modified as desired without departing from the spirit of the present technology. In other words, any other configurations, algorithms, etc. for implementing the present technology may be adopted.

[0187] In this disclosure, terms such as "about," "approximately," "almost," and "roughly" may be used as appropriate to facilitate understanding of the description. However, there is no clear difference between using and not using terms such as "about," "approximately," "almost," and "approximately." In other words, in this disclosure, concepts that define shape, size, positional relationship, state, etc., such as "center," "middle," "uniform," and "equal," are concepts that include "substantially center," "substantially central," "substantially uniform," and "substantially equal." For example, states that fall within a predetermined range (e.g., a range of ±10%) based on "completely centered," "completely central," "completely uniform," and "completely equal" are also included. Therefore, even if terms such as "approximately," "almost," and "approximately" are not used, concepts expressed by adding "approximately," "almost," and "approximately" may be included. Conversely, states expressed by adding terms such as "approximately," "almost," and "approximately" do not necessarily exclude perfect states.

[0188] In the present disclosure, expressions using "than", such as "greater than A" and "smaller than A", are expressions that comprehensively include both concepts that include the case where it is equivalent to A and concepts that do not include the case where it is equivalent to A. For example, "greater than A" is not limited to cases that do not include equivalent to A, but also includes "A or greater". Furthermore, "smaller than A" is not limited to "less than A" but also includes "A or less". When implementing the present technology, specific settings and the like can be appropriately adopted from the concepts included in "greater than A" and "smaller than A" so that the effects described above can be achieved.

[0189] It is also possible to combine at least two of the features of the present technology described above. That is, the various features described in each embodiment may be arbitrarily combined without distinguishing between the embodiments. Furthermore, the various effects described above are merely examples and are not intended to be limiting, and other effects may also be achieved.

[0190] Note that the present technology can also be configured as follows. (1) An information processing system including: an estimation unit that estimates a shooting position and orientation for each of a plurality of captured images taken for generating a 3D model of a target object in real space based on a common field of view with other captured images; and a reliability calculation unit that calculates, for the estimated shooting position and orientation, a reliability regarding an influence on estimation accuracy of the shooting position and orientation caused by shooting motion when capturing the captured image. (2) The information processing system described in (1), in which the reliability calculation unit calculates, for each of the plurality of captured images, a consistency of a geometric relationship with the other captured images having the common field of view, and calculates the reliability based on the calculated consistency of the geometric relationship. (3) The information processing system described in (1) or (2), in which the shooting motion is a pure rotational shooting motion. (4) The information processing system according to any one of (1) to (3), further comprising: an acquisition unit that acquires a captured image of the target object in response to shooting of the target object, wherein the estimation unit estimates the capturing position and orientation of the new captured image based on the common field of view with an existing captured image that has already been acquired and whose capturing position and orientation have been estimated, each time a new captured image is acquired, and the reliability calculation unit calculates the reliability for the capturing position and orientation of the new captured image. (5) The information processing system according to (4), wherein the estimation unit selects the existing captured image having corresponding feature points as a corresponding point pair with respect to the new captured image, and estimates the capturing position and orientation of the new captured image by calculating a geometric relationship of the new captured image with respect to the existing captured image based on the corresponding point pair. (6) The information processing system according to (5), wherein the reliability calculation unit calculates the reliability of the new captured image based on the consistency of the geometric relationship with the existing captured image that has the common field of view, has already had its reliability calculated, and has its reliability higher than a predetermined threshold.(7) The information processing system according to (6), wherein the reliability calculation unit calculates, using the corresponding point pair as a reference, a degree of consistency of the geometric relationship with the existing captured image whose reliability is higher than a predetermined threshold. (8) The information processing system according to (6) or (7), wherein the reliability calculation unit selects, as an anchor image, an existing captured image whose capture timing is closest to that of the new captured image from among the existing captured images whose reliability has already been calculated and whose reliability is higher than a predetermined threshold, and calculates the reliability based on whether or not there is a common field of view with the anchor image and the degree of consistency of the geometric relationship with the anchor image. (9) The information processing system according to (8), wherein, when there is a common field of view, the reliability calculation unit calculates the degree of consistency of the geometric relationship with the anchor image as the reliability. (10) The information processing system according to any one of (1) to (9), further comprising an information presentation unit that calculates a target position and orientation for photographing the target object and presents the target position and orientation to a user when the reliability is smaller than a predetermined threshold. (11) The information processing system according to (8) or (9), further comprising an information presentation unit that calculates a target position and orientation for photographing the target object and presents the target position and orientation to a user when the reliability is smaller than a predetermined threshold, wherein the information presentation unit calculates the target position and orientation based on the photographing position and orientation of the anchor image. (12) The information processing system according to (11), wherein the target position and orientation include a target photographing position and a target photographing orientation, and the information presentation unit calculates the target photographing position and the target photographing orientation based on the photographing position and orientation of the anchor image so as not to result in a pure rotational photographing motion. (13) The information processing system according to (12), wherein the information presenting unit calculates, as the target shooting position, a shooting position that is moved a predetermined distance in one direction from the shooting position of the anchor image included in the shooting position and orientation of the anchor image.(14) The information processing system according to (13), wherein the one direction is a direction from a shooting position of the existing captured image captured immediately before the anchor image, which is included in the shooting position and attitude of the existing captured image captured immediately before the anchor image, to the shooting position of the anchor image. (15) The information processing system according to (14), wherein the information presenter calculates the target position so that a distance from the shooting position of the anchor image to the target position is greater than a predetermined minimum movement distance and less than a predetermined maximum movement distance. (16) The information processing system according to any one of (1) to (15), further comprising a generator that generates a 3D model of the target object based on the plurality of captured images, the shooting positions and attitudes of each of the plurality of captured images, and the reliability calculated for the shooting positions and attitudes. (17) The information processing system according to any one of (1) to (16), wherein the estimation unit estimates the shooting position and orientation by executing SfM (Structure from Motion) processing, and the generation unit generates the 3D model by executing MVS (Multi View Stereo) processing. (18) An information processing method executed by a computer system to estimate a shooting position and orientation for each of a plurality of captured images taken for generating a 3D model of a target object in real space, based on a common field of view with other captured images, and calculate, for the estimated shooting position and orientation, a reliability regarding an influence on the estimation accuracy of the shooting position and orientation caused by shooting motion when capturing the captured image. (19) A program causing a computer system to estimate a shooting position and orientation for each of a plurality of captured images taken for generating a 3D model of a target object in real space, based on a common field of view with other captured images, and calculate, for the estimated shooting position and orientation, a reliability regarding an influence on the estimation accuracy of the shooting position and orientation caused by shooting motion when capturing the captured image.

[0191] REFERENCE SIGNS LIST 1... User 2... Photographing device 3... Target object 7... Photographed image 9... 3D model 14, 25... 3D model generation system 15... Smartphone 17... Display 41... View cone 60... Computer

Claims

1. An information processing system comprising: an estimation unit that estimates the shooting position and orientation for each of a plurality of captured images taken to generate a 3D model of a target object in real space based on a common field of view with other captured images; and a reliability calculation unit that calculates the reliability of the estimated shooting position and orientation regarding the influence on the estimation accuracy of the shooting position and orientation caused by the shooting motion when taking the captured images.

2. An information processing system according to claim 1, wherein the reliability calculation unit calculates, for each of the plurality of captured images, the degree of consistency of the geometric relationship with the other captured images that have the common field of view, and calculates the reliability based on the calculated degree of consistency of the geometric relationship.

3. An information processing system according to claim 1, wherein the shooting motion is a pure rotational shooting motion.

4. An information processing system according to claim 1, further comprising an acquisition unit that acquires a photographed image of the target object in response to photographing the target object, wherein the estimation unit, each time a new photographed image is acquired, estimates the photographing position and orientation of the new photographed image based on the common field of view with an existing photographed image that has already been acquired and whose photographing position and orientation have been estimated, and the reliability calculation unit calculates the reliability of the photographing position and orientation of the new photographed image.

5. An information processing system according to claim 4, wherein the estimation unit selects an existing captured image that has corresponding feature points as a pair of corresponding points for the new captured image, and estimates the shooting position and orientation of the new captured image by calculating the geometric relationship between the new captured image and the existing captured image based on the pair of corresponding points.

6. An information processing system according to claim 5, wherein the reliability calculation unit calculates the reliability of the new captured image based on the consistency of the geometric relationship with the existing captured image that has the common field of view, has already had its reliability calculated, and has its reliability higher than a predetermined threshold.

7. An information processing system according to claim 6, wherein the reliability calculation unit calculates the consistency of the geometric relationship with the existing captured image whose reliability is higher than a predetermined threshold value, using the corresponding point pair as a reference.

8. An information processing system according to claim 6, wherein the reliability calculation unit selects, as an anchor image, from among the existing photographed images for which the reliability has already been calculated and the reliability is higher than a predetermined threshold, the existing photographed image whose photographing timing is closest to that of the new photographed image, and calculates the reliability based on the presence or absence of a common field of view with the anchor image and the degree of consistency of the geometric relationship with the anchor image.

9. An information processing system according to claim 8, wherein the reliability calculation unit calculates the degree of consistency of the geometric relationship with the anchor image as the reliability when the common field of view exists.

10. An information processing system according to claim 1, further comprising an information presentation unit that, when the reliability is smaller than a predetermined threshold, calculates a target position and orientation for photographing the target object and presents it to a user.

11. An information processing system according to claim 8, further comprising an information presentation unit that, when the reliability is smaller than a predetermined threshold, calculates a target position and orientation for photographing the target object and presents the calculated target position and orientation to a user, wherein the information presentation unit calculates the target position and orientation based on the photographing position and orientation of the anchor image.

12. An information processing system according to claim 11, wherein the target position and orientation include a target shooting position and a target shooting direction, and the information presentation unit calculates the target shooting position and the target shooting direction based on the shooting position and orientation of the anchor image so as not to result in a pure rotational shooting motion.

13. An information processing system according to claim 12, wherein the information presentation unit calculates, as the target shooting position, a shooting position that is moved a predetermined distance in one direction from the shooting position of the anchor image included in the shooting position and orientation of the anchor image.

14. An information processing system according to claim 13, wherein the one direction is a direction from the shooting position of the existing image captured immediately before the anchor image, which is included in the shooting position and orientation of the existing image captured immediately before the anchor image, toward the shooting position of the anchor image.

15. An information processing system according to claim 14, wherein the information presentation unit calculates the target position so that the distance from the shooting position of the anchor image to the target position is greater than a predetermined minimum movement distance and less than a predetermined maximum movement distance.

16. An information processing system according to claim 1, further comprising: a generation unit that generates a 3D model of the target object based on the plurality of captured images, the capturing positions and orientations of each of the plurality of captured images, and the reliability calculated for the capturing positions and orientations.

17. An information processing system according to claim 1, wherein the estimation unit estimates the shooting position and orientation by executing SfM (Structure from Motion) processing, and the generation unit generates the 3D model by executing MVS (Multi View Stereo) processing.

18. An information processing method implemented by a computer system, which estimates the shooting position and orientation for each of multiple images taken to generate a 3D model of a target object in real space based on a common field of view with other images, and calculates the reliability of the estimated shooting position and orientation regarding the effect on the estimation accuracy of the shooting position and orientation caused by the shooting motion when taking the images.

19. A program that causes a computer system to execute the following steps: For each of multiple images taken to generate a 3D model of a target object in real space, estimate the shooting position and orientation based on a common field of view with other images; and calculate the reliability of the estimated shooting position and orientation regarding the effect on the estimation accuracy of the shooting position and orientation caused by the shooting motion when taking the images.