Information processing system, three-dimensional model generation device, optical device, information processing method, and program

By using FBK photography and depth compositing technology, the 3D model generation process has been optimized, solving the problems of resource waste and increased time caused by poor shooting conditions in existing technologies, and achieving efficient and high-quality 3D model generation.

CN121909487APending Publication Date: 2026-04-21FUJIFILM CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIFILM CORP
Filing Date
2024-09-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to capture images under optimal photographic conditions when generating 3D models, leading to an increase in the number of shots and time-consuming photography setup, which affects model quality.

Method used

Focus bracketing (FBK) photography is used to acquire images with different focus positions from multiple shooting directions. 3D models are generated by deep image synthesis. Computer processing is then used to evaluate images and set conditions to optimize the shooting process.

Benefits of technology

It improves the quality and efficiency of 3D models, reduces shooting time and resource waste, and generates high-precision 3D models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121909487A_ABST
    Figure CN121909487A_ABST
Patent Text Reader

Abstract

The invention provides an information processing system, a three-dimensional model generation device, an optical device, an information processing method, and a program capable of providing a good 3D model. This information processing system is provided with one or more processors that perform: a process for acquiring two or more images having different focus positions in each of a plurality of imaging directions of an object; generating a depth-synthesized image by synthesizing images acquired from the same imaging direction and having different focus positions; and generating a 3D model of at least a portion of the object on the basis of the depth-synthesized image obtained in each of the imaging directions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an information processing system, a three-dimensional model generation device, an optical device, an information processing method, and a program for generating 3D models. Background Technology

[0002] Patent document 1 discloses a method for generating a 3D view of an object, wherein image data is captured from multiple viewpoints around the object, the quality of the image data is analyzed, an image dataset is created based on the image data, the image dataset is filtered, data reference parameters are generated, and the image dataset is uploaded to a server via a network.

[0003] Previous technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Publication No. 2020-527814 Summary of the Invention

[0006] -The technical problem that the invention aims to solve-

[0007] One embodiment of the present invention provides an information processing system, a three-dimensional model generation device, an optical device, an information processing method, and a program capable of providing good 3D models.

[0008] -Means used to solve technical problems-

[0009] The information processing system of the first method has one or more processors, wherein the one or more processors perform the following processing: acquiring two or more images with different focus positions in multiple photographic directions of the object; synthesizing images with different focus positions acquired from the same photographic direction to generate a depth composite image; and generating a 3D model of at least a part of the object based on the depth composite image acquired in each photographic direction.

[0010] The second type of information processing system, in the first type, involves one or more processors performing the following processing: determining whether the image is appropriate for the generation of the 3D model.

[0011] In the third method of information processing system, in the first or second method, one or more processors perform the following processing: determining whether the image is suitable for the generation of a depth-synthesized image.

[0012] In the fourth method of information processing system, in any one of the methods from the first to the third method, one or more processors perform the following processing: determining whether the depth-synthesized image is appropriate for the generation of the 3D model.

[0013] In the fifth method of information processing system, in any one of the methods 2 to 4, one or more processors perform image evaluation to determine whether it is appropriate or not.

[0014] The information processing system of the sixth method, in the fifth method, involves one or more processors performing the following processing: instructing the re-capture of the image based on the image evaluation or appropriateness judgment result.

[0015] In the information processing system of the seventh method, in any one of the methods from the first to the sixth method, one or more processors perform the following processing: determining the conditions for acquiring images of different focus positions of the object.

[0016] In the information processing system of the eighth method, in the seventh method, one or more processors perform the following processing: before capturing an image, conditions are determined based on a pre-image of the object to be captured.

[0017] In the information processing system of the ninth method, in the eighth method, one or more processors perform the following processing: based on the pre-image, calculate the width of the object, and determine the conditions based on the width.

[0018] In the 8th method, the information processing system of the 10th method involves one or more processors performing the following processing: obtaining a pre-3D model of the object based on the pre-image, and determining conditions based on the pre-3D model.

[0019] The information processing system of method 11, in method 10, involves one or more processors performing the following processing: generating a pre-3D model using the visual hull method.

[0020] In the information processing system of the 12th method, in the 7th method, one or more processors perform the following processing: calculate the photographic coverage of the object area for each photographic direction, and determine the conditions based on the calculation results.

[0021] In the 7th method, the information processing system of the 13th method involves one or more processors performing the following processing: determining the area of ​​the object that does not need to be photographed according to each photographic direction, and determining the conditions.

[0022] In the 7th method, the information processing system of the 14th method involves one or more processors performing the following processing: determining conditions in a direction different from the photographic direction based on conditions determined in one direction.

[0023] In the 8th method, the information processing system of the 15th method involves one or more processors performing the following processing: extracting feature points of the object based on the pre-image, and determining conditions based on the feature points.

[0024] In the 15th method, the information processing system of the 16th method involves one or more processors performing the following processing: determining the initial focus position of the image used to create the depth-synthesized image based on feature points.

[0025] The information processing system of method 17, in method 7, involves one or more processors performing the following processing: determining conditions based on the purpose required for the 3D model.

[0026] In the information processing system of method 18, in method 7, one or more processors perform the following processing: determine conditions based on the distance information to the object.

[0027] The 3D model generation apparatus of the 19th method includes an optical device capable of capturing images of an object and an information processing device. The 3D model generation apparatus includes one or more processors contained in either the optical device or the information processing device. The one or more processors perform the following processing: acquiring two or more images with different focus positions in multiple photographic directions of the object; synthesizing images with different focus positions acquired from the same photographic direction to generate a depth-synthesized image; and generating a 3D model of at least a portion of the object based on the depth-synthesized image acquired in each photographic direction.

[0028] In the 19th method, the 3D model generation apparatus of the 20th method involves one or more processors performing the following process: determining whether it is appropriate to apply the image to the generation of the 3D model.

[0029] In the 21st method, in either the 19th or 20th method, one or more processors perform the following process: determining whether the image is suitable for generating a depth-synthesized image.

[0030] In any one of the methods from method 19 to method 21, the 3D model generation apparatus of method 22 performs the following process: determining whether it is appropriate to apply the depth-synthesized image to the generation of the 3D model.

[0031] In any one of the methods from method 20 to method 22, the three-dimensional model generation apparatus of method 23 uses one or more processors to perform image evaluation for determining whether it is appropriate.

[0032] In the 23rd method, the 3D model generation device of the 24th method has one or more processors performing the following processing: based on the image evaluation or the judgment result of whether it is appropriate or not, instructing the image to be re-captured.

[0033] In any of the 19th to 24th modes, the 3D model generation apparatus of mode 25 performs the following process by one or more processors: determining the conditions for acquiring images of different focus positions of the object.

[0034] The optical device of the 26th method has one or more processors and is capable of photographing an object, wherein the one or more processors perform the following processing: acquiring two or more images with different focus positions in multiple photographic directions of the object; synthesizing images with different focus positions acquired from the same photographic direction to generate a depth-synthesized image; and generating a 3D model of at least a portion of the object based on the depth-synthesized image acquired in each photographic direction.

[0035] In the optical device of the 27th method, in the 26th method, one or more processors perform the following process: determining the conditions for acquiring images of different focus positions of the object.

[0036] The information processing method of the 28th method is executed by an information processing system having one or more processors, wherein the one or more processors perform the following processing: acquiring images of the object with different focus positions in multiple photographic directions; synthesizing images with different focus positions acquired from the same photographic direction to generate a depth composite image; and generating a 3D model based on the depth composite image acquired in each photographic direction.

[0037] The program of method 29 executes an information processing method of an information processing system having one or more processors, which causes one or more processors to perform the following processing: acquiring images of an object with different focus positions in multiple photographic directions; synthesizing images with different focus positions acquired from the same photographic direction to generate a depth composite image; and generating a 3D model based on the depth composite image acquired in each photographic direction. Attached Figure Description

[0038] Figure 1 This is a diagram that represents an overview of the system that generates 3D models.

[0039] Figure 2 It represents and generates 3D models. Figure 1 A diagram outlining the different systems.

[0040] Figure 3 It is a block diagram representing the general structure of a camera.

[0041] Figure 4 It is a block diagram that represents the general structure of a computer.

[0042] Figure 5 This is a flowchart illustrating the method for generating a 3D model.

[0043] Figure 6This is an example of FBK photography.

[0044] Figure 7 This is a block diagram representing the functions associated with the generation of depth-synthesized images.

[0045] Figure 8 This is a block diagram representing the functions associated with 3D model generation.

[0046] Figure 9 This is a flowchart illustrating the method for determining which FBK images should be excluded.

[0047] Figure 10 This is a flowchart illustrating the method for determining which depth-synthesized images should be excluded.

[0048] Figure 11 This is a flowchart illustrating a method for generating a 3D model, including pre-photographed images.

[0049] Figure 12 This diagram illustrates an example of the conditions set for FBK photography.

[0050] Figure 13 This is an illustration of an example of interval culling photography.

[0051] Figure 14 This is a diagram illustrating an example of matching an object with a still image (2D image).

[0052] Figure 15 This diagram illustrates the steps for determining the preferred conditions for other FBK photography.

[0053] Figure 16 This diagram illustrates the steps for determining the preferred conditions for other FBK photography. Detailed Implementation

[0054] Hereinafter, preferred embodiments of the present invention will be described with reference to the accompanying drawings.

[0055] [summary]

[0056] In photogrammetry-based 3D modeling, the goal is to capture images of the objects to be modeled at higher resolutions, without any jitter or blur in the captured images.

[0057] To improve the quality of the created 3D models, more detailed images (higher pixel count) are needed. Using a wide-angle or macro lens near the subject results in a shallow depth of field. Considering camera shake, a fast shutter speed is desirable. High sensitivity noise negatively impacts the image quality of the 3D model, so low sensitivity is preferred. Therefore, it's ultimately necessary to reduce the f-stop (open side) to achieve a shallow depth of field. Thus, to create better 3D models, images for depth compositing need to be captured under optimal camera settings (photographic conditions). However, this sometimes leads to an increase in the number of shots or time-consuming setup processes, thus increasing the overall time required for photography.

[0058] <Implementation Method>

[0059] Figure 1 This is a diagram illustrating a system 1 for generating 3D models according to an implementation method. System 1 includes a camera 10 and a computer 20. An object 30, which serves as the 3D model object, is positioned on a camera platform 40.

[0060] Camera 10 is capable of moving around object 30. When moving around object 30, camera 10 is configured to capture images of object 30 from multiple directions. When generating a 3D model of the entire object 30, camera 10 moves around object 30 (e.g., 360° or more). When generating a 3D model of a portion of object 30, camera 10 moves within a specific range of object 30 (less than 360°). Camera 10 can be moved around object 30 by a user holding camera 10 or by a moving body supporting camera 10. The moving body is, for example, an arm mounted on camera stand 40. This arm is configured to move around camera stand 40 via a motor or the like. Furthermore, if object 30 is large, the moving body can be a vehicle or a drone.

[0061] The camera 10 is configured to capture images of the object 30 from various positions while moving around it. The camera 10 can capture both still and moving images. Furthermore, the camera 10 is configured to perform focus bracketing (FBK) photography. By performing FBK photography, the camera 10 can capture two or more images with different focus positions in one photographic direction of the object 30. The image obtained through FBK photography is called an FBK image. FBK photography refers to photography in which multiple images are acquired by moving the focus position in one photographic direction.

[0062] Camera 10 is capable of capturing images of object 30 from multiple photographic directions. Camera 10 is capable of capturing two or more images with different focus positions from each of the multiple photographic directions. Furthermore, camera 10 is configured to store the images obtained from capturing object 30. Camera 10 is configured to capture object 30 manually or automatically upon user instruction (operation). Furthermore, camera 10 may include a rangefinder device capable of measuring the distance to object 30.

[0063] Computer 20 has a display and a keyboard. The display is an example of a display device that shows various information. The keyboard is an example of an input device that allows a user to input instructions. Computer 20 is configured to process various types of data, including images, input and output various types of data, and store various types of data. Computer 20 is configured to store programs for performing its functions and is capable of executing programs.

[0064] The computer 20 is configured to perform depth synthesis on FBK images acquired from the camera 10 by executing a program. Depth synthesis is a technique that synthesizes multiple FBK images with different focus positions to generate a composite image of the object 30 that is in overall focus.

[0065] Computer 20 is configured to generate a 3D model from multiple images of object 30 by executing a program using photogrammetry. Photogrammetry is a technique that analyzes multiple images of object 30 taken from different angles and synthesizes the multiple images to generate (restore) a three-dimensional shape or structure (3D model).

[0066] Object 30 is any object for which a 3D model is to be generated; its shape and size are not particularly limited. Object 30 simply needs to be an object with a physical shape.

[0067] The camera stand 40 has multiple markings 41 on its mounting surface (upper surface). These markings 41 serve as indicators of the photographic direction. Furthermore, the distance between two markings 41 becomes a reference for the size of the 3D model. The camera stand 40 has a cylindrical upper surface with a flat surface. The shape of the camera stand 40 is not limited as long as it can accommodate the object 30. The camera stand 40 can be a cuboid. However, considering factors such as the size of the object 30 and its placement, the camera stand 40 is not essential.

[0068] Figure 2 This is a diagram illustrating the general structure of System 2, which generates 3D models differently from System 1. Figure 2 In the text, the parts that are the same as those in System 1 above are marked with the same symbols, and their descriptions are omitted. Figure 2 System 2 includes a camera 10 and a computer 20. The camera 10 is fixed to a tripod 44 and is in a stationary state. The tripod 44 is an example of a support component that keeps the camera 10 stationary.

[0069] The object 30 is positioned on a camera stage 42 with a marking 43. Unlike the camera stage 40 in System 1, the camera stage 42 is rotatable. The camera stage 42 allows the object 30 to rotate at any angle.

[0070] When generating a 3D model of the entire object 30, the camera stage 42 rotates the object 30 (e.g., 360° or more). When generating a 3D model of a portion of the object 30, the camera stage 42 rotates the object 30 within a specific range (less than 360°). The camera stage 42 includes a motor or the like and is capable of rotating at any speed.

[0071] The camera 10 is configured to capture images at various positions of the object 30 while it is rotated via the camera stage 42. The camera 10 can manually or automatically capture images of the object 30 from multiple shooting directions.

[0072] The camera 10 and computer 20 of System 2 are basically the same as those of System 1 and its camera 10 and computer 20.

[0073] [camera]

[0074] Figure 3 This is a block diagram representing the general structure of camera 10. For example... Figure 3 As shown, the camera 10 includes a lens assembly 100. The lens assembly 100 includes an imaging optical system 102 comprising a lens group 104 and an aperture 106, a lens drive unit 110, an aperture drive unit 112, etc.

[0075] The lens device 100 can be a device that can be attached to or detached from the camera 10, or it can be an integrated device with the camera 10.

[0076] The lens group 104 includes at least a focusing lens that can move along the optical axis. This focusing lens is a focusing lens. The imaging optical system 102 focuses by moving the focusing lens back and forth along the optical axis. The focusing lens is driven by the lens drive unit 110. The focusing lens can be moved to a focusing position that is the focus position for focusing on the object 30.

[0077] The aperture 106 is, for example, a variable aperture. The amount of light passing through the camera optical system 102 is adjusted by the aperture 106. The aperture 106 is driven by the aperture drive unit 112.

[0078] like Figure 3 As shown, the camera 10 includes an imaging element 130, a shutter 132, a shutter drive unit 134, a memory 136, a digital signal processing unit 138, an input / output interface 140, a display unit 142, an operation unit 144, and a system control unit 146.

[0079] The imaging element 130 is, for example, a CMOS (Complementary Metal-Oxide Semiconductor) type image sensor having a predetermined color filter arrangement (e.g., Bayer arrangement). In the camera 10 of the embodiment, the imaging element 130 is configured to include a driving unit, an ADC (Analog to Digital Converter), and a signal processing unit. The imaging element 130 is driven by the built-in driving unit. Furthermore, the signal of each pixel is converted into a digital signal by the built-in ADC. Moreover, the signal of each pixel is subjected to correlation double sampling processing, gain processing, correction processing, etc., as needed by the built-in signal processing unit. The signal processing can be configured to process either the analog signal of each pixel or the digital signal of each pixel.

[0080] In addition to CMOS image sensors, the imaging element 130 can also be composed of organic thin-film imaging elements, XY address type, CCD (Charged Coupled Device) type image sensors.

[0081] The shutter 132 is positioned between the aperture 106 and the imaging element 130. The shutter 132 is driven by the shutter drive unit 134. The shutter drive unit 134 controls the opening and closing of the shutter 132 and controls the exposure time (shutter speed) in the imaging element 130.

[0082] The memory 136 includes flash memory, ROM (Read-only Memory), RAM (Random Access Memory), auxiliary storage devices, etc. The flash memory and ROM store camera control programs and various data required for executing focus bracketing shooting programs, image processing programs, camera control, etc., in focus bracketing shooting mode. The RAM temporarily stores photographic data and functions as a working area processed by the system control unit 146. Furthermore, it temporarily stores camera control programs, image processing programs, etc., stored in the flash memory, etc. Additionally, the system control unit 146 may have a portion of the memory 136 (RAM) built into it.

[0083] The digital signal processing unit 138 performs signal processing such as offset processing, gamma correction processing, de-mosaic processing, and RGB / YCrCb conversion processing on the image obtained by shooting, and generates image data.

[0084] The input / output interface 140 includes a connection section for connecting to an external display device, a connection section for connecting to an external recording device, a card connection section for inserting and removing a memory card, and a communication section for connecting to a network. For example, the input / output interface 140 can be compatible with USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface) (HDMI is a registered trademark), etc.

[0085] The display unit 142 serves as a playback monitor for displaying captured images and a live view monitor for displaying live view images during shooting. It also functions as a setting monitor when making various settings. The display unit 142 is composed of displays such as LCD (Liquid Crystal Display) or OLED (Organic Light Emitting Diode).

[0086] The operation unit 144 is configured to include various operating components for operating the camera 10. These operating components include a power button, a shutter button, and various types of operation buttons. Among these operation buttons are buttons for turning the image correction mechanism 150 on and off. Furthermore, if the display unit 142 is configured as a display unit with a touch panel, the operation components constituting the operation unit 144 include a touch panel. The operation unit 144 outputs signals corresponding to the operation of each operating component to the system control unit 146. For example, the user can set FBK shooting conditions from the operation unit 144.

[0087] The system control unit 146 centrally controls the entire camera 10. Furthermore, the system control unit 146 calculates various physical quantities required for control. The system control unit 146 may be configured as, for example, a microcomputer equipped with a processor and memory. The processor may be, for example, a CPU (Central Processing Unit). The system control unit 146 is capable of executing FBK photography programs. The system control unit 146 includes control of the shake correction mechanism 150, control of the ranging mechanism 160, etc.

[0088] The camera 10 is equipped with a shake correction mechanism 150. The shake correction mechanism 150 is a BIS (Body Image Stabilizer) control method shake correction mechanism that performs shake correction by displacing (including rotating) the imaging element 130 in a plane orthogonal to the optical axis in the opposite direction to the shake direction. The shake correction mechanism 150 includes a shake control unit 152, an imaging element drive unit 154, a shake detection unit 156, and a position detection unit 158.

[0089] The imaging element drive unit 154 may include an actuator for moving the imaging element 130. The shake detection unit 156 may include an accelerometer, a gyroscope, or the like. The position detection unit 158 ​​may include a Hall element that generates a voltage signal corresponding to the position of the imaging element 130. The shake control unit 152 controls the imaging element drive unit 154 based on the signals from the shake detection unit 156 and the position detection unit 158, causing the imaging element 130 to displace in a plane perpendicular to the optical axis to counteract camera shake. A shake correction mechanism may be provided in the lens assembly 100.

[0090] Camera 10 may include a ranging mechanism 160. The ranging mechanism 160 is capable of determining the distance or direction to an object. The ranging mechanism 160 includes a ranging device 162 and a ranging device control unit 164. The ranging device 162 is driven by the ranging device control unit 164. The ranging mechanism 160 may be, for example, a LiDAR (Light Detection and Ranging) system. The LiDAR includes a laser light source and a light sensor. The LiDAR illuminates an object with a laser beam and uses the light sensor to detect the laser beam that has illuminated the object 30 and reflected back. The LiDAR measures the time from when the laser beam illuminates the object 30 to when it reflects back. Thus, the LiDAR can determine the distance or direction to the object 30. Camera 10 is an example of the optical device of the present invention.

[0091] [computer]

[0092] Figure 4 This is a block diagram illustrating an example of the hardware structure of computer 20. For example... Figure 4 As shown, the computer 20 is configured to include a CPU 200, RAM 202, ROM 204, auxiliary storage device 206, input / output interface (IF) 207, input device 208, and display device 209. The ROM 204 and / or auxiliary storage device 206 store programs and various data executed by the CPU 200. The auxiliary storage device 206 is, for example, a HDD (hard disk drive) or an SSD (solid state drive). The input device 208 is, for example, a keyboard, mouse, or touch panel. The display device 209 is, for example, an LCD or OLED. Images captured by the camera 10 are input to the computer 20 via the input / output interface 207 and stored, for example, in the auxiliary storage device 206.

[0093] The computer 20 stores a program that generates a 3D model based on multiple images of the object 30 in the auxiliary storage device 206. The CPU 200 can generate a 3D model by executing the program using photogrammetry.

[0094] Methods for generating 3D models

[0095] Next, the method for generating a 3D model using camera 10, computer 20 and camera stand 40 will be explained. Figure 5 It is a flowchart representing the method for generating a 3D model (information processing method).

[0096] like Figure 5 As shown, the user prepares a camera 10 capable of FBK photography, a computer 20 capable of depth compositing and 3D model creation, and an object 30. The user places the object 30 on the photography platform 40 (step S1).

[0097] The user moves the camera 10 to a designated shooting position to begin photography (step S2). The user fixes the object 30 to the photography platform 40, and while moving the camera 10 around the object 30, performs FBK photography of the object 30 from various shooting directions. The user can arbitrarily determine the shooting position to begin photography. Here, the shooting position is designated as P, and a specific shooting position is designated as Pm (m is an integer). For example, P1 represents the starting position of photography.

[0098] The user performs FBK photography of the object 30 using camera 10 at the photography position (step S3). The user sets camera 10 to a 2D image photography mode for generating a 3D model. At the photography position (P1), the user includes the object 30, which is placed on the photography platform 40, within the photography range of camera 10 from the photography direction (θ1). Here, the photography direction is set as θ, and the photography direction at a specific photography position Pm is set as θm. The photography direction (θm) is the angle representing the relative positional relationship between camera 10 and object 30 at the photography position (Pm).

[0099] If the user half-presses the shutter button of the camera 10, the camera 10 sets the FBK shooting conditions in that shooting direction based on the information of the camera 10's photography equipment and the analysis results of the real-time viewfinder image.

[0100] The following describes the steps for automatically determining the conditions for FBK photography using camera 10. The system control unit 146 of camera 10 analyzes the live view image and extracts the region of the object 30. Known contour extraction processing can be applied to this region extraction. Camera 10 can exclude objects outside the region of the object 30 from the processed objects. This can be achieved through masking or occlusion.

[0101] While moving the focusing lens of camera 10 to change the focus position, the deepest and shallowest distances relative to camera 10 are calculated for the focus position where the contrast within the contour is maximized. Thus, camera 10 acquires the distance between itself and the focus position when focusing on the deepest part of the object 30 (denoted as the deepest focus position distance) and the distance between itself and the focus position when focusing on the shallowest part (denoted as the shallowest focus position distance). In other words, while changing the focus position, the distances between the shallowest and deepest points of the object 30 that are focused on are searched. Alternatively, the shallowest and deepest focus positions can be searched for any position where the focus is not on the object 30.

[0102] The system control unit 146 calculates the depth of field width based on the F-number setting of the lens group 104, the distance to the deepest focal point, and the distance to the shallowest focal point. Next, the system control unit 146 calculates the required focal position and number of shots, and sets the shooting conditions for FBK photography. Once the shooting conditions for FBK photography are set, the camera 10 can notify the user that the setting is complete.

[0103] If the system detects that a user capable of taking photos has pressed the shutter button on the camera 10, the camera 10 will perform FBK photography with the set focus position and number of shots. The system control unit 146 controls the lens drive unit and, while moving the focusing lens, captures images of two or more objects 30 with different focus positions.

[0104] Figure 6 This is an example of FBK photography. Figure 6 This illustrates a scenario where four images are taken from one photographic direction using FBK photography from the same shooting position. FP represents the focus position, and D represents the depth of field. Figure 6 Figure 6-1 shows the case where the focal position of the photographic image is the shallowest focal position FP determined by the system control unit 146. Figure 6 Figure 6-4 shows the case where the focal position of the photographic image is determined by the system control unit 146 at the deepest focal position.

[0105] like Figure 6 As shown in Figure 6-1, the focal position is determined to be the shallowest focal position, and camera 10 takes the first image of object 30. Next, as... Figure 6 As shown in Figure 6-2, the camera 10 takes a second image of the object 30 at the next focal position according to the photographic conditions. The system control unit 146 calculates the focal position of subsequent images based on the focal position and the number of photographs, using a step increment.

[0106] Next, as Figure 6As shown in Figure 6-3, camera 10 takes the third image of object 30 at the next focal position. Finally, as... Figure 6 As shown in Figure 6-4, camera 10 takes the fourth image of object 30 at the deepest focus position and ends FBK photography. Camera 10 can notify the user that FBK photography has ended. Regarding the movement of the focus position, it can be moved from the shallowest focus position to the deepest focus position, or from the deepest focus position to the shallowest focus position, or it can be moved randomly.

[0107] The FBK image captured by FBK is associated with the shooting direction (θ) and which image it is, and stored in memory 136. For example, the shooting direction (θ) and the FBK image are associated and stored as shown in Img(θ,1), Img(θ,2), Img(θ,3)...Img(θ,n) (n is an integer representing the number of shots).

[0108] Furthermore, when storing images, the camera 10 can store information related to the photographic conditions (meta-information). Meta-information may include photographic direction data relative to the object 30 obtained from the camera 10's sensors, etc. When the ranging mechanism 160 measures the distance to the object 30 in the photographic direction, the meta-information may include the distance to the object 30.

[0109] Next, after step S3, a determination is made as to whether the photograph of object 30 has been completed (step S4). If the result of this determination is that the photograph has not been completed (step S4: No), the camera 10 is moved to the next shooting position (P) (step S5). Then, returning to step S3, the user, at the next shooting position (P), performs FBK photography of object 30 in the shooting direction (θ) using the camera 10 (step S3). The process from step S3 to step S5 is repeated until the photograph of object 30 is completed.

[0110] In step S4, the system control unit 146 can determine whether the photography action in the set photography direction has been completed. The system control unit 146 can display the completed image on the display unit 142 so that the user can easily make a judgment.

[0111] On the other hand, after the object 30 is photographed (step S4: yes), the camera 10 ends the FBK photography.

[0112] Next, the computer 20 acquires two or more images (FBK images) with different focus positions in multiple photographic directions of the object 30 stored in the camera 10 (step S6). The computer 20 acquires the FBK images from the memory 136 of the camera 10 via the input / output interface 207 in a wired or wireless manner, or from a recording medium such as a memory card.

[0113] All FBK images, such as the shooting direction (θ) and the order of the images, are stored in the auxiliary storage device 206 of computer 20 after being associated with each other as shown in Table 1 below. In Table 1, the number of images for each shooting direction is n, but the number of images taken in each shooting direction may be different.

[0114]

[0115] Next, the computer 20 performs depth synthesis on the FBK images for each photographic direction and generates a depth-synthesized image for each photographic direction (step S7). Furthermore, the depth synthesis technique itself is a known technique. A summary of this process is as follows.

[0116] Figure 7 The functional boxes related to the generation of depth-synthesized images by the CPU200 are shown. For example... Figure 7 As shown, the image acquisition unit 210 acquires all FBK images (Img(θ1,1), Img(θ1,2)...Img(θ1,n), Img(θ2,1), Img(θ2,2)...Img(θ2,n), ..., Img(θm,1), Img(θm,2)...Img(θm,n)) from the auxiliary storage device 206 in each photographic direction.

[0117] The alignment unit 211 aligns the multiple FBK images used for depth synthesis in each photographic direction. For example, in the event of hand shakiness, the multiple FBK images can be aligned in the same coordinate system through geometric transformations such as affine transformations. The presence or absence of hand shakiness can be determined, for example, based on information from the shake detection unit 156 of the camera 10.

[0118] The region extraction unit 212 of the CPU 200 extracts the focus-aligned pixels (focus areas) from the FBK images (e.g., Img(θ1,1), Img(θ1,2)...Img(θ1,n)) in each photographic direction. When extracting the focus-aligned areas, the region extraction unit 212 extracts the area with the highest contrast value from the areas corresponding to the same position in multiple FBK images in the photographic direction as the focus area, and uses it for the generation of the depth-synthesized image.

[0119] The depth-synthesized image generation unit 213 synthesizes the focus areas extracted from each FBK image to generate a single depth-synthesized image (Img_syn(θ1), Img_syn(θ2)...Img_syn(θm)) that is aligned to the overall focus. The method for extracting and synthesizing the focus areas is not particularly limited and can be any known technique. If depth synthesis is completed in each photographic direction, the CPU 200 stores the depth-synthesized image (Img_syn(θ1), Img_syn(θ2)...Img_syn(θm)) in the auxiliary storage device 206.

[0120] Next, computer 20 generates a 3D model based on multiple depth-synthesized images (step S8). Furthermore, the technique of generating a 3D model using photogrammetry is a well-known technique. A summary of this process is as follows.

[0121] Figure 8 The diagram shows the function blocks associated with the 3D model generation of the CPU200. For example... Figure 8 As shown, the image acquisition unit 220 acquires depth-synthesized images (Img_syn(θ1), Img_syn(θ2) ... Img_syn(θm)) for each photographic direction from the auxiliary storage device 206.

[0122] The point cloud data generation unit 221 analyzes multiple depth-synthesized images and generates 3D point cloud data of feature points. The point cloud data generation unit 221 extracts feature points from each depth-synthesized image. Next, the point cloud data generation unit 221 matches corresponding feature points between different depth-synthesized images, using these corresponding feature points as counterparts. The point cloud data generation unit 221 estimates the camera parameters (e.g., fundamental matrix, basic matrix, and intrinsic parameters) and estimates the shooting position and pose based on the estimated camera parameters. Then, the 3D position of the feature points of the object 30 is determined. Clustering adjustment is performed as needed. The estimated 3D coordinates of the feature points are combined to generate point cloud data (point cloud).

[0123] The 3D surface model generation unit 222 processes the data to generate a 3D surface model of the subject based on the 3D point cloud data of the object 30 generated by the point cloud data generation unit 221. Specifically, it generates surface patches (mesh) based on the generated 3D point cloud and then generates a 3D surface model. As a result, the surface undulations can be represented with a small number of points.

[0124] The 3D model generation unit 223 generates a textured 3D model (three-dimensional model) by performing texture mapping on the 3D patch model generated by the 3D patch model generation unit 222. The 3D model generation unit 223 gives the object 30 a realistic appearance by mapping textures onto the mesh.

[0125] The generated 3D model data is stored in auxiliary storage device 206, etc. Furthermore, the 3D model data is displayed on display device 209 as needed.

[0126] In this embodiment, the case of system 1, in which the object 30 is placed on the camera platform 40 and the camera 10 moves around the object 30, has been described. However, this is not a limitation; 3D models can also be generated in system 2.

[0127] In this implementation, a 3D model is generated by photogrammetry based on multiple depth composite images that are focused from the front to the back of the object 30, thus enabling the generation of a high-precision 3D model.

[0128] When creating a 3D model, the distance information between the two markers 41 on the camera platform 40 can be used as a reference for the size of the 3D model. Based on the multiple markers 41 on the camera platform 40, the positional relationship between the camera 10 and the object 30 can be determined, and a 3D model can be generated. Additionally, the computer 20 is an example of the information processing system of this invention.

[0129] [Preferred Implementation]

[0130] Next, a preferred embodiment will be described. In the preferred embodiment, the computer 20 determines FBK images that should be excluded when generating depth-synthesized images. Furthermore, the computer 20 determines depth-synthesized images that should be excluded when generating 3D models.

[0131] Figure 9 A flowchart is shown to determine which FBK images should be excluded when generating depth-synthesized images. For example... Figure 9 As shown, multiple FBK images are acquired in each photographic direction (step S11). Specifically, the image acquisition unit 210 acquires multiple FBK images.

[0132] Next, the alignment of multiple FBK images is performed in each photographic direction (step S12). Specifically, the alignment unit 211 performs the alignment. For example, alignment is performed by selecting a reference image from multiple FBK images and extracting feature points from the reference image. The position of the corresponding point in the remaining FBK images is tracked, and based on the result, the remaining FBK images are subjected to parallel translation, rotation, and scaling operations such as affine transformation. The alignment of multiple FBK images is performed. Even with some hand tremors, alignment can be performed on the FBK images.

[0133] Next, FBK images that cannot be aligned are excluded from the objects of depth compositing (step S13). Specifically, the alignment unit 211 excludes FBK images from the objects of depth compositing. In step S12, the alignment unit 211 can determine that FBK images that cannot be aligned with the reference image are not suitable for depth compositing. The alignment unit 211 associates the information indicating that it is not suitable for depth compositing with the FBK image and stores it in the auxiliary storage device 206.

[0134] Next, the focus-aligned region is extracted from the FBK image (step S14). Specifically, the region extraction unit 212 extracts the focus-aligned region from the FBK image. As described above, the region with the highest contrast value is extracted from the region corresponding to the same position in the FBK image as the focus region and used for generating the depth-synthesized image. However, the FBK image excluded from the depth-synthesized object is not included. This shortens the extraction time of the focus position.

[0135] Next, FBK images whose focus positions cannot be extracted are excluded from the depth compositing objects (step S15). Specifically, in step S14, the region extraction unit 212 can determine that FBK images whose focus positions cannot be extracted are not suitable for depth compositing (excluded). The region extraction unit 212 associates the information indicating that it is not suitable for depth compositing with the FBK images and stores it in the auxiliary storage device 206.

[0136] Finally, the focus areas extracted from each FBK image are synthesized to generate a depth-synthesized image (step S16). Specifically, the depth-synthesized image generation unit 213 generates a depth-synthesized image. However, the FBK images excluded from the depth-synthesizing object in step S13 or step S15 are not included in the FBK images. Furthermore, the depth-synthesizing image generation unit 213 can also extract whether there is camera shake (shooting conditions) from the FBK images and exclude images determined to have camera shake from the depth-synthesizing object. Whether there is camera shake can be determined based on high-frequency components of the FBK image, etc. In the step of excluding FBK images from the depth-synthesizing object, the exclusion can be determined based on the image evaluation of the FBK image.

[0137] By excluding images unsuitable for depth synthesis, high-quality depth-synthesized images can be generated. Therefore, high-quality 3D models can be generated from high-quality depth-synthesized images.

[0138] In addition, Figure 9In the illustrated process, when FBK images are determined to be excluded from the deep composited object, the computer 20 can also prompt the camera 10 to retake the image. Information such as the position and orientation of the camera 10 that captured the FBK image can be obtained based on information accompanying the acquisition of the FBK image. The computer 20 can then notify the user of the position and orientation of the camera 10. The user can easily determine the position of the object 30 that needs to be retaken as an FBK image.

[0139] Figure 10 A flowchart is shown to determine which FBK images should be excluded when generating a 3D model. For example... Figure 10 As shown, a depth-synthesized image is acquired (step S21). Specifically, the image acquisition unit 220 acquires multiple depth-synthesized images.

[0140] Next, feature points are extracted from multiple depth-synthesized images (step S22). Feature points are detected from the multiple depth-synthesized images as corresponding points (step S23). Specifically, the point cloud data generation unit 221 extracts feature points from multiple depth-synthesized images and detects consistent feature points as corresponding points by comparing the feature points of the multiple depth-synthesized images.

[0141] Next, the depth-synthesized image is excluded from the objects of the 3D model (step S24). Specifically, the point cloud data generation unit 221 can determine that depth-synthesized images from which feature points cannot be extracted in step S22 or from which corresponding points cannot be detected in step S23 are not suitable for the 3D model. Furthermore, the point cloud data generation unit 221 can determine images excluded from the objects of the 3D model based on high-frequency components, etc., of the depth-synthesized image. The point cloud data generation unit 221 associates the information indicating that the image is not suitable for the generation of the 3D model with the depth-synthesized image and stores it in the auxiliary storage device 206.

[0142] Next, a 3D model is generated based on the depth-synthesized image (step S25). Specifically, the point cloud data generation unit 221 combines the camera parameters, shooting position and pose of the camera, and the estimated 3D coordinates of the feature points, and generates point cloud data based on the depth-synthesized image. The 3D patch model generation unit 222 and the 3D model generation unit 223 then generate the 3D model. The depth-synthesized image does not include depth-synthesized images excluded from the objects in the 3D model. Since depth-synthesized images unsuitable for generating a 3D model are excluded, a high-quality 3D model can be generated.

[0143] In addition, Figure 10In the illustrated process, when a depth-synthesized image is determined to be excluded from the object in the 3D model, the computer 20 can also prompt the camera 10 to retake the image. Information such as the position and orientation of the camera 10 that captured the FBK image suitable for depth synthesis can be obtained based on information accompanying the acquisition of the depth-synthesized image. The computer 20 can then notify the user of the position and orientation of the camera 10. The user can easily determine the position of the object 30 for which the FBK image needed to be retaken for the generation of the depth-synthesized image.

[0144] <Settings for FBK Photography>

[0145] The conditions for performing FBK photography are explained.

[0146] Depth of field (DOF) refers to the area within a defined range in front of and behind the focusing position in a captured image. Within this range, the subject is captured relatively clearly, resulting in an image that is in focus. Conversely, areas outside this range are blurred, resulting in an image that is out of focus. Depth of field is a physically and optically determined value based on the lens's focal length, pixel size (the size of one pixel in an imaging element), the camera's F-number (aperture value), the distance from the pixel sensor to the subject (focus position), and the allowable circle of confusion. Therefore, depth of field varies depending on the F-number and the distance between the subject and the camera.

[0147] In FBK photography, the same object is photographed at multiple focal points to obtain multiple FBK images. These multiple FBK images are then depth-synthesized to obtain an image with extended depth of field for the object (depth-synthesized image).

[0148] In FBK photography, when shooting from multiple focal points (assuming a fixed distance between the subject and the camera), a greater depth of field occurs when shooting with a larger F-number (a reduced focal length). As a result, the number of FBK images captured is reduced, thus decreasing the workload and movement between FBK images.

[0149] On the other hand, the amount of light reaching the imaging element is reduced, so it is necessary to take pictures under either of the following photographic conditions: (1) increase the shutter speed to compensate for not increasing the ISO sensitivity, thereby suppressing high sensitivity noise, or (2) shorten the shutter speed to compensate for increasing the ISO sensitivity, thereby suppressing hand shake.

[0150] However, (1) and (2) are trade-offs, so in case (1) there is a problem of hand tremor, and in case (2) there is a problem of high sensitivity noise.

[0151] Furthermore, the photographic environment also affects photographic conditions. For example, in bright environments such as those with artificial lighting, increasing the F-stop or decreasing the ISO sensitivity will not affect image quality.

[0152] In the implementation, conditions are set for FBK photography that can reduce the workload of photography in each shooting direction while maintaining the quality of the generated 3D model. For example, FBK photography can be determined by taking into account F-stop, ISO sensitivity, and shutter speed.

[0153] When the depth of the subject is short relative to the shooting direction, even with a slightly shallow depth of field (narrow width), fewer shots can be taken, reducing workload. Therefore, the computer 20 can set the conditions for FBK photography as a small F-stop, short shutter speed (exposure time), and low ISO sensitivity. As examples of conditions for setting FBK photography, a depth of field setting that is small F-stop, short shutter speed, and low ISO sensitivity can be considered within the range that reduces the number of FBK image shots.

[0154] The conditions for these FBK photographs can be pre-stored in tabular form in the auxiliary storage device 206 of the computer 20. Furthermore, an AI (Artificial Intelligence) model can be mounted on the CPU of the computer 20 to learn from the prepared learning data.

[0155] Furthermore, the explanation assumes that FBK photography is performed under the premise of changing the F-value for each photography direction. However, if the image quality is inconsistent for each photography direction, it may affect the quality of the generated 3D model. Therefore, pre-photography can be performed before FBK photography, and a common FBK photography setting can be determined based on the object 30 (regardless of the photography direction). Moreover, the settings can be set "separately" for each photography direction to achieve optimal FBK images from each direction. Alternatively, the settings can be set "commonly" to achieve optimal settings when considering the images from each photography direction "overall." The FBK photography settings can be appropriately determined according to the purpose of generating 3D data.

[0156] As another example of setting conditions in FBK photography, there is a beginner mode. In beginner mode, to reduce the workload of shooting, the f-stop is increased, the depth of field is increased, and the number of shots is reduced. On the other hand, the incidence of camera shake is higher than that of experienced photographers. Therefore, even if some high-sensitivity noise is generated, it is possible to set the condition of increasing the shutter speed as a condition for FBK photography.

[0157] In FBK photography settings, as another example, after pre-shooting, regardless of the shooting direction, the depth of field of the subject (30) is very short. That is, even with a small subject (30), the number of shots won't increase due to the shallow depth of field, allowing for a reduction in the f-number. Reducing the f-number allows for a shorter shutter speed to prevent camera shake. It also allows for adjusting ISO sensitivity based on ambient lighting to improve image quality. Furthermore, FBK photography conditions can be set to achieve the optimal balance.

[0158] Next, the situation of pre-photographing the object 30 by camera 10 will be explained. Figure 11 This is a flowchart illustrating a method for generating a 3D model, including pre-photographed images.

[0159] Prepare a camera 10 capable of FBK photography, a computer 20 capable of depth compositing and 3D model creation, and an object 30. For example... Figure 11 As shown, the user places object 30 on camera platform 40 (step S41).

[0160] The user moves the handheld camera 10 to a designated shooting position to begin pre-shooting. The object 30 is fixed on the shooting platform 40. While moving the handheld camera 10 around the object 30, the user takes pre-shooting images of the object 30 from various shooting directions (step S42). Pre-shooting refers to taking pictures of the object 30 before performing FBK photography (step S45). Pre-shooting can be a still image of the object 30 or a moving image. It can also be a live view image. The distance between the object 30 and the camera 10 can also be obtained.

[0161] Next, computer 20 acquires a pre-image obtained from pre-photographing by camera 10, and acquires information for determining the conditions for FBK photography (step S43). The information that can be acquired includes the position and orientation of camera 10 at a certain coordinate location, and a pre-3D model of the object 30 obtained from the pre-photographing. The pre-3D model refers to a 3D model generated based on the pre-image before generating the 3D model. Regarding the pre-3D model, if the approximate size of the object 30 is known, the accuracy of the 3D model generated in step S50 is not required. When generating the pre-3D model, for example, if rapid shooting is desired, the parallax volume intersection method, which has a fast processing speed, is preferred. Furthermore, the parallax volume intersection method itself is a known technique, which estimates the 3D model based on information about the outline of the object 30. The acquired information is stored in auxiliary storage device 206.

[0162] The user moves the camera 10 to a designated shooting position to begin photography (step S44). From this position, the user performs FBK photography of the object 30 using the camera 10 (step S45). The pre-3D model obtained in step S43 has a completed three-dimensional actual shape. Therefore, the actual depth distance when viewing the object 30 from various directions can be obtained from the pre-3D model. When the user performs FBK photography of the object 30 from a certain direction, the FBK photography conditions (e.g., depth of field, number of FBK images, etc.) can be automatically set based on these distances.

[0163] The distance to the object 30 is calculated by the CPU 200 of the computer 20. Furthermore, the FBK photography settings are executed by the system control unit 146 (CPU) of the camera 10. However, this is not a limitation; either the CPU 200 or the system control unit 146 can execute the settings.

[0164] Figure 12 This diagram illustrates an example of setting conditions for FBK photography. A pre-3D model is generated before performing FBK photography. Figure 12 In step 12-1, the depth distance L1 of the object 30 in that direction is obtained from the pre-3D model. Since the distance L1 is short, the camera 10 or computer 20 determines the conditions for FBK photography, such as the F-number, ISO sensitivity, and shutter speed, which result in a shallower depth of field.

[0165] Similarly, in Figure 12 In step 12-2, the depth distance L2 of the object 30 in that direction is obtained from the pre-3D model. Since the distance L2 is relatively long, the camera 10 or computer 20 determines the conditions for FBK photography, such as the F-number, ISO sensitivity, and shutter speed, which affect the depth of field. Based on these FBK photography conditions, the number of FBK images can be reduced.

[0166] In addition, the case where the depth of object 30 is determined based on a pre-3D model has been explained, but the depth of object 30 in each direction can also be determined based on a pre-image (e.g., a 2D image: a still image).

[0167] Next, we will explain the interval rejection photography in FBK photography. Interval rejection photography refers to omitting FBK photography from a specific direction when FBK photography is performed from multiple directions. By performing interval rejection photography, the number of FBK photography sessions can be reduced, thereby reducing the workload.

[0168] Figure 13 This is an illustration of an example of interval culling photography. Figure 13 Figure 13-1 shows the state of FBK photography of the object 30 from the first photographic direction by camera 10. Figure 13Figure 13-2 shows the state in which the object 30 is photographed by camera 10 from a second photographing direction different from the first photographing direction. Figure 13 Figure 13-3 shows the state in which the object 30 is photographed by camera 10 from a third photographing direction different from the first and second photographing directions.

[0169] For example, based on a pre-photograph or pre-3D model, camera 10 or computer 20 can envision an FBK image obtained by FBK photography in the photographic directions shown from the first photographic direction (13-1) to the third photographic direction (13-3). It can be determined that the FBK image obtained in the second photographic direction (13-2) is included in the FBK images obtained in the first photographic direction (13-1) and the third photographic direction (13-3).

[0170] The camera 10 or computer 20 can determine the area to be ignored from the FBK image of the object 30 in the second shooting direction (13-2). Furthermore, the camera 10 or computer 20 can display the FBK images envisioned from the first shooting direction (13-1) to the third shooting direction (13-3) on the display device 209. The user can determine the area to be ignored based on the envisioned FBK images.

[0171] The camera 10 or computer 20 can calculate the photographic coverage of the object area of ​​the object 30 based on the FBK images acquired from the first photographic direction (13-1) to the third photographic direction (13-3). Based on this calculation result, the conditions for FBK photography can be determined. For example, based on the photographic coverage, it can be determined whether the FBK image acquired from the second photographic direction (13-2) is included in the FBK images acquired from the first photographic direction (13-1) and the third photographic direction (13-3). If sufficient FBK images can be acquired by shooting from the first photographic direction (13-1) and the third photographic direction (13-3), then the second photographic direction (13-2) can be set as an object outside the FBK photography area. Furthermore, the number of FBK images taken when performing FBK photography from the second photographic direction (13-2) can be reduced.

[0172] When performing pre-photography or generating a pre-3D model, during FBK photography (step S45), it is preferable to use camera 10 to photograph the object 30 from the position and direction of the pre-photography. This makes it easier to utilize the information from the pre-photography or pre-3D model.

[0173] exist Figure 13In sections 13-1 to 13-3, the steps for interval elimination photography during FBK photography were explained. Next, using the same diagrams, another step in determining the FBK photography conditions will be explained. FBK photography is performed from the photography directions shown in the first photography direction (13-1) to the third photography direction (13-3). First, the conditions for FBK photography performed from the first photography direction (13-1) are determined. Next, the conditions for FBK photography performed from the second photography direction (13-2) are determined. At this point, it can be determined that the second photography direction (13-2) has not changed significantly relative to the first photography direction (13-1). As a result, the conditions for FBK photography from the second photography direction (13-2) can be made the same as those from the first photography direction (13-1). There is no need to spend time determining the conditions for FBK photography from the second photography direction (13-2).

[0174] On the other hand, it can be determined that the third imaging direction (13-3) changes significantly relative to the first imaging direction (13-1). The determination of the FBK imaging conditions from the third imaging direction (13-3) is not affected by the FBK imaging conditions from the first imaging direction (13-1). That is, it is possible to determine the FBK imaging conditions in different directions based on the FBK imaging conditions determined in one imaging direction. Furthermore, whether there has been a significant change can be determined by tracking feature points, etc.

[0175] Figure 14 This is a diagram illustrating an example of matching object 30 with a still image (2D image).

[0176] For example, Figure 14 Figure 14-1 shows the state where the user wants to perform FBK photography on the object 30 using the camera 10. A live view image 31 of the object 30 is displayed on the display unit 142 of the camera 10. The camera 10 stores still images (2D images) taken for pre-photography or pre-3D modeling in the memory 136. The camera 10 uses template matching to find and display the 2D image 32 most similar to the image of the object 30 displayed in the live view image 31. Information about the position and orientation of the camera 10 when capturing the 2D image 32 is obtained. Therefore, the current distance between the object 30 and the camera 10 can also be determined and used when determining the conditions for FBK photography. The camera 10 displays the 2D image 32, but it is also possible to notify the user of the completion of template matching without displaying the 2D image 32.

[0177] Figure 14Figure 14-2 shows the state where the user wants to perform FBK photography on the object 30 using the camera 10. The live view image 31 of the object 30 is displayed on the display unit 142 of the camera 10. As shown in Figure 14-2, the display unit 142 displays the outline of the 2D image 33. The 2D image 33 prompts the user to move in a direction (arrow direction) relative to the object 30 that makes the outline of the 2D image 33 consistent. Thus, the current distance between the object 30 and the camera 10 can also be determined, and this information can be used when determining the conditions for FBK photography.

[0178] Next, the conditions for determining other preferred FBK photography will be explained. Figure 15 This diagram illustrates the steps for determining the preferred conditions for other FBK photography.

[0179] like Figure 15 As shown in Figure 15-1, object 30A differs from object 30 in that it has a feature part 30B. For example... Figure 15 As shown in Figure 15-2, in the first photographic direction, the camera 10 extracts the feature 30B based on the live view image, which serves as a preview image, focuses on the feature 30B, and uses this position as a reference. The conditions for FBK photography are determined by shifting the focus position back and forth based on the feature 30B.

[0180] Similarly, as Figure 15 As shown in Figure 15-3, in the second photographic direction, the camera 10 extracts feature 30B based on a real-time viewfinder image serving as a pre-image, focuses on feature 30B, and uses this position as a reference. The conditions for FBK photography are determined by shifting the focus position back and forth using feature 30B as a reference. That is, as an FBK photography condition, the initial focus position of the FBK image used to create the depth-synthesized image is determined to be feature 30B. Feature 30B becomes an important element when generating a 3D model; therefore, by setting the initial focus position of the FBK image to feature 30B, a high-precision depth-synthesized image, crucial for generating the 3D model, can be obtained.

[0181] Next, the conditions for determining other preferred FBK photography will be explained. Figure 16 This diagram illustrates the steps involved in determining the preferred conditions for other FBK photography. Figure 16 In this process, FBK photography is performed on the object 30 from a first shooting direction and FBK photography is performed on the object 30 from a second shooting direction. When FBK photography is performed from the first shooting direction, the width distance L3 of the object 30 can be obtained. Then, when FBK photography is performed from the second shooting direction, the number of FBK shots, depth of field, and shooting interval can be determined using the distance L3.

[0182] Return to Figure 11 After step S45, a determination is made as to whether the photographing of object 30 has been completed (step S46). If the result of this determination is that the photographing has not been completed (step S46: No), the camera 10 is moved to the next shooting position (P) (step S47). At the next shooting position (P), the user performs FBK photography of object 30 using camera 10 in the shooting direction (θ) (step S45). The process from step S45 to step S47 is repeated until the photographing of object 30 is completed.

[0183] Upon completion of the photographing of object 30 (step S46: Yes), camera 10 ends FBK photography.

[0184] Next, the computer 20 acquires two or more images (FBK images) with different focus positions in multiple photographic directions of the object 30 stored in the camera 10 (step S48).

[0185] Next, the computer 20 performs depth synthesis on the FBK images in each shooting direction to generate a depth-synthesized image in each shooting direction (step S49).

[0186] Next, the computer 20 generates a 3D model based on multiple depth-synthesized images (step S50). Once the generation of the 3D model is complete, the process ends.

[0187] In the above embodiments, systems 1 and 2 constituting a 3D model generation apparatus including a camera 10 and a computer 20 for generating 3D models have been described. The camera 10 includes a system control unit 146, and the computer 20 includes a CPU. The functions of the system control unit 146 and the CPU 200 are not particularly limited. For example, the determination of FBK photography conditions can be performed by either the CPU 200 or the system control unit 146. That is, as long as the camera 10 and the computer 20 can generate a 3D model, either the CPU 200 or the system control unit 146 can perform its function.

[0188] Furthermore, the present invention includes an information processing system comprising a computer 20 for generating 3D models. It also includes an optical apparatus for performing FBK photography using only the camera 10, depth synthesis of the FBK images, and generation of the 3D model. It includes a 3D model generation apparatus comprising the camera 10 (optical apparatus) and the computer 20 (information processing apparatus). While an example of FBK photography using one camera 10 is shown, FBK photography can also be performed using multiple cameras 10.

[0189] [Structure of System Control Unit 146 and CPU 200]

[0190] The functions of the system control unit 146 and CPU 200 are implemented by various processors. These processors include general-purpose processors (CPUs and / or GPUs, Graphics Processing Units) that execute programs and function as various processing units; programmable logic devices (PLDs) whose circuit structures can be modified after manufacturing; and application-specific integrated circuits (ASICs) with circuit structures specifically designed to perform specific processes. The terms "program" and "software" have the same meaning.

[0191] A processing unit can be composed of one of these various processors, or it can be composed of two or more processors of the same or different types. For example, a processing unit can be composed of multiple FPGAs or a combination of a CPU and an FPGA. Furthermore, multiple processing units can also be composed of a single processor. As examples of multiple processing units composed of a single processor, firstly, in computers used for clients and servers, a single processor is composed of a combination of one or more CPUs and software, and this processor functions as multiple processing units. Secondly, in systems-on-chips (SoCs), a processor that implements the overall system functionality including multiple processing units is used, represented by a single IC (Integrated Circuit) chip. In this way, various processing units are constructed as hardware structures using one or more of the aforementioned processors.

[0192] The present invention has been described above, but the present invention is not limited to the examples above. Of course, various improvements or modifications can be made without departing from the spirit of the present invention.

[0193] Symbol Explanation

[0194] 1-System, 2-System, 10-Camera, 20-Computer, 30-Object, 30A-Object, 30B-Feature Unit, 31-Live View Image, 32-2D Image, 33-2D Image, 40-Camera Stand, 41-Marker, 42-Camera Stand, 43-Marker, 44-Tripod, 100-Lens Assembly, 102-Camera Optical System, 104-Lens Group, 106-Aperture, 110-Lens Drive Unit, 112-Aperture Drive Unit, 130-Imaging Element, 132-Shutter, 134-Shutter Drive Unit, 136-Memory, 138-Digital Signal Processing Unit, 140-Input / Output Interface, 142-Display Unit, 144-Operation Unit, 146-System Control Unit, 150-Shake Correction Mechanism 152-Jitter control unit, 154-Imaging element drive unit, 156-Detection unit, 158-Position detection unit, 160-Range measuring mechanism, 162-Range measuring device, 164-Range measuring device control unit, 200-CPU, 202-RAM, 204-ROM, 206-Auxiliary storage device, 207-Input / output interface, 208-Input device, 209-Display device, 210-Image acquisition unit, 211-Alignment unit, 212-Region extraction unit, 213-Depth composite image generation unit, 220-Image acquisition unit, 221-Point cloud data generation unit, 222-3D patch model generation unit, 223-3D model generation unit, D-Depth of field, FP-Focus position, L1-Distance, L2-Distance, L3-Distance.

Claims

1. An information processing system having one or more processors, wherein, The one or more processors perform the following processing: Acquire two or more images of the object with different focus positions from multiple photographic directions; A depth-synthesized image is generated by synthesizing images acquired from the same photographic direction but with different focus positions; and A 3D model of at least a portion of the object is generated based on the depth-synthesized image obtained in each of the said photographic directions.

2. The information processing system according to claim 1, wherein, The one or more processors perform the following processing: Determine whether it is appropriate to apply the image to the generation of the 3D model.

3. The information processing system according to claim 1, wherein, The one or more processors perform the following processing: Determine whether the image is appropriate for the generation of the depth-synthesized image.

4. The information processing system according to claim 1, wherein, The one or more processors perform the following processing: Determine whether the depth-synthesized image is appropriate for the generation of the 3D model.

5. The information processing system according to any one of claims 2 to 4, wherein, The one or more processors perform image evaluation to determine whether the image is appropriate or not.

6. The information processing system according to claim 5, wherein, The one or more processors perform the following processing: Based on the image evaluation or the judgment of whether it is appropriate, the image is instructed to be retaken.

7. The information processing system according to claim 1, wherein, The one or more processors perform the following processing: Determine the conditions for acquiring images of the object at different focus positions.

8. The information processing system according to claim 7, wherein, The one or more processors perform the following processing: Before taking the image, the conditions are determined based on a pre-image obtained by taking the image of the object.

9. The information processing system according to claim 8, wherein, The one or more processors perform the following processing: Based on the pre-image, the width of the object is determined, and based on the width, the conditions are determined.

10. The information processing system according to claim 8, wherein, The one or more processors perform the following processing: Based on the pre-image, a pre-3D model of the object is obtained, and based on the pre-3D model, the conditions are determined.

11. The information processing system according to claim 10, wherein, The one or more processors perform the following processing: The pre-3D model is generated using the visual volume cross method.

12. The information processing system according to claim 7, wherein, The one or more processors perform the following processing: Calculate the photographic coverage of the object area for the object in each photographic direction, and determine the conditions based on the calculation results.

13. The information processing system according to claim 7, wherein, The one or more processors perform the following processing: Determine the areas where the object does not need to be photographed according to each of the photographic directions, and determine the conditions.

14. The information processing system according to claim 7, wherein, The one or more processors perform the following processing: Based on the conditions determined in one direction of the photographic direction, the conditions in a direction different from the one direction are determined.

15. The information processing system according to claim 8, wherein, The one or more processors perform the following processing: Based on the pre-image, feature points of the object are extracted, and the conditions are determined based on the feature points.

16. The information processing system according to claim 15, wherein, The one or more processors perform the following processing: Based on the feature points, the initial focus position of the image used to create the depth-synthesized image is determined as the condition.

17. The information processing system according to claim 7, wherein, The one or more processors perform the following processing: The conditions are determined based on the objectives required for the 3D model.

18. The information processing system according to claim 7, wherein, The one or more processors perform the following processing: The conditions are determined based on the distance information to the object.

19. A three-dimensional model generation device, comprising an optical device capable of capturing an image of an object and an information processing device, wherein, The 3D model generation device includes one or more processors included in either the optical device or the information processing device. The one or more processors perform the following processing: Acquire two or more images of the object with different focus positions from multiple photographic directions; A depth-synthesized image is generated by synthesizing images acquired from the same photographic direction but with different focus positions; and A 3D model of at least a portion of the object is generated based on the depth-synthesized image obtained in each of the said photographic directions.

20. The three-dimensional model generation apparatus according to claim 19, wherein, The one or more processors perform the following processing: Determine whether it is appropriate to apply the image to the generation of the 3D model.

21. The three-dimensional model generation apparatus according to claim 19, wherein, The one or more processors perform the following processing: Determine whether the image is appropriate for the generation of the depth-synthesized image.

22. The three-dimensional model generation apparatus according to claim 19, wherein, The one or more processors perform the following processing: Determine whether the depth-synthesized image is appropriate for the generation of the 3D model.

23. The three-dimensional model generation apparatus according to any one of claims 20 to 22, wherein, The one or more processors perform image evaluation to determine whether the image is appropriate or not.

24. The three-dimensional model generation apparatus according to claim 23, wherein, The one or more processors perform the following processing: Based on the image evaluation or the judgment of whether it is appropriate, the image is instructed to be retaken.

25. The three-dimensional model generation apparatus according to claim 19, wherein, The one or more processors perform the following processing: Determine the conditions for acquiring images of the object at different focus positions.

26. An optical device having one or more processors and capable of photographing an object, wherein, The one or more processors perform the following processing: Two or more images with different focus positions are acquired from multiple photographic directions of the object; A depth-synthesized image is generated by synthesizing images acquired from the same photographic direction but with different focus positions; and A 3D model of at least a portion of the object is generated based on the depth-synthesized image obtained in each of the said photographic directions.

27. The optical device according to claim 26, wherein, The one or more processors perform the following processing: Determine the conditions for acquiring images of the object at different focus positions.

28. An information processing method, executed by an information processing system having one or more processors, wherein, The one or more processors perform the following processing: Acquire images of the object with different focus positions from multiple photographic directions; A depth-synthesized image is generated by synthesizing images acquired from the same photographic direction but with different focus positions; and A 3D model is generated based on the depth-synthesized image obtained in each of the said photographic directions.

29. A program that executes an information processing method of an information processing system having one or more processors, wherein the one or more processors perform the following processing: Acquire images of the object with different focus positions from multiple photographic directions; A depth-synthesized image is generated by synthesizing images acquired from the same photographic direction but with different focus positions; and A 3D model is generated based on the depth-synthesized image obtained in each of the said photographic directions.

30. A recording medium that is non-transitory and computer-readable, and which records the program of claim 29.

Citation Information

Patent Citations

  • System and method for creating and displaying interactive 3D representations of real objects

    JP2020527814A