IMAGING DEVICE, SERVER DEVICE, AND 3D DATA GENERATION METHOD

The image processing system facilitates volumetric imaging and 3D model generation using smartphones and tablets, addressing the limitations of dedicated studios by determining device capability and generating 3D models for free viewpoint display.

JP7757974B2Active Publication Date: 2025-10-22SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022555353
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-07
Filing Date
2021-09-24
Publication Date
2025-10-22
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

Existing volumetric photography requires dedicated studios and specialized equipment, limiting its accessibility for users with common electronic devices like smartphones and tablets.

Method used

An image processing system that enables volumetric imaging using smartphones and tablets by transmitting device information to a server, determining capability for volumetric photography, and generating 3D model data from multiple device captures.

Benefits of technology

Enables volumetric photography using commonly owned devices, allowing for 3D model generation and free viewpoint image display on various devices, overcoming the need for dedicated studios and equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007757974000001
    Figure 0007757974000001
  • Figure 0007757974000002
    Figure 0007757974000002
  • Figure 0007757974000003
    Figure 0007757974000003
Patent Text Reader

Abstract

A technology of the present invention relates to an image capture device, a server device, and a 3D data generation method that enables volumetric imaging using a simple image capture device. The server device comprises a control unit that receives information pertaining to imaging from a plurality of image capture devices, and on the basis of the received information, determines whether the plurality of image capture devices can perform volumetric imaging. Each of the image capture devices comprises a control unit that transmits information of the host image capture device pertaining to image capture to the server device, and acquires a candidate for a setting value pertaining to volumetric imaging transmitted from the server device on the basis of determination results as to whether volumetric imaging is possible. This disclosure can be applied to an image processing system or the like that provides volumetric imaging and playback by an image capture device such as a smartphone, for example.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present technology relates to an imaging device, a server apparatus, and a 3D data generation method, and more particularly to an imaging device, a server apparatus, and a 3D data generation method that enable volumetric imaging using a simple imaging device. [Background technology]

[0002] There is a technology that generates a 3D model of a subject from video images captured from multiple viewpoints, and then generates a virtual viewpoint image of the 3D model according to any viewing position, thereby providing images from any viewpoint. This technology is also known as volumetric capture.

[0003] When capturing moving images for generating a 3D model, the subject must be captured from different directions (viewpoints), so multiple image capture devices for capturing the subject must be placed in different locations, the positional relationships between the image capture devices must be determined, and each of the multiple image capture devices must be synchronized to capture the images (see, for example, Patent Document 1). This type of image capture for generating a 3D model is also referred to as volumetric image capture below.

[0004] Currently, volumetric photography is generally performed in dedicated studios using dedicated equipment. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 2019-87791 Summary of the Invention [Problem to be solved by the invention]

[0006] However, it is desirable to be able to easily perform volumetric photography using electronic devices with photography functions that are commonly owned by users, such as smartphones and tablets.

[0007] The present technology has been made in view of such circumstances, and makes it possible to perform volumetric imaging using a simple imaging device. [Means for solving the problem]

[0008] The imaging device of the first aspect of the present technology includes a control unit that transmits its own information related to imaging to a server device and acquires candidate setting values ​​related to volumetric imaging transmitted from the server device based on a determination result of whether volumetric imaging is possible.

[0009] In a first aspect of the present technology, information about the device itself related to photography is sent to a server device, and candidate setting values ​​for volumetric photography are obtained that are sent from the server device based on the determination result of whether volumetric photography is possible.

[0010] A server device according to a second aspect of the present technology includes a control unit that receives information related to photography from a plurality of photography devices and determines, based on the received information, whether the plurality of photography devices are capable of volumetric photography.

[0011] In a second aspect of the present technology, information related to photography is received from a plurality of photography devices, and based on the received information, it is determined whether the plurality of photography devices are capable of volumetric photography.

[0012] A 3D data generation method of a third aspect of the present technology receives information related to photography from a plurality of imaging devices, determines whether the plurality of imaging devices are capable of volumetric photography based on the received information, and generates 3D model data of a subject from images captured by the plurality of imaging devices determined to be capable of volumetric photography.

[0013] In a third aspect of the present technology, information regarding photography is received from a plurality of imaging devices, and based on the received information, it is determined whether the plurality of imaging devices are capable of volumetric photography, and 3D model data of the subject is generated from images captured by the plurality of imaging devices determined to be capable of volumetric photography.

[0014] The image capturing device and the server device may be independent devices, or may be internal blocks constituting a single device. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a block diagram showing an example configuration of an image processing system to which the present disclosure is applied. [Figure 2] FIG. 1 is a diagram illustrating volumetric imaging and playback. [Figure 3] FIG. 2 is a diagram illustrating an example of a data format of 3D model data. [Figure 4] FIG. 2 is a block diagram illustrating an example of a detailed configuration of each device in the image processing system. [Figure 5] 10 is a flowchart illustrating volumetric shooting and playback processing by the image processing system. [Figure 6] 6 is a flowchart illustrating details of the grouping process of the imaging devices in step S1 of FIG. 5. [Figure 7] FIG. 10 is a diagram showing an example of a two-dimensional code screen. [Figure 8] FIG. 10 is a diagram illustrating an example of camera calibration processing. [Figure 9] 6 is a flowchart illustrating details of the camera calibration process of the imaging device in step S3 of FIG. 5. [Figure 10] FIG. 10 is a diagram illustrating an example of camera calibration processing. [Figure 11] FIG. 10 is a diagram illustrating an example of camera calibration processing. [Figure 12]6 is a flowchart illustrating details of the volumetric imaging setting process in step S4 of FIG. 5. [Figure 13] FIG. 10 is a diagram illustrating an example of capability information. [Figure 14] 6 is a flowchart illustrating details of the synchronized photographing process of the photographing devices in step S5 of FIG. 5. [Figure 15] 6 is a flowchart illustrating the details of the synchronized photographing process of the other photographing devices in step S5 of FIG. 5. [Figure 16] 6 is a flowchart illustrating details of the offline modeling process in step S7 of FIG. 5. [Figure 17] FIG. 10 is a diagram showing an example of 3D model data of an object. [Figure 18] 6 is a flowchart illustrating details of the content playback process in step S8 of FIG. 5. [Figure 19] 6 is a flowchart illustrating details of the real-time modeling reproduction process in step S9 of FIG. 5. [Figure 20] 10 is a flowchart illustrating an auto-calibration process. [Figure 21] FIG. 10 is a diagram illustrating another example of the camera calibration process. [Figure 22] 10 is a flowchart illustrating a camera calibration process performed by the control device. [Figure 23] FIG. 10 is a diagram illustrating an example of a feedback screen. [Figure 24] 10 is a flowchart illustrating a camera calibration process performed by a cloud server. [Figure 25] 25 is a flowchart illustrating details of the camera parameter calculation process in step S224 of FIG. 24. [Figure 26] FIG. 1 is a block diagram illustrating a configuration example of an embodiment of a computer to which the present disclosure is applied. DETAILED DESCRIPTION OF THE INVENTION

[0016] Hereinafter, with reference to the accompanying drawings, a description will be given of an embodiment of the present technology. In this specification and the drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted. The description will be given in the following order. 1. Example of image processing system configuration 2. Volumetric Capture and Playback Overview 3. Detailed configuration example of each device in the image processing system 4. Volumetric shooting playback processing flowchart 5. Flowchart of imaging device grouping process 6. Flowchart of camera calibration process for imaging device 7. Flowchart of shooting setting process for volumetric shooting 8. Flowchart of synchronized shooting process of shooting devices 9. Flowchart of offline modeling process 10. Content playback process flowchart 11. Flowchart of real-time modeling playback process 12. Flowchart of auto-calibration process 13. Camera calibration process without using calibration board images 14. Computer Configuration Example

[0017] <1. Example of image processing system configuration> FIG. 1 shows an example of the configuration of an image processing system according to the present disclosure, which provides volumetric image capture and playback using an image capture device such as a smartphone.

[0018] The image processing system 1 includes N (N>1) image capturing devices 11, a cloud server (server device) 12, and a playback device 13.

[0019] Each of the N imaging devices 11 performs volumetric imaging of a predetermined subject as an object OBJ, and transmits the resulting image data to the cloud server 12. The imaging devices 11 are, for example, electronic devices with imaging capabilities, such as smartphones, tablets, digital cameras, and game consoles. In the example of FIG. 1, N=5, that is, five imaging devices 11-1 to 11-5 are shown capturing images of the object OBJ, but the number of imaging devices 11 is arbitrary.

[0020] The cloud server 12 generates 3D model data (3D model data) MO of the object OBJ based on the image data obtained by volumetric photography. In addition, in response to a request from the playback device 13, the cloud server 12 transmits the stored 3D model data MO of a specific object OBJ to the playback device 13.

[0021] The playback device 13 acquires and plays back 3D model data MO of a predetermined object OBJ from the cloud server 12, thereby generating a free viewpoint image of the 3D model of the object OBJ viewed from any viewing position, and displays the image on a predetermined display. The playback device 13 is configured with an electronic device equipped with a display, such as a smartphone or tablet 13A, a personal computer 13B, or a head-mounted display (HMD) 13C.

[0022] The N imaging devices 11 and the cloud server 12 are connected via a predetermined network. The cloud server 12 and the playback device 13 are also connected via a predetermined network. The network connecting them is any communication network, and may be a wired communication network, a wireless communication network, or a combination of both. The network may be a single communication network or multiple communication networks. For example, the network may include any communication network or communication channel of any communication standard, such as the Internet, a public telephone network, a wide area communication network for wireless mobile devices such as a so-called 4G line or 5G line, a WAN (Wide Area Network), a LAN (Local Area Network), a wireless communication network that communicates in accordance with the Bluetooth (registered trademark) standard, a short-range wireless communication channel such as NFC (Near Field Communication), an infrared communication channel, or a wired communication network in accordance with standards such as HDMI (registered trademark) (High-Definition Multimedia Interface) or USB (Universal Serial Bus).

[0023] <2. Volumetric Shooting and Playback Overview> Volumetric imaging and playback will now be briefly described with reference to FIGS.

[0024] For example, as shown in FIG. 2, a predetermined shooting space in which a subject such as a person is placed is imaged from the periphery by a plurality of camera devices, thereby obtaining a plurality of captured images. The captured images are, for example, composed of moving images. In the example of FIG. 2, three camera devices CAM1 to CAM3 are arranged to surround the subject #Ob1, but the number of camera devices CAM is not limited to three and can be any number. The number of camera devices CAM during shooting is the number of known viewpoints when generating a free-viewpoint image, so the more camera devices CAM there are, the more accurately the free-viewpoint image can be expressed. The subject #Ob1 is assumed to be a person performing a predetermined action.

[0025] Using images captured from multiple camera devices CAM in different directions, a 3D object MO1, which is a 3D model of the subject #Ob1 to be displayed in the captured space, is generated (3D modeling). For example, the 3D object MO1 is generated using a technique such as Visual Hull, which carves out the three-dimensional shape of the subject using images captured in different directions.

[0026] Then, data (3D model data) of one or more 3D objects among the one or more 3D objects present in the shooting space is transmitted to a playback device and played back. That is, the playback device renders the 3D objects based on the acquired 3D object data, and a 3D shape image is displayed on the viewer's viewing device. Figure 2 shows an example in which the viewing device is a display D1 or a head-mounted display (HMD) D2.

[0027] The playback side can request only the 3D object to be viewed from one or more 3D objects present in the shooting space and display it on the viewing device. For example, the playback side can imagine a virtual camera whose shooting range is the viewer's viewing range, and request only the 3D objects captured by the virtual camera from the many 3D objects present in the shooting space and display them on the viewing device. The viewpoint of the virtual camera (virtual viewpoint) can be set to any position so that the viewer can view the subject from any viewpoint in the real world. A background image representing a specified space can be composited with the 3D object as appropriate.

[0028] Figure 3 shows an example of a general data format for 3D model data.

[0029] 3D model data is generally expressed by 3D shape data that represents the 3D shape (geometry information) of the subject, and texture data that represents color information of the subject.

[0030] 3D shape data is expressed, for example, in a point cloud format, which represents the three-dimensional position of a subject as a collection of points; a 3D mesh format called a polygon mesh, which represents the vertices and the connections between the vertices; or a voxel format, which represents the data as a collection of cubes called voxels.

[0031] Texture data can be in a multi-texture format, which is stored in the captured images (2D texture images) captured by each camera CAM, or in a UV mapping format, which is stored in a 2D texture image that is attached to each point or polygon mesh, which is 3D shape data, expressed in a UV coordinate system.

[0032] As shown in the upper part of Figure 3, the format for describing 3D model data is a view-dependent format in which color information can change depending on the virtual viewpoint (position of the virtual camera), with 3D shape data and a multi-texture format held in multiple captured images P1 to P8 captured by each camera CAM.

[0033] In contrast, the format for describing 3D model data, as shown in the bottom row of Figure 3, is a view-independent format in which the color information is the same depending on the virtual viewpoint (position of the virtual camera), using a UV mapping format in which 3D shape data and the texture information of the subject are mapped to a UV coordinate system.

[0034] Volumetric photography, which is the photography used to generate 3D model data, has traditionally been carried out in dedicated studios using specialized equipment.

[0035] However, the image processing system 1 in FIG. 1 is a system that enables volumetric photography to be performed using electronic devices such as smartphones and tablets that are commonly owned by users.

[0036] <3. Detailed configuration example of each device in the image processing system> FIG. 4 is a block diagram showing an example of a detailed configuration of the image capturing device 11, the cloud server 12, and the playback device 13. As shown in FIG.

[0037] (Imaging device 11) The photographing device 11 includes a communication unit 31 , a message transmission / reception unit 32 , a control unit 33 , an audio output unit 34 , a speaker 35 , an image output unit 36 ​​, a display 37 , a flash output unit 38 , and a flash 39 .

[0038] The photographing device 11 also includes a touch sensor 40 , a gyro sensor 41 , an acceleration sensor 42 , a GPS sensor 43 , a sensor input unit 44 , and a synchronization signal generation unit 45 .

[0039] The photographing device 11 further includes one or more cameras 51 (51A to 51C), a microphone 52, a camera input / output unit 53, an image processing unit 54, an image compression unit 55, an audio input unit 56, an audio processing unit 57, an audio compression unit 58, and a stream transmission unit 59.

[0040] The communication unit 31 is composed of various communication modules such as carrier communication for wireless mobile devices such as so-called 4G lines and 5G lines, wireless communication such as Wi-Fi (registered trademark), and wired communication such as 1000BASE-T, and communicates messages and data with the cloud server 12.

[0041] Message transmitting / receiving unit 32 communicates messages with cloud server 12 via communication unit 31. Message transmitting / receiving unit 32 supports an instant messaging system that allows message communication with cloud server 12, regardless of the type of communication of communication unit 31. Examples of such instant messaging systems include Jabber, XMPP, and SIP (Session Initiation Protocol).

[0042] The control unit 33 controls the overall operation of the imaging device 11 based on messages received by the message transmission / reception unit 32 and user operations detected by an operation unit (not shown). For example, the control unit 33 communicates messages with the cloud server 12 via the message transmission / reception unit 32 to notify the user of the operation, or starts volumetric imaging (capturing an image of a subject) in response to a request from the cloud server 12 and transmits image data obtained by the imaging to the cloud server 12. The control unit 33 also transmits capability information of the imaging device 11 to the cloud server 12 and supplies sensor information, which is the detection results of various sensors provided in the imaging device 11, to the stream transmission unit 59. The control unit 33 supplies setting value information of the imaging function provided by the cloud server 12 to the camera input / output unit 53, the image processing unit 54, and the image compression unit 55, and performs predetermined settings. The setting value information of the imaging function includes setting values ​​for, for example, exposure time, resolution, compression method, bit rate, etc.

[0043] An application program for performing volumetric photography (hereinafter referred to as a volumetric photography app) is installed in the photography device 11, and by launching and executing the volumetric photography app, the control unit 33 performs processing to control the overall operation of the photography device 11.

[0044] The audio output unit 34 outputs an audio signal to the speaker 35 under the control of the control unit 33. The speaker 35 outputs sound based on the audio signal supplied from the audio output unit 34. The image output unit 36 ​​outputs an image signal to the dispenser 37 under the control of the control unit 33. The display 37 displays an image based on the image signal supplied from the image output unit 36. The flash output unit 38 outputs a light emission control signal to the flash 39 under the control of the control unit 33. The flash 39 emits light based on the light emission control signal from the flash output unit 38.

[0045] The touch sensor 40 detects the touch position when the user touches the display 37 and supplies the detected position to the sensor input unit 44 as sensor information. The gyro sensor 41 detects angular velocity and supplies the detected position to the sensor input unit 44 as sensor information. The acceleration sensor 42 detects acceleration and supplies the detected position to the sensor input unit 44 as sensor information. The GPS sensor 43 receives a signal from a GPS, which is one of the GNSS (Global Navigation Satellite Systems). The GPS sensor 43 detects the current position of the imaging device 11 based on the received GPS signal and supplies the detected position to the sensor input unit 44 as sensor information. The sensor input unit 44 acquires the sensor information supplied from the touch sensor 40, gyro sensor 41, acceleration sensor 42, and GPS sensor 43 and supplies the information to the control unit 33.

[0046] The synchronization signal generation unit 45 generates a synchronization signal based on an instruction from the control unit 33 and supplies it to one or more cameras 51. The synchronization signal generation unit 45 can use the following methods to generate a synchronization signal, depending on the type of network connecting the N image capturing devices 11 and the cloud server 12.

[0047] The first synchronization method is to synchronize using clock information from carrier communications such as 5G lines. Carrier communications have highly accurate clocks for communication, so this clock information can be used to adjust the time and achieve synchronization.

[0048] The second synchronization method is to use the time information contained in GPS signals. GPS signals contain time information with such high precision that they are used as the grandmaster clock of the Precision Time Protocol (PTP). This clock information can be used to adjust the time and achieve synchronization.

[0049] The third synchronization method is to synchronize time using multicast communication over Wi-Fi wireless communication. When connected to the same access point, multicast packets are sent to detect the timing and achieve synchronization. 802.11AC and WiFi Time Sync can also be used.

[0050] The fourth synchronization method is to synchronize time using multicast communication over carrier communications. Because 5G carrier communications have low latency down to 1 ms, multicast packets can be sent over carrier communications to detect timing and achieve synchronization.

[0051] Cameras 51A to 51C are of different camera types (kinds), for example, camera 51A is an RGB camera that generates an RGB image based on the light reception results of receiving visible light (RGB), camera 51B is an RGB-D camera that generates a depth image that stores a depth value as distance information to the subject as the pixel value of each pixel along with the RGB image, and camera 51C is an IR camera that generates an IR image based on the light reception results of receiving infrared light (IR). When there is no particular distinction between RGB images, depth images, and IR images, they are referred to as camera images.

[0052] Cameras 51A to 51C perform predetermined settings for exposure time, gain, resolution, etc. based on setting value information supplied from control unit 33 via camera input / output unit 53. Cameras 51A to 51C also output camera images obtained as a result of shooting to camera input / output unit 53.

[0053] The microphone 52 collects ambient sounds and outputs them to the audio input unit 56 .

[0054] The camera input / output unit 53 supplies the setting value information supplied from the control unit 33 to each of the cameras 51A to 51C. The camera input / output unit 53 also supplies the camera images supplied from the cameras 51A to 51C to the image processing unit .

[0055] The image processing unit 54 performs predetermined image signal processing such as demosaic processing, color correction processing, distortion correction processing, and color space conversion processing on the camera image (RAW data) supplied from the camera input / output unit 53.

[0056] The image compression unit 55 performs a predetermined compression encoding process on the image signal from the image processing unit 54 based on setting values ​​specified by the control unit 33, and supplies the compressed and encoded image stream to the stream transmission unit 59. The setting values ​​specified by the control unit 33 include, for example, parameters such as the compression method and bit rate.

[0057] The stream transmission unit 59 transmits the image stream from the image compression unit 55, the audio stream from the audio compression unit 58, and the sensor information from the control unit 33 to the cloud server 12 via the communication unit 31. The sensor information from the control unit 33 is stored in the image stream of the camera image for each frame and transmitted, for example.

[0058] The audio input unit 56 acquires audio input from the microphone 52 and supplies it to the audio processing unit 57. The audio processing unit 57 performs predetermined audio signal processing, such as noise removal processing, on the audio signal from the audio input unit 56. The audio compression unit 58 performs predetermined compression encoding processing on the audio signal from the audio processing unit 57 based on setting values ​​specified by the control unit 33, and supplies the compressed and encoded audio stream to the stream transmission unit 59. The setting values ​​specified by the control unit 33 include, for example, parameters such as the compression method and bit rate.

[0059] (Cloud Server 12) The cloud server 12 includes a controller 101 , a message transmitting / receiving unit 102 , a communication unit 103 , a stream receiving unit 104 , a calibration unit 105 , a modeling task generating unit 106 , and a task storage unit 107 .

[0060] The cloud server 12 also includes an offline modeling unit 108 , a content management unit 109 , a content storage unit 110 , a real-time modeling unit 111 , a stream transmission unit 112 , and an auto-calibration unit 113 .

[0061] The controller 101 controls the overall operation of the cloud server 12. For example, the controller 101 transmits a predetermined message to each imaging device 11 or playback device 13 via the message transmission / reception unit 102, thereby causing the device to perform a predetermined operation such as an imaging operation or a playback operation. The controller 101 also controls the offline modeling unit 108 and the like to perform offline modeling, and controls the real-time modeling unit 111 and the like to perform real-time modeling. In offline modeling, the generation of a 3D model by volumetric imaging and the playback of the 3D model based on the generated 3D model data (display of a free viewpoint image) are performed at different times. On the other hand, in real-time modeling, the generation of a 3D model and the playback of the 3D model based on the generated 3D model data are performed as a series of processes. The voxel size and bounding box size, which are modeling parameters used when generating a 3D model, are determined by the controller 101 according to the imaging target area of ​​each imaging device 11 and user settings. The voxel size represents the size of a voxel, and the bounding box size represents the processing range in which a search is performed to determine whether a voxel of a 3D object exists.

[0062] The message transmitting / receiving unit 102 communicates messages with the image capturing device 11 or the playback device 13 via the communication unit 103. The message transmitting / receiving unit 102 corresponds to the message transmitting / receiving unit 32 of the image capturing device 11, for example.

[0063] The communication unit 103 is composed of various communication modules such as carrier communication, wireless communication such as Wi-Fi (registered trademark), and wired communication such as 1000BASE-T, and communicates messages and data with the imaging device 11 or playback device 13.

[0064] The stream receiving unit 104 receives the image streams and audio streams transmitted from each imaging device 11 via the communication unit 103 and performs decoding processing corresponding to the predetermined compression encoding method executed by the imaging device 11. The stream receiving unit 104 supplies the camera image or audio signal after decoding processing to at least one of the calibration unit 105, the modeling task generation unit 106, and the real-time modeling unit 111. The stream receiving unit 104 also supplies sensor information that has been stored in the image stream and transmitted to the calibration unit 105, for example.

[0065] The calibration unit 105 performs calibration based on the camera images and sensor information of each imaging device 11 in accordance with instructions from the controller 101, and calculates the camera parameters of each camera 51. The camera parameters include external parameters and internal parameters, but if the internal parameters are set to fixed values, only the external parameters are calculated. The calculated camera parameters are supplied to the modeling task generation unit 106, the offline modeling unit 108, the real-time modeling unit 111, and the auto-calibration unit 113.

[0066] In response to an instruction from the controller 101, the modeling task generation unit 106 generates a modeling task, which is a task for generating a 3D model, based on the image stream from the stream reception unit 104 and the camera parameters from the calibration unit 105, and stores the generated modeling task in the task storage unit 107. The task storage unit 107 stores the modeling tasks sequentially supplied from the modeling task generation unit 106 as a task queue.

[0067] The offline modeling unit 108 sequentially retrieves one or more modeling tasks stored as a task queue in the task storage unit 107, and executes offline 3D modeling processing. Through the 3D modeling processing, 3D model data of the object is generated and supplied to the content management unit 109. During the 3D modeling processing, the offline modeling unit 108 may request the auto-calibration unit 113 to perform calibration processing at that time, as necessary, based on sensor information associated with image data of each frame of a predetermined imaging device 11, and may update the calibration information.

[0068] The content management unit 109 stores and manages the 3D model data of the object generated by the offline modeling unit 108 as content in the content storage unit 110. The content management unit 109 acquires 3D model data of a predetermined object, which is content specified by the controller 101, from the content storage unit 110 and transmits the 3D model data to the playback device 13 via the stream transmission unit 112. The content storage unit 110 stores the 3D model data of the object, which is content.

[0069] The real-time modeling unit 111 acquires the camera image supplied from the stream receiving unit 104 and executes 3D modeling processing in real time. By the 3D modeling processing, data of the 3D model of the object (3D model data) is generated frame by frame and sequentially supplied to the stream transmitting unit 112. In the 3D modeling processing, the real-time modeling unit 111 may request the auto-calibration unit 113 to perform calibration processing at that time as needed, based on sensor information associated with image data of each frame of a predetermined imaging device 11, and may update the calibration information.

[0070] The stream sending unit 112 sends the 3D model data of the object supplied from the content management unit 109 or the real-time modeling unit 111 to the playback device 13 .

[0071] When it is determined that the camera parameters calculated by the calibration unit 105 need to be updated based on the sensor information associated with the image data of each frame from the offline modeling unit 108 or the real-time modeling unit 111, the auto-calibration unit 113 executes a calibration process and updates the camera parameters.

[0072] (Playback Device 13) The playback device 13 includes a communication unit 151, a message transmission / reception unit 152, a control unit 153, a sensor 154, a sensor input unit 155, a stream reception unit 156, a playback unit 157, an image output unit 158, a display 159, an audio output unit 160, and a speaker 161.

[0073] The communication unit 151 is configured with various communication modules for carrier communication, wireless communication such as Wi-Fi (registered trademark), wired communication such as 1000BASE-T, and the like, and communicates messages and data with the cloud server 12.

[0074] Message transmitting / receiving unit 152 communicates messages with cloud server 12 via communication unit 103. Message transmitting / receiving unit 152 corresponds to message transmitting / receiving unit 102 of cloud server 12.

[0075] The control unit 153 controls the overall operation of the playback device 13 based on messages received by the message transmission / reception unit 152 and viewer operations detected by an operation unit (not shown). For example, the control unit 153 causes the message transmission / reception unit 152 to transmit a message requesting predetermined content based on the viewer's operation. The control unit 153 also causes the playback unit 157 to play back 3D model data of an object that is transmitted from the cloud server 12 in response to the content request. When causing the playback unit 157 to play back the 3D model data of the object, the control unit 153 controls the playback unit 157 to play back the 3D model data of the object based on sensor information supplied from the sensor input unit 155.

[0076] The sensor 154 is configured with a gyro sensor and an acceleration sensor, detects the viewing position of the viewer, and supplies the same as sensor information to the sensor input unit 155. The sensor input unit 155 supplies the sensor information supplied from the sensor 154 to the control unit 153.

[0077] The stream receiving unit 156 is a receiving unit corresponding to the stream sending unit 112 of the cloud server 12 , and receives the 3D model data of the object sent from the cloud server 12 , and supplies it to the playback unit 157 .

[0078] The playback unit 157 plays back the 3D model of the object from the 3D model data of the object so that it fits within the viewing range based on the viewing position supplied from the control unit 153. A free viewpoint image of the 3D model obtained as a result of the playback is supplied to the image output unit 158, and the audio is supplied to the audio output unit 160.

[0079] The image output unit 158 ​​supplies the free viewpoint image from the playback unit 157 to the display 159, which displays the free viewpoint image. The display 159 displays the free viewpoint image supplied from the image output unit 158.

[0080] The audio output unit 160 outputs the audio signal from the reproduction unit 157 from the speaker 161. The speaker 161 outputs sound based on the audio signal from the audio output unit 160.

[0081] The image capturing device 11, the cloud server 12, and the playback device 13 of the image processing system 1 each have the configuration described above.

[0082] The processing performed by the image processing system 1 will be described in detail below.

[0083] <4. Volumetric shooting playback processing flowchart> FIG. 5 is a flowchart of the volumetric shooting and playback process performed by the image processing system 1.

[0084] First, in step S1, the image processing system 1 executes a grouping process for the image capturing devices 11, which groups a plurality of (N) image capturing devices 11 participating in volumetric image capturing into one group. Details of this process will be described later with reference to FIGS. 6 and 7.

[0085] In step S2, the image processing system 1 executes a calibration shooting setting process for setting the shooting device 11 for performing the camera calibration process. Each shooting device 11 sets the exposure time, resolution, etc. to predetermined values ​​according to its own shooting functions and settable ranges, for example. The setting values ​​for this camera calibration process may be determined in advance, or may be set (selected) by the user based on its own capability information, as in the volumetric shooting setting process in step S4 described below.

[0086] In step S3, the image processing system 1 executes a camera calibration process for the imaging devices 11 to calculate the camera parameters of each of the plurality of imaging devices 11. Details of this process will be described later with reference to FIGS.

[0087] In step S4, the image processing system 1 executes a volumetric imaging setting process for setting the imaging device 11 for volumetric imaging. Details of this process will be described later with reference to FIGS.

[0088] In step S5, the image processing system 1 executes synchronized photographing processing for the photographing devices 11, in which the photographing timings of the plurality of photographing devices 11 are synchronized to perform photographing. Details of this processing will be described later with reference to FIGS. 14 and 15.

[0089] In step S6, the image processing system 1 determines whether to perform offline modeling or real-time modeling. Offline modeling is a 3D modeling process in which a 3D model is generated at a timing separate from volumetric imaging, while real-time modeling is a 3D modeling process in which a 3D model is generated in synchronization with volumetric imaging. Whether to perform offline modeling or real-time modeling is determined by a user's designation for a specific imaging device 11 when grouping the imaging devices 11, for example.

[0090] If it is determined in step S6 that offline modeling is to be performed, the following processes of steps S7 and S8 are executed.

[0091] In step S7, the image processing system 1 executes offline modeling processing to generate 3D model data based on the image stream of the subject captured by the multiple imaging devices 11. Details of this processing will be described later with reference to FIGS. 16 and 17.

[0092] In step S8, the image processing system 1 executes a content playback process in which the 3D model data generated in the offline modeling process is used as content and a free viewpoint image of the 3D model is displayed on the playback device 13. Details of this process will be described later with reference to FIG.

[0093] On the other hand, if it is determined in step S6 that real-time modeling is to be performed, the following processing in step S9 is executed.

[0094] In step S9, the image processing system 1 generates 3D model data based on the image stream of the subject captured by the multiple imaging devices 11, transmits the generated 3D model data to the playback device 13, and executes real-time modeling playback processing to display a free viewpoint image of the 3D model. Details of this processing will be described later with reference to FIG. 19.

[0095] This completes the volumetric shooting and playback process by the image processing system 1.

[0096] The processes of the steps in FIG. 5 may be executed consecutively, or may be executed at different times with a predetermined time interval between each step.

[0097] <5. Flowchart of imaging device grouping process> Next, with reference to the flowchart of FIG. 6, the grouping process of the imaging devices 11 executed as step S1 of FIG. 5 will be described in detail.

[0098] Of the multiple imaging devices 11 participating in the volumetric imaging, one predetermined imaging device 11 is referred to as a master device, and the other imaging devices 11 are referred to as participating devices.

[0099] First, in step S21, the photographing device 11, which is the master device, starts a volumetric photographing app, inputs a user ID, a password, and a group name, and transmits a group registration request to the cloud server 12. The user ID may be input by the user, or may be input automatically (without user input) using the terminal name or MAC address of the photographing device 11. When the group registration request is transmitted to the cloud server 12, location information of the master device is also transmitted to the cloud server 12. The location information may be information that allows determination in step S24, described later, of whether the multiple photographing devices 11 constituting the group are capable of volumetric photographing. The location information may be, for example, moving longitude information detected by the GPS sensor 43, latitude and longitude information obtained by triangulation using the electric field strength of a base station of a carrier communication, or latitude and longitude information obtained by transmitting Wi-Fi wireless communication access point information (e.g., SSID and MAC address information) to a Wi-Fi location information service.

[0100] In step S22, the master device acquires the two-dimensional code returned from the cloud server 12 in response to the group registration request, and displays it on the display 37.

[0101] FIG. 7 shows an example of a screen of a two-dimensional code displayed on the display 37 of the master device in step S22.

[0102] 7 displays the user ID (USERID), password (PASSWD), and group name (GroupName) entered by the user in step S21, along with the text "Volumetric Modeling Service," which is the name of the service provided by cloud server 12. Also displayed is a two-dimensional code along with the text "Please take a picture of the following two-dimensional code with your camera."

[0103] In step S23, each participating device, which is an imaging device 11 other than the master device, launches a volumetric imaging app and captures the two-dimensional code displayed on the display 37 of the master device. The participating device that captured the two-dimensional code transmits group identification information that identifies the registered group and its own location information to the cloud server 12. Alternatively, when a participating device captures the two-dimensional code, the volumetric imaging app may be automatically launched, and the group identification information and location information of the participating device may be transmitted to the cloud server 12. The group identification information may be, for example, a user ID or a group name.

[0104] In step S24, the cloud server 12 determines whether the master device and participating devices in the same group are capable of volumetric imaging, using the group identification information and position information transmitted from the participating devices. The controller 101 of the cloud server 12 determines whether the participating device is capable of volumetric imaging, for example, based on whether the participating device is located in a nearby position within a certain range from the position of the master device. If the participating device is located in a nearby position within a certain range from the position of the master device, the cloud server 12 determines that volumetric imaging is possible.

[0105] If it is determined in step S24 that the participating device is capable of volumetric photography, the cloud server 12 registers the participating device in the group indicated by the group identification information, i.e., the group of the master device, and sends a registration completion message to the participating device.

[0106] On the other hand, if it is determined in step S24 that the participating device is not capable of volumetric imaging, the cloud server 12 transmits a group registration error to the participating device.

[0107] In step S27, the image processing system 1 determines whether the group registration is complete, and if it is determined that the group registration is not yet complete, the processes of steps S23 to S27 described above are repeated. For example, if an operation to end the display of the two-dimensional code screen in Fig. 7 and end the registration work is performed on the master device, it is determined that the group registration is complete.

[0108] If it is determined in step S27 that the group registration is complete, the grouping process of the imaging device 11 in FIG. 6 ends.

[0109] In the above-described grouping process, the group identification information is obtained by recognizing the two-dimensional barcode and transmitted to the cloud server 12, but the group identification information may also be input by the user and transmitted to the cloud server 12.

[0110] Alternatively, group registration of participating devices may be performed by an operation similar to the WPS button of Wi-Fi (registered trademark). As a specific process, for example, when the master device presses a group registration button on the screen, a message indicating that group registration is now open and location information are transmitted to the cloud server 12. When a user of a participating device presses the group registration button on the screen, a message indicating that the group registration button has been pressed and location information are transmitted from the participating device to the cloud server 12. Within a certain time after the group registration button on the master device is pressed, the cloud server 12 checks the location information of the participating device whose group registration button was pressed, determines whether the participating device can be registered in the group of the master device, and transmits a registration completion message or a group registration error to the participating device.

[0111] Furthermore, in the grouping process described above, whether the participating devices are capable of volumetric imaging is determined based only on the position information of the master device and the participating devices, but the determination may be made based on other information.

[0112] For example, the determination may be made based on whether the master device and participating devices are connected to the same access point. Alternatively, for example, the master device and participating devices may each capture an image of the same subject, and volumetric photography may be determined to be possible if the captured image shows the same subject. Alternatively, for example, the master device may capture an image of a participating device, and volumetric photography may be determined to be possible if the captured image shows a specific portion of the participating device (e.g., a display image on a display). Conversely, the participating device may capture an image of the master device. Furthermore, the determination may be made by combining two or more of the above-mentioned determinations.

[0113] The above-described grouping process allows users to easily register common electronic devices that they normally use, such as smartphones and tablets, as imaging devices 11 that perform volumetric imaging, enabling them to operate in conjunction with other imaging devices 11 registered in the group.

[0114] <6. Flowchart of camera calibration process for imaging device> Next, the camera calibration process of the image capturing device 11, which is executed as step S3 in FIG. 5, will be described in detail.

[0115] In dedicated volumetric photography studios, calibration boards are prepared in advance for camera calibration, but it is necessary to be able to perform calibration even without such specially prepared boards.

[0116] 8, in the image processing system 1, a calibration board image is displayed on one photographing device 11, and all the other photographing devices 11 photograph the displayed calibration board image and detect feature points, and this process is performed sequentially on all the grouped photographing devices 11, thereby executing a camera calibration process for calculating the camera parameters of the photographing device 11. As the camera 51 of the photographing device 11, a camera 51 on the same surface as the display 37 is used.

[0117] 8, of the five photographing devices 11, a calibration board image is displayed on the photographing device 11-5, and the other photographing devices 11-1 to 11-4 photograph the calibration board image of the photographing device 11-5 and detect feature points. The photographing devices 11-1 to 11-4 also display calibration board images in turn, and the other photographing devices 11 photograph the displayed calibration board images and detect feature points.

[0118] The camera calibration process of the image capturing device 11, which is executed as step S3 in FIG. 5, will be described in detail with reference to the flowchart in FIG.

[0119] First, in step S41, the cloud server 12 selects one photographing device 11 that will display the calibration board image. The cloud server 12 transmits a command (message) to the selected photographing device 11 (hereinafter also referred to as the selected photographing device 11) to display the calibration board image.

[0120] In step S42, the selected photographing device 11 receives the command to display the calibration board image from the cloud server 12, and displays the calibration board image on the display 37.

[0121] In step S43, each of the photographing devices 11 other than the selected photographing device 11 photographs the calibration board image displayed on the display 37 of the selected photographing device 11 in synchronization with each other.

[0122] In step S44, each of the photographing devices 11 other than the selected photographing device 11 detects feature points of the photographed calibration board image, and transmits feature point information, which is information on each of the detected feature points, to the cloud server 12.

[0123] In step S45, the calibration unit 105 of the cloud server 12 acquires the feature point information of the calibration board image transmitted from each imaging device 11 via the communication unit 103 etc., and stores it.

[0124] In step S46, the cloud server 12 determines whether the calibration board image has been displayed on all of the grouped imaging devices 11.

[0125] If it is determined in step S46 that the calibration board image is not being displayed on all of the photographing devices 11, the process returns to step S41, and the processes of steps S41 to S46 described above are executed again. That is, one of the photographing devices 11 that has not yet displayed the calibration board image is selected, and the other photographing devices 11 each capture the displayed calibration board image, detect feature points, and transmit the feature point information to the cloud server 12.

[0126] On the other hand, if it is determined in step S46 that the calibration board image has been displayed on all of the photographing devices 11, the processing proceeds to step S47, and the calibration unit 105 of the cloud server 12 uses the stored feature point information to estimate the three-dimensional positions of the feature points and the internal and external parameters of the photographing devices 11 for each photographing device 11.

[0127] Methods for calculating camera parameters of multiple imaging devices from images captured by the devices include, for example, bundle adjustment and algorithms for solving nonlinear optimization problems, as disclosed in, for example, "Iwamoto Yuki, Sugaya Yasuyuki, Kanaya Kenichi. Implementation and Evaluation of Bundle Adjustment for 3D Reconstruction." Computer Vision and Image Media (CVIM) 2011.19 (2011): 1-8."

[0128] In step S48, the calibration unit 105 of the cloud server 12 supplies the estimated internal parameters and external parameters of each imaging device 11 to the offline modeling unit 108 and the real-time modeling unit 111, and the camera calibration process ends.

[0129] By the above-described camera calibration process, it is possible to calculate the camera parameters of each of the grouped imaging devices 11 even if a specially prepared calibration board or the like does not exist.

[0130] In the above-described camera calibration process, the multiple grouped photographing devices 11 are assumed to display the same calibration board image, and the multiple photographing devices 11 are caused to display the calibration board image in sequence.

[0131] However, when the calibration board images displayed on the grouped multiple imaging devices 11 are different as shown in Fig. 10, the multiple imaging devices 11 simultaneously display the calibration board images, and each imaging device 11 can simultaneously detect feature points of the calibration board images displayed on the other imaging devices 11. The calibration board images may be images with different pattern shapes or different colors. The calibration board images shown in Fig. 10 are examples of chess pattern board images and dot pattern board images, and are images with different numbers of grids (MxN, KxL) and dot sizes and densities.

[0132] Alternatively, one of the grouped multiple photographing devices 11, such as photographing device 11-6 in Figure 11, may be used to display a calibration board image, and the user may take photographs with each photographing device 11 while moving photographing device 11-6 on which the calibration board image is displayed like a calibration board, thereby detecting feature points of the calibration board image.

[0133] <7. Flowchart of shooting setting process for volumetric shooting> Next, the volumetric imaging setting process executed as step S4 in FIG. 5 will be described in detail with reference to the flowchart in FIG.

[0134] First, in step S61, each of the grouped imaging devices 11 transmits its own capability information to the cloud server 12.

[0135] The capability information is information about the photographing functions and settable ranges of the photographing device 11, and may include information about the following items, for example. 1. Camera type (RGB, RGB-D, IR) 2. Camera settings 1 (exposure time, gain, zoom) 3. Camera settings 2 (shooting resolution, ROI, frame rate) 4. Image encoding settings (output resolution, compression encoding method, bit rate) 5. Camera synchronization method (Type 1, Type 2, Type 3) 6. Camera calibration method (Type 1, Type 2, Type 3)

[0136] The ROI in camera setting 2 represents the range that can be set as the region of interest within the shooting range indicated by the shooting resolution.

[0137] FIG. 13 shows an example of capability information that a certain imaging device 11 has.

[0138] According to the capability information in FIG. 13, the imaging device 11 has a device identification name of “Camera 1”, a camera type of “RGB”, an exposure time that can be set to any of “500”, “1000”, or “10000”, a frame rate that can be set to any of “30”, “60”, or “120”, and an imaging resolution that can be set to any of “3840x2160”, “1920x1080”, or “640x480”.

[0139] In addition, with regard to image encoding settings, the photographing device 11 can set the compression encoding method to either "H.264" or "H.265", the output resolution to either "3840x2160", "1920x1080", or "640x480", and the bit rate to either "1M", "3M", "5M", "10M", "20M", or "50M".

[0140] Furthermore, the imaging device 11 can set the camera synchronization method to one of "Type 1", "Type 2", or "Type 3", and can set the calibration method to one of "Type 1", "Type 2", or "Type 3".

[0141] In step S62, the cloud server 12 generates candidate setting values ​​for each imaging device 11 suitable for volumetric imaging based on the capability information of all imaging devices 11 set in one group. The setting values ​​for each imaging device 11 may include common setting values ​​for all imaging devices 11 set in one group, and different setting values. For example, the camera synchronization method may be set in common for all imaging devices 11, but the camera type may be set separately, such as RGB for some imaging devices 11 and IR for the remaining imaging devices 11.

[0142] In step S63, the cloud server 12 determines whether volumetric photography, i.e., photography for generating a 3D model, is possible. Whether volumetric photography is possible or not will be described in detail later, but for example, it can be determined by superimposing the photography target areas of all the photography devices 11 set in one group and checking whether there is an area common to a certain number or more of the photography devices 11.

[0143] If it is determined in step S63 that volumetric photography is possible, processing proceeds to step S64, in which the cloud server 12 transmits candidate setting values ​​to each photographing device 11, and each photographing device 11 receives the candidate setting values ​​and presents them to the user.

[0144] The user refers to the displayed setting value candidates and selects a desired setting value candidate. Then, in step S65, each imaging device 11 transmits the setting value candidate designated by the user to the cloud server 12.

[0145] In step S66, the cloud server 12 acquires candidate setting values ​​specified by the user transmitted from each photographing device 11, transmits a setting signal to each photographing device 11 to set the setting values ​​specified by the user, and terminates the photographing setting process for volumetric photographing.

[0146] On the other hand, if it is determined in step S63 that volumetric photography is not possible, the process proceeds to step S67, and the cloud server 12 transmits to each photography device 11 a message indicating that volumetric photography is not possible, and notifies the user of this. This ends the photography setting process for volumetric photography.

[0147] As described above, in the shooting setting process for volumetric shooting, it is determined whether volumetric shooting is possible based on the capability information of each shooting device 11, and a predetermined setting value is selected from among the setting values ​​that make volumetric shooting possible.

[0148] Below, we will explain an example of determining setting value candidates for each imaging device 11 and an example of determining whether volumetric imaging is possible based on the capability information of each imaging device 11. It is assumed that the capability information of each imaging device 11 has already been acquired. It is also assumed that the desired frame rate for generating 3D model data and the desired processing time, i.e., how many minutes of content to process, have been set as user request values.

[0149] [STEP 1] The controller 101 of the cloud server 12 first determines the camera synchronization method based on the capability information of each image capture device 11. For example, when the camera synchronization methods of the image capture devices 11-1 to 11-N satisfy the following conditions, the controller 101 selects the camera synchronization method with the highest priority among those supported by all of the image capture devices 11. When the priority is type1 > type2 > ... > type K, Type2 is selected. Photographic Device 11-1: Type1, Type2, ... TypeK Photographic Device 11-2: Type2, ... TypeK ... Photographic Device 11-N: Type1, Type2, ... TypeK

[0150] [STEP 2] Next, the controller 101 measures the communication bandwidth with each of the image capturing devices 11 and determines the maximum bit rate [Mbps]. It is assumed that the maximum bit rates determined for each of the image capturing devices 11-1 to 11-N are as follows: Shooting device 11-1: X0[Mbps] Shooting device 11-2: X1[Mbps] ... Camera 11-N: X n [Mbps]

[0151] [STEP 3] Next, the controller 101 calculates possible combinations of camera type, resolution, frame rate, compression encoding method, etc. within the range of the maximum bit rate of each imaging device 11, and determines candidate settings for each imaging device 11. Camera 11-1: Candidate 1 (RGB, 3840x2160, 30, H.265, 40M) Candidate 2 (RGB, 3840x2160, 60, H.265, 50M) Candidate 3 (RGB, 1920x1080, 30, H.265, 20M) Candidate 4 (RGB-D, 1920x1080, 30, H.265, 40M) Camera 11-2: Candidate 1 (RGB, 3840x2160, 30, H.265, 40M) Candidate 2 (RGB, 3840x2160, 60, H.265, 50M) Candidate 3 (RGB, 1920x1080, 120, H.265, 20M) Candidate 4 (RGB-D, 1920x1080, 30, H.265, 40M) ...

[0152] [STEP 4] The controller 101 determines whether volumetric imaging is possible using the determined candidates for each setting value. The controller 101 can determine whether volumetric imaging is possible by the following method.

[0153] For example, the controller 101 calculates the imaging target area of ​​each imaging device 11 based on the calibration result of each imaging device 11, and determines whether volumetric imaging is possible based on whether a common imaging area exists when the imaging target areas of each imaging device 11 are superimposed. The common imaging area may be an area common to the imaging target areas of all imaging devices 11, or may be an area common to the imaging target areas of a certain number or more of imaging devices 11.

[0154] For example, the controller 101 calculates modeling parameters of a 3D model based on the resolution and imaging target area of ​​each imaging device 11, and determines whether volumetric imaging is possible based on whether the calculated modeling parameters are within a predetermined range. The modeling parameters include a voxel size and a bounding box size. More specifically, the controller 101 determines the voxel size (mm square) to be used when generating a 3D model based on the resolution and imaging target area of ​​each imaging device 11. The controller 101 then determines whether volumetric imaging is possible based on whether the calculated voxel size is between the lower limit value Voxel_L and the upper limit value Voxel_U of the voxel size (Voxel_L<=Voxel size<=Voxel_U). The controller 101 may add multiple voxel sizes and bounding box sizes between the lower limit value Voxel_L and the upper limit value Voxel_U of the voxel size as setting value candidates.

[0155] Furthermore, for example, the controller 101 estimates the processing time required for 3D modeling from candidate setting values ​​for each imaging device 11, and determines whether volumetric imaging is possible based on whether the estimated 3D modeling processing time is within a predetermined time. More specifically, the controller 101 estimates the decoding processing time for 3D model data from the resolution, frame rate, compression encoding method, bit rate, etc. of each imaging device 11, and estimates the rendering processing time for a free viewpoint image based on the number of voxels to be processed calculated from the voxel size and bounding box size. In the case of offline modeling, the controller 101 determines whether the estimated 3D modeling processing time is within the processing time requested by the user, and in the case of real-time modeling, determines whether the estimated 3D modeling processing time is equal to or less than a predetermined value that allows real-time processing.

[0156] [STEP 5] If it is determined that volumetric imaging is possible, the controller 101 presents the candidates for the setting values ​​determined in [STEP 3] to the user in a predetermined order of priority.

[0157] For example, if image quality is prioritized, candidate settings are presented in the following order: (1) setting values ​​that result in smaller voxel size, (2) setting values ​​that result in higher resolution, and (3) setting values ​​that result in higher frame rate.

[0158] For example, if frame rate is prioritized, candidate settings are presented in the following order: (1) setting values ​​that increase frame rate, (2) setting values ​​that decrease voxel size, and (3) setting values ​​that increase resolution.

[0159] As described above, the cloud server 12 determines setting value candidates for each imaging device 11 based on the capability information of each imaging device 11, and automatically determines whether volumetric imaging is possible.

[0160] When a user uses a smartphone, tablet, or other device that they normally use as the imaging device 11 for volumetric imaging, the performance of the electronic devices owned by the user will vary widely. The imaging setting process for volumetric imaging described above makes it possible to optimally determine the setting values ​​for the imaging functions according to the performance of the imaging device 11 participating in the volumetric imaging.

[0161] <8. Flowchart of synchronized shooting process of shooting devices> Next, the synchronized photographing process of the photographing device 11 executed in step S5 of FIG. 5 will be described in detail.

[0162] There are at least two types of synchronized photographing processes of the photographing device 11: a process described with reference to Fig. 14 and a process described with reference to Fig. 15. The synchronized photographing process of Fig. 14 is a process for synchronously photographing using clock information in the photographing device 11, and the synchronized photographing process of Fig. 15 is a process for the case where the clock information in the photographing device 11 cannot be used. The synchronized photographing process for the case where the clock information in the photographing device 11 cannot be used is a photographing process for enabling the photographed image data to be synchronized afterwards.

[0163] First, a description will be given of a synchronized photographing process when photographing is performed in synchronization using clock information in the photographing device 11, with reference to the flowchart in Fig. 14. When the process in Fig. 14 starts, the clock information in the photographing device 11 is synchronized with high precision by the first to fourth synchronization methods described above.

[0164] In step S81, the cloud server 12 acquires the current time from each of the grouped imaging devices 11.

[0165] In step S82, the cloud server 12 determines the time obtained by adding a certain time to the most advanced time among the acquired times of the imaging devices 11 as the capture start time.

[0166] In step S83, the cloud server 12 controls the camera 51 of each photographing device 11 to be in a standby state by transmitting a predetermined command to each photographing device 11 via the network. The standby state is a state in which, when a synchronization signal is input, photographing can be performed in response to the synchronization signal.

[0167] In step S84, the cloud server 12 transmits the capture start time determined in step S82 to each image capturing device 11.

[0168] In step S85, each imaging device 11 acquires the capture start time transmitted from the cloud server 12 and sets it in the synchronization signal generation unit 45. When the capture start time arrives, the synchronization signal generation unit 45 starts generating and outputting a synchronization signal. The synchronization signal is output at the frame rate set in the imaging setting process of FIG. 8 and supplied to the camera 51.

[0169] In step S86, the camera 51 of each imaging device 11 captures an image of the subject to be used as a 3D model based on the input synchronization signal. Image data (image stream) of the subject obtained by capturing the image is temporarily stored in the imaging device 11 or transmitted to the cloud server 12 in real time.

[0170] In step S87, each photographing device 11 determines whether an instruction to end photographing has been issued. The end of photographing is determined by a user operating a predetermined photographing device 11 such as a master device, and the end instruction may be issued from that photographing device 11 to the other photographing devices 11, or may be issued via the cloud server 12.

[0171] If it is determined in step S87 that an instruction to end shooting has not been issued, steps S86 and S87 are repeated, that is, shooting based on the synchronization signal continues.

[0172] On the other hand, if it is determined in step S87 that an instruction to end the photographing has been issued, the process proceeds to step S88, in which each photographing device 11 stops photographing the subject, thereby terminating the synchronized photographing process. The generation of the synchronization signal by the synchronization signal generator 45 also stops.

[0173] Next, with reference to the flowchart of FIG. 15, a synchronized photographing process when the clock information in the photographing device 11 cannot be used will be described.

[0174] First, in step S101, the cloud server 12 transmits a predetermined command to each of the photographing devices 11 via the network, causing each of the photographing devices 11 to start photographing.

[0175] In step S102, each imaging device 11 captures an image of a subject to be used as a 3D model based on a synchronization signal generated by the synchronization signal generating unit 45 based on the internal clock. Image data (image stream) of the subject obtained by the capture is temporarily stored in the imaging device 11 or transmitted to the cloud server 12 in real time.

[0176] In step S103, the cloud server 12 determines whether a predetermined time has elapsed since the start of image capture.

[0177] If it is determined in step S103 that a predetermined time has elapsed since the start of shooting, then in step S104, the cloud server 12 sends a command to fire the flash 39 to a predetermined shooting device 11, for example, the master device.

[0178] In step S105, the predetermined photographing device 11 to which the flash emission command has been transmitted emits the flash 39 based on the emission command.

[0179] On the other hand, if it is determined in step S103 that the predetermined time after the fixed time has elapsed since the start of shooting has not arrived, steps S104 and S105 are skipped. Therefore, the processing of steps S104 and S105 is executed only at the predetermined time after the fixed time has elapsed since the start of shooting.

[0180] In step S106, each photographing device 11 determines whether an instruction to end photographing has been issued. If it is determined in step S106 that an instruction to end photographing has not been issued, the above-described steps S102 to S106 are repeated. That is, photographing based on the synchronization signal continues.

[0181] On the other hand, if it is determined in step S106 that an instruction to end the photographing has been issued, the process proceeds to step S107, in which each photographing device 11 stops photographing the subject, thereby terminating the synchronized photographing process. The synchronization signal generation unit 45 also stops generating synchronization signals.

[0182] 15, flash images are included in the image stream captured by the grouped multiple imaging devices 11. When performing 3D modeling processing, the images captured by the multiple imaging devices 11 can be synchronized by aligning the time information of the image streams captured by each of the multiple imaging devices 11 so that frames captured by the flash have the same timestamp.

[0183] In addition to using the flash 39 during shooting to align the timestamps, the captured images may also be synchronized by outputting sound from the speaker 35 and aligning the time information of the image stream so that the frames in which the sound is recorded have the same timestamp.

[0184] <9. Flowchart of offline modeling process> Next, the offline modeling process executed as step S7 in FIG. 5 will be described in detail with reference to the flowchart in FIG.

[0185] First, in step S121, each of the plurality of photographing devices 11 determines whether or not photographing has ended, and waits until it is determined that photographing has ended.

[0186] If it is determined in step S121 that the photographing has ended, the process proceeds to step S122, and each photographing device 11 transmits an image stream of photographed images of the subject via a predetermined network to the cloud server 12. The photographed images of the subject undergo predetermined image signal processing such as demosaicing in the image processing unit 54, and then undergo predetermined compression encoding processing in the image compression unit 55, and are then transmitted as an image stream from the stream transmission unit 59.

[0187] In step S123, the stream receiving unit 104 of the cloud server 12 receives the image streams transmitted from the image capturing devices 11 via the communication unit 103, and supplies them to the modeling task generating unit .

[0188] In step S124, the modeling task generation unit 106 of the cloud server 12 acquires the camera parameters of each imaging device 11 from the calibration unit 105, generates a modeling task together with the image stream from the stream reception unit 104, and stores the generated modeling task in the task storage unit 107. The task storage unit 107 stores the modeling tasks sequentially supplied from the modeling task generation unit 106 as a task queue.

[0189] In step S125, the offline modeling unit 108 acquires one of the modeling tasks stored in the task storage unit 107.

[0190] In step S126, the offline modeling unit 108 acquires the captured image of the i-th frame from the image stream of each image capture device 11 of the acquired modeling task. Here, the initial value of the variable i, which is the frame number of the image stream, is set to "1."

[0191] In step S127, the offline modeling unit 108 generates a silhouette image from the captured image of the i-th frame acquired from each imaging device 11 and a preset background image. The silhouette image is an image that represents the subject area as a silhouette, and can be generated, for example, by using a background subtraction method that calculates the difference between the captured image and the background image.

[0192] In step S128, the offline modeling unit 108 generates shape data representing the 3D shape of the object, for example, by Visual Hull, from the N silhouette images corresponding to each imaging device 11. Visual Hull is a technique for projecting the N silhouette images according to camera parameters and carving out a three-dimensional shape. The shape data representing the 3D shape of the object is represented, for example, by voxel data that indicates whether an image belongs to the object in three-dimensional lattice (voxel) units.

[0193] In step S129, the offline modeling unit 108 converts the shape data representing the 3D shape of the object from voxel data into mesh-format data called polygon mesh. To convert the polygon mesh data format, which is easy to render on a display device, an algorithm such as the marching cubes method can be used.

[0194] In step S130, the offline modeling unit 108 generates a texture image corresponding to the shape data of the object. When the multi-texture format described with reference to Fig. 3 is adopted, the captured images captured by each capturing device 11 are used as the texture image. On the other hand, when the UV mapping format described with reference to Fig. 3 is adopted, a UV mapping image corresponding to the shape data of the object is generated as the texture image.

[0195] In step S131, the offline modeling unit 108 determines whether the current frame is the final frame of the image stream of each image capture device 11 of the acquired modeling task.

[0196] If it is determined in step S131 that the current frame is not the final frame of the image stream of each image capture device 11 for the acquired modeling task, the process proceeds to step S132. Then, in step S132, the variable i is incremented by 1, and then the above-described steps S126 to S131 are repeated. That is, shape data and a texture image of the object are generated for the captured image of the next frame in the image stream of each image capture device 11 for the acquired modeling task.

[0197] On the other hand, if it is determined in step S131 that the current frame is the final frame of the image stream of each imaging device 11 for the acquired modeling task, the process proceeds to step S133, where the offline modeling unit 108 supplies the 3D model data of the object to the content management unit 109. The content management unit 109 stores and manages the 3D model data of the object supplied from the offline modeling unit 108 in the content storage unit 110 as content.

[0198] 17 shows an example of 3D model data of an object stored in the content storage unit 110. The 3D model data of an object has shape data and texture images of the object for each frame.

[0199] This completes the offline modeling process.

[0200] <10. Content Playback Processing Flowchart> Next, the content playback process executed as step S8 in FIG. 5 will be described in detail with reference to the flowchart in FIG.

[0201] A viewer, who is a user of the playback device 13, performs an operation to specify content to be played back. In step S141, the message transmitting / receiving unit 152 of the playback device 13 requests the cloud server 12 for 3D model data of the content specified for playback by the user.

[0202] In step S142, the cloud server 12 transmits the 3D model data of the specified content to the playback device 13. More specifically, the controller 101 of the cloud server 12 receives a message requesting 3D model data of predetermined content via the message transmitting / receiving unit 102. The controller 101 instructs the content management unit 109 to transmit the requested content to the playback device 13. The content management unit 109 acquires the 3D model data of the specified content from the content storage unit 110, and transmits it to the playback device 13 via the stream transmitting unit 112.

[0203] In step S143, the playback device 13 acquires the shape data and texture image of the k-th frame of the content transmitted from the cloud server 12. Here, the initial value of the variable k, which is the frame number of the 3D model data of the transmitted content, is set to "1."

[0204] In step S144, the control unit 153 of the playback device 13 determines a virtual viewpoint position, which is the viewer's viewing position relative to the object, based on the sensor information detected by the sensor 154. Information on the determined virtual viewpoint position is supplied to the playback unit 157.

[0205] In step S145, the playback unit 157 of the playback device 13 performs rendering processing to draw an image of the 3D model object based on the determined virtual viewpoint position. That is, the playback unit 157 generates an object image of the 3D model object viewed from the virtual viewpoint position, and outputs the object image to the display 159 via the image output unit 158.

[0206] In step S146, the playback unit 157 determines whether the current frame is the final frame of the acquired content.

[0207] If it is determined in step S146 that the current frame is not the final frame of the acquired content, the process proceeds to step S147. Then, in step S147, the variable k is incremented by 1, and then the above-described steps S143 to S146 are repeated. That is, an object image viewed from the virtual viewpoint is generated and displayed based on the shape data and texture image of the object in the next frame of the acquired content.

[0208] On the other hand, if it is determined in step S146 that the current frame is the final frame of the acquired content, the content playback process ends.

[0209] <11. Real-time modeling playback process flowchart> Next, with reference to the flowchart of FIG. 19, the real-time modeling reproduction process executed as step S9 in FIG. 5 will be described in detail.

[0210] The state in which the real-time modeling playback process is executed is when it is determined in step S6 of Figure 5 that real-time modeling is to be performed, and therefore, images captured by volumetric photography in the image stream are sequentially transmitted from each of the multiple photography devices 11 to the cloud server 12.

[0211] In step S161, the stream receiving unit 104 of the cloud server 12 receives, via the communication unit 103, the captured images of the image streams sequentially transmitted from the image capturing devices 11, and supplies them to the real-time modeling unit 111.

[0212] In step S162, the real-time modeling unit 111 acquires one frame at a time of the captured images of each imaging device 11 that have been supplied and accumulated from the stream receiving unit 104. More specifically, the real-time modeling unit 111 acquires, for each of the multiple imaging devices 11, one frame with the oldest time information from the accumulated captured images of each imaging device 11.

[0213] In step S163, the real-time modeling unit 111 generates a silhouette image from the acquired images captured by each imaging device 11 and a preset background image.

[0214] In step S164, the real-time modeling unit 111 generates shape data representing the 3D shape of the object, for example, by visual hull, from the N silhouette images corresponding to each imaging device 11. The shape data representing the 3D shape of the object is represented, for example, by voxel data.

[0215] In step S165, the real-time modeling unit 111 converts the shape data representing the 3D shape of the object from voxel data into mesh-format data called a polygon mesh.

[0216] In step S166, the real-time modeling unit 111 generates a texture image corresponding to the shape data of the object.

[0217] In step S167, the real-time modeling unit 111 transmits the shape data and texture image of the object to the playback device 13 via the stream transmission unit 112.

[0218] In step S168 , the playback unit 157 of the playback device 13 acquires the shape data and texture image of the object transmitted from the cloud server 12 via the stream receiving unit 156 .

[0219] In step S169, the control unit 153 of the playback device 13 determines a virtual viewpoint position, which is the viewer's viewing position relative to the object, based on the sensor information detected by the sensor 154. Information on the determined virtual viewpoint position is supplied to the playback unit 157.

[0220] In step S170, the playback unit 157 of the playback device 13 performs rendering processing to draw an image of the 3D model object based on the determined virtual viewpoint position. That is, the playback unit 157 generates an object image of the 3D model object viewed from the virtual viewpoint position, and outputs the object image to the display 159 via the image output unit 158.

[0221] In step S171, the control unit 153 of the playback device 13 determines whether or not reception of the shape data and texture image of the 3D model object has been completed.

[0222] If it is determined in step S171 that the reception of the shape data and texture image of the 3D model object has not yet been completed, the process proceeds to step S162, and the processes of steps S162 to S171 described above are repeated.

[0223] On the other hand, if it is determined in step S171 that the reception of the shape data and texture image of the 3D model object has been completed, the real-time modeling reproduction process ends.

[0224] Note that, although the above-mentioned real-time modeling reproduction process has been described as the processes of the cloud server 12 and the reproduction device 13 being executed consecutively, it goes without saying that in reality, the processes are executed independently in each of the cloud server 12 and the reproduction device 13. In the real-time modeling reproduction process, the cloud server 12 sequentially generates and transmits 3D model data in frame units, which is made up of the shape data and texture images shown in Fig. 17, and the reproduction device 13 sequentially receives, reproduces (draws), and displays the 3D model data in frame units.

[0225] <12. Auto-calibration process flowchart> When shooting using a smartphone or tablet as the playback device 13, it is not possible to assume that the camera will be firmly fixed as in a conventional dedicated studio, and it is possible that the position will shift due to vibrations such as camera shake.

[0226] Therefore, during the above-mentioned real-time modeling playback process, the cloud server 12 can acquire sensor information such as the gyro sensor 41 and acceleration sensor 42 of the photographing device 11, determine the positional deviation of the photographing device 11, and perform an auto-calibration process to update the camera parameters.

[0227] The auto-calibration process will be described with reference to the flowchart of Fig. 20. This process is executed simultaneously with the real-time modeling reproduction process of Fig. 19, for example.

[0228] First, in step S181, the imaging device 11 acquires sensor information from various sensors such as the gyro sensor 41, the acceleration sensor 42, and the GPS sensor 43, and transmits the information together with the captured image as frame data to the cloud server 12. The sensor information is stored in, for example, header information of each frame, and transmitted on a frame-by-frame basis.

[0229] In step S182, the stream receiving unit 104 of the cloud server 12 receives, via the communication unit 103, the captured images of the image stream sequentially transmitted from each imaging device 11, and supplies them to the real-time modeling unit 111. The real-time modeling unit 111 supplies, to the auto-calibration unit 113, the sensor information stored in the frame data of the captured images.

[0230] In step S183, the auto-calibration unit 113 determines, based on the sensor information from the real-time modeling unit 111, whether any of the imaging devices 11 performing volumetric imaging has moved by a predetermined value or more.

[0231] If it is determined in step S183 that there is a photographing device 11 that has moved by more than a predetermined value, the processing of the following steps S184 to S186 is executed, and if it is determined that there is no photographing device 11 that has moved by more than a predetermined value, the processing of steps S184 to S186 is skipped.

[0232] In step S184, the auto-calibration unit 113 extracts feature points from the images captured by each imaging device 11.

[0233] In step S185, the auto-calibration unit 113 calculates the camera parameters of the imaging device 11 to be updated, using feature points extracted from the captured image of each imaging device 11. The imaging device 11 to be updated is the imaging device 11 determined to have moved by a predetermined value or more in step S183. The camera parameters calculated here are the external parameters, out of the internal parameters and external parameters, and the internal parameters are left unchanged.

[0234] In step S186, the auto-calibration unit 113 supplies the calculated camera parameters of the imaging device 11 to be updated to the real-time modeling unit 111. The real-time modeling unit 111 uses the updated camera parameters of the imaging device 11 to generate 3D model data.

[0235] In step S187, the cloud server 12 determines whether the image capturing has ended. For example, if the captured images are still being transmitted from each of the image capturing devices 11, the cloud server 12 determines that the image capturing has not ended, and if the captured images are no longer being transmitted from each of the image capturing devices 11, the cloud server 12 determines that the image capturing has ended.

[0236] If it is determined in step S187 that the photographing has not yet ended, the process returns to step S181, and the processes of steps S181 to S187 described above are repeated.

[0237] On the other hand, if it is determined in step S187 that the photographing has ended, the auto-calibration process ends.

[0238] In the above-described auto-calibration process, the camera parameters are calculated and updated only for the photographing device 11 that is determined to have moved by a predetermined value or more, but when updating the camera parameters, it is also possible to calculate the camera parameters for all the photographing devices 11. Furthermore, not only the external parameters but also the internal parameters may be calculated.

[0239] Furthermore, when calculating the camera parameters of the imaging device 11 to be updated, the external parameters may be calculated by limiting the range of possible values ​​of the external parameters based on the amount of deviation calculated from the sensor information. The calculation of the external parameters is a nonlinear optimization calculation of the three-dimensional positions of the feature points and the internal and external parameters of the imaging device 11, so by limiting the range of possible values ​​of the external parameters or fixing the internal parameters, the amount of calculation can be reduced and processing can be speeded up.

[0240] 13. Camera calibration without using calibration board images In the camera calibration process described with reference to FIGS. 8 to 11, instead of using a calibration board, the camera parameters are calculated by displaying a calibration board image on the photographing device 11 and extracting feature points.

[0241] However, when the display size of a smartphone or tablet is taken into consideration, it may not be possible to sufficiently detect feature points.

[0242] Therefore, a camera calibration process other than the method of displaying the calibration board image on the photographing device 11 (hereinafter referred to as second camera calibration process) will be described.

[0243] In the second camera calibration process described below, the image processing system 1 calculates camera parameters by capturing an image of an arbitrary subject and extracting its feature points. For example, as shown in Fig. 21, the camera parameters are calculated based on the result of capturing an image of a person and a table as subjects.

[0244] In the second camera calibration process, a control device 15 is provided that controls the entire camera calibration process, in addition to the multiple (five) imaging devices 11-1 to 11-5 that capture images of the subject. The control device 15 may be any device having a configuration equivalent to that of the playback device 13, and the playback device 13 may also be used.

[0245] The second camera calibration process will be described with reference to Fig. 22 to Fig. 25. Fig. 22 is a flowchart of the camera calibration process performed by the control device 15 in the second camera calibration process, and Fig. 24 and Fig. 25 are flowcharts of the camera calibration process performed by the cloud server 12.

[0246] First, the camera calibration process performed by the control device 15 will be described with reference to the flowchart of FIG.

[0247] In step S201, the control device 15 transmits a calibration image capture start message to the cloud server 12 to start the calibration image capture.

[0248] After transmitting the calibration shooting start message, the control device 15 waits until a predetermined message is transmitted from the cloud server 12. While the camera parameters of each photographing device 11 are being calculated, feedback information for the camera parameter calculation is transmitted from the cloud server 12 to the control device 15. Then, when the calculation of the camera parameters of each photographing device 11 is completed, a calibration completion message is transmitted from the cloud server 12 to the control device 15.

[0249] In step S202, the control device 15 receives the message sent from the cloud server 12.

[0250] In step S203, the control device 15 determines whether the calibration is complete. If the received message is a calibration completion message, the control device 15 determines that the calibration is complete, and if the received message is feedback information, the control device 15 determines that the calibration is not yet complete.

[0251] If it is determined in step S203 that the calibration has not yet been completed, the process proceeds to step S204, where the control device 15 displays a feedback screen on the display based on the feedback information.

[0252] The calibration unit 105 of the cloud server 12 extracts feature points for each captured image captured by each imaging device 11 and associates the extracted feature points with multiple captured images corresponding to each imaging device 11. That is, the calibration unit 105 associates each extracted feature point with the imaging devices 11. Then, the calibration unit 105 calculates camera parameters by solving a nonlinear optimization problem that minimizes the error for the three-dimensional position, internal parameters, and external parameters of each feature point using the multiple feature points associated with each other between the imaging devices 11. If there are few feature points extracted from the captured image or if there are few associated feature points between the imaging devices 11, the cloud server 12 transmits feedback information to the control device 15 and provides feedback to the user to redo the imaging using the imaging device 11.

[0253] FIG. 23 shows an example of a feedback screen displayed on the control device 15 based on feedback information from the cloud server 12.

[0254] The feedback screen in FIG. 23 shows an example in which the number of feature points associated with each other between the image capturing devices 11-2 and 11-4 is small.

[0255] 23 displays a photographed image 201 photographed by the photographing device 11-2 and a photographed image 202 photographed by the photographing device 11-4. In each of the photographed images 201 and 202, a circle (◯) is superimposed on the feature points that have been matched, and a cross (×) is superimposed on the feature points that could not be matched.

[0256] The feedback screen also displays a feedback message 203 to improve the accuracy of calibration when redoing photography. In the example of Fig. 23, the feedback message 203 displays "Adjust the subject so that a part common to both camera 2 and camera 4 is captured."

[0257] Furthermore, the feedback screen displays an object image 204 that would appear if a 3D model of the subject were generated using the current calibration results. The virtual viewpoint position of the object image 204 can be changed by swiping on the screen.

[0258] Returning to FIG. 22, if it is determined in step S203 that the calibration is complete, the process proceeds to step S205, and the control device 15 sends a calibration shooting completion message to the cloud server 12 to end the calibration shooting, thereby ending the camera calibration process.

[0259] Next, the camera calibration process performed by the cloud server 12 will be described with reference to the flowchart of FIG.

[0260] First, in step S221, the cloud server 12 receives a calibration image capture start message transmitted from the control device 15.

[0261] In step S222, the cloud server 12 starts synchronous photographing for each of the photographing devices 11. The synchronous photographing process for each of the photographing devices 11 can be performed in the same manner as in FIG.

[0262] In step S223, the cloud server 12 acquires the captured images transmitted from each of the image capturing devices 11.

[0263] In step S224, the cloud server 12 executes a camera parameter calculation process to calculate camera parameters using the captured images transmitted from each of the image capturing devices 11.

[0264] FIG. 25 is a flowchart showing the details of the camera parameter calculation process in step S224.

[0265] In this process, first, in step S241, the calibration unit 105 of the cloud server 12 extracts feature points from each of the images captured by each imaging device 11.

[0266] Next, in step S242, the calibration unit 105 determines whether the extracted feature points are sufficient. For example, if the number of extracted feature points in one captured image is equal to or greater than a predetermined value, the extracted feature points are determined to be sufficient, and if the number is less than the predetermined value, the extracted feature points are determined to be insufficient.

[0267] If it is determined in step S242 that the extracted feature points are not sufficient, the process proceeds to step S247, which will be described later.

[0268] On the other hand, if it is determined in step S242 that the extracted feature points are sufficient, the processing proceeds to step S243, and the calibration unit 105 matches the feature points between the imaging devices 11 and detects corresponding feature points between the imaging devices 11.

[0269] Next, in step S244, the calibration unit 105 determines whether there are sufficient corresponding feature points. For example, if the number of corresponding feature points in one captured image is equal to or greater than a predetermined value, it is determined that there are sufficient corresponding feature points, and if the number is less than the predetermined value, it is determined that there are insufficient corresponding feature points.

[0270] If it is determined in step S244 that there are not enough corresponding feature points, the process proceeds to step S247, which will be described later.

[0271] On the other hand, if it is determined in step S244 that the corresponding feature points are sufficient, the process proceeds to step S245, where the calibration unit 105 calculates the three-dimensional positions of the corresponding feature points and also calculates the camera parameters, i.e., the internal parameters and external parameters, of each imaging device 11. The three-dimensional positions of each feature point and the camera parameters of each imaging device 11 are calculated by solving a nonlinear optimization problem that minimizes the error.

[0272] In step S246, the calibration unit 105 determines whether the error in the nonlinear optimization calculation when the camera parameters are calculated is sufficiently small, equal to or less than a predetermined value.

[0273] If it is determined in step S246 that the error in the nonlinear optimization calculation is not equal to or less than the predetermined value, the process proceeds to step S247, which will be described later.

[0274] In step S247, the calibration unit 105 generates feedback information including a captured image with feature points and a feedback message.

[0275] If it is determined in step S242 above that the extracted feature points are insufficient, the detected feature points are superimposed as circles (◯) on the captured images 201 and 202 on the feedback screen of Fig. 23 as feedback information generated in the processing of step S247. On the other hand, if it is determined in step S244 above that the corresponding feature points are insufficient, the corresponding feature points are superimposed as circles (◯) on the captured images 201 and 202 on the feedback screen of Fig. 23 as feedback information generated in the processing of step S247, and the feature points that could not be matched are superimposed as crosses (×). On the other hand, if it is determined in step S246 above that the error is large, the image captured by the imaging device 11 with the large error is displayed on the feedback screen of Fig. 23 as feedback information generated in the processing of step S247.

[0276] On the other hand, if it is determined in step S246 that the error in the nonlinear optimization calculation is sufficiently small, that is, equal to or less than the predetermined value, the process proceeds to step S248, where the calibration unit 105 generates a message indicating that calibration has been completed.

[0277] By the process of step S247 or S248, the camera parameter calculation process ends, and the process returns to FIG. 24 and proceeds to step S225.

[0278] In step S225 of FIG. 24, the calibration unit 105 determines whether the calibration is complete, that is, whether a message indicating the completion of the calibration is generated.

[0279] If it is determined in step S225 that the calibration is not completed, that is, that the feedback information has been generated, the process proceeds to step S226, where the calibration unit 105 transmits the generated feedback information as a message. After step S226, the process returns to step S224, and the processes from step S224 onwards are repeated.

[0280] On the other hand, if it is determined in step S225 that the calibration is completed, that is, that a calibration completion message has been generated, the process proceeds to step S227, where the cloud server 12 transmits a message indicating that image capture has ended to each photographing device 11. Each photographing device 11 receives the message indicating that image capture has ended and ends synchronous photographing.

[0281] In step S228, the cloud server 12 transmits the generated calibration completion message to the control device 15, and ends the camera calibration process.

[0282] According to the second camera calibration process described above, it is possible to calibrate the camera even when there is no calibration board or calibration board image.

[0283] <14. Computer configuration example> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that make up the software are installed on a computer. Here, the term "computer" includes microcomputers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.

[0284] FIG. 26 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0285] In the computer, a CPU (Central Processing Unit) 301, a ROM (Read Only Memory) 302, and a RAM (Random Access Memory) 303 are interconnected by a bus 304.

[0286] An input / output interface 305 is further connected to the bus 304. To the input / output interface 305, an input unit 306, an output unit 307, a storage unit 308, a communication unit 309, and a drive 310 are connected.

[0287] The input unit 306 includes a keyboard, mouse, microphone, touch panel, input terminal, etc. The output unit 307 includes a display, speaker, output terminal, etc. The storage unit 308 includes a hard disk, RAM disk, non-volatile memory, etc. The communication unit 309 includes a network interface, etc. The drive 310 drives a removable recording medium 311 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0288] In the computer configured as above, the CPU 301 performs the above-described series of processes by, for example, loading a program stored in the storage unit 308 into the RAM 303 via the input / output interface 305 and the bus 304 and executing the program. The RAM 303 also stores data necessary for the CPU 301 to execute various processes as needed.

[0289] The program executed by the computer (CPU 301) can be provided by being recorded on a removable recording medium 311 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0290] In this specification, the steps described in the flowcharts may be performed in chronological order in the order described, but they do not necessarily have to be processed in chronological order, and may be performed in parallel or at any necessary timing, such as when a call is made.

[0291] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device with multiple modules housed in a single housing, are both systems.

[0292] The embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure.

[0293] Furthermore, for example, the multiple technologies of the present disclosure can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies of the present disclosure can also be implemented in combination. For example, part or all of the technology described in any embodiment can be implemented in combination with part or all of the technology described in another embodiment. Furthermore, part or all of any of the above-described technologies of the present disclosure can also be implemented in combination with other technologies not described above.

[0294] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.

[0295] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.

[0296] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0297] The effects described in this specification are merely examples and are not limiting, and there may be effects other than those described in this specification.

[0298] The present technology can have the following configurations. (1) a control unit that transmits its own information related to photography to a server device and acquires candidate setting values ​​related to volumetric photography transmitted from the server device based on a determination result of whether volumetric photography is possible; Filming device. (2) The control unit presents candidates for the setting value acquired from the server device to the user and allows the user to select one. The photographing device according to (1) above. (3) The control unit further controls to transmit identification information of a group consisting of a plurality of imaging devices that perform volumetric imaging and its own location information, and to receive information indicating whether or not it is possible to register the group. The photographing device according to (1) or (2) above. (4) The control unit acquires, as the setting value candidates, camera synchronization methods supported by all image capturing devices that perform volumetric image capturing. The photographing device according to any one of (1) to (3). (5) The camera synchronization method generates a synchronization signal based on time information from carrier communication, time information from a GPS signal, or timing detection in wireless communication or multicast communication from carrier communication. The photographing device according to (4) above. (6) The control unit controls the volumetric photography so that a flash or sound is output at a predetermined time to generate an image including the flash or sound, and synchronizes the images photographed by each photographing device by detecting the image including the flash or sound. The photographing device according to any one of (1) to (5). (7) The control unit controls the camera calibration process so as to capture a predetermined image displayed by another image capture device, detect feature points of the predetermined image, and transmit the detected feature points to the server device. The photographing device according to any one of (1) to (6). (8) The control unit controls a plurality of imaging devices that perform volumetric imaging so as to capture the displayed predetermined images in sequence. The photographing device according to (7) above. (9) The control unit controls a plurality of imaging devices that perform volumetric imaging so as to simultaneously capture the predetermined images that are displayed. The photographing device according to (7) above. (10) The predetermined images displayed simultaneously are different images from the plurality of photographing devices. The photographing device according to (9) above. (11) the control unit transmits sensor information in units of frames together with a captured image of a predetermined subject; The sensor information is used to determine whether to update the camera parameters. The photographing device according to any one of (1) to (10). (12) a control unit that receives information about photography from a plurality of photography devices and determines whether the plurality of photography devices are capable of volumetric photography based on the received information; Server device. (13) The control unit generates setting value candidates for volumetric imaging of the imaging device based on the received information, and determines whether volumetric imaging is possible based on the generated setting value candidates. The server device according to (12) above. (14) The control unit calculates a photographing target area of ​​each photographing device, and when the photographing target areas of the photographing devices are superimposed, determines whether volumetric photographing is possible based on whether or not there is a common area among a predetermined number or more of photographing devices. The server device according to any one of (12) to (13). (15) The control unit calculates modeling parameters of a 3D model from the resolution of each imaging device and the imaging target area, and determines whether volumetric imaging is possible based on whether the calculated modeling parameters are within a predetermined range. The server device according to any one of (12) to (14). (16) The control unit estimates a processing time required for 3D modeling from the information of each imaging device, and determines whether volumetric imaging is possible based on whether the estimated 3D modeling processing time is within a predetermined time. The server device according to any one of (12) to (15). (17) The camera further includes a calibration unit that acquires predetermined images captured by the plurality of imaging devices and calculates camera parameters. The server device according to any one of (12) to (16). (18) The image processing device further includes a modeling unit that generates 3D model data of the subject from the images captured by the plurality of imaging devices and transmits the generated 3D model data to a playback device. The server device according to any one of (12) to (17). (19) The image capturing device further includes a calibration unit that updates camera parameters when it is determined that the image capturing device has moved based on sensor information received on a frame-by-frame basis along with the captured image of the subject. The server device according to (18) above. (20) receiving information about photography from a plurality of photography devices, and determining whether the plurality of photography devices are capable of volumetric photography based on the received information; Generate 3D model data of the subject from images captured by the plurality of imaging devices determined to be capable of volumetric imaging. 3D data generation method. [Explanation of symbols]

[0299] 1 Image processing system, 11 Photography device, 12 Cloud server, 13A Smartphone or tablet, 13B Personal computer, 13 Playback device, 13C Head-mounted display (HMD), 15 Control device, 33 Control unit, 45 Synchronization signal generation unit, 51 Camera, 101 Controller, 105 Calibration unit, 106 Modeling task generation unit, 108 Offline modeling unit, 109 Content management unit, 111 Real-time modeling unit, 113 Auto-calibration unit, 153 Control unit, 157 Playback unit, 301 CPU, 302 ROM, 303 RAM, 306 Input unit, 307 Output unit, 308 Memory unit, 309 Communication unit, 310 Drive

Claims

1. a control unit that transmits its own information related to photography to a server device, and acquires candidate setting values ​​related to volumetric photography that are transmitted from the server device based on a determination result of whether volumetric photography is possible; The control unit further controls to transmit identification information of a group consisting of a plurality of imaging devices that perform volumetric imaging and its own location information, and to receive information indicating whether or not it is possible to register the group. Filming device.

2. The control unit presents candidates for the setting value acquired from the server device to the user and allows the user to select one. The imaging device according to claim 1 .

3. The control unit acquires, as the setting value candidates, camera synchronization methods supported by all image capturing devices that perform volumetric image capturing. The imaging device according to claim 1 .

4. The camera synchronization method generates a synchronization signal based on time information from carrier communication, time information from a GPS signal, or timing detection in wireless communication or multicast communication from carrier communication. The imaging device according to claim 3 .

5. The control unit controls the volumetric photography so that a flash or sound is output at a predetermined time to generate an image including the flash or sound, and synchronizes the images photographed by each photographing device by detecting the image including the flash or sound. The imaging device according to claim 1 .

6. The control unit controls the camera calibration process so as to capture a predetermined image displayed by another image capture device, detect feature points of the predetermined image, and transmit the detected feature points to the server device. The imaging device according to claim 1 .

7. The control unit controls a plurality of imaging devices that perform volumetric imaging so as to capture the displayed predetermined images in sequence. The imaging device according to claim 6 .

8. The control unit controls a plurality of imaging devices that perform volumetric imaging so as to simultaneously capture the predetermined images that are displayed. The imaging device according to claim 6 .

9. The predetermined images displayed simultaneously are different images from the plurality of photographing devices. The imaging device according to claim 8 .

10. the control unit transmits sensor information in units of frames together with a captured image of a predetermined subject; The sensor information is used to determine whether to update the camera parameters. The imaging device according to claim 1 .

11. a control unit that receives information related to photography from a plurality of photography devices and determines whether the plurality of photography devices are capable of volumetric photography based on the received information; The control unit generates setting value candidates for volumetric imaging of the imaging device based on the received information, and determines whether volumetric imaging is possible based on the generated setting value candidates. Server device.

12. A control unit that receives information related to photography from a plurality of photography devices and determines whether the plurality of photography devices are capable of volumetric photography based on the received information, The control unit calculates modeling parameters of a 3D model from the resolution of each imaging device and the imaging target area, and determines whether volumetric imaging is possible based on whether the calculated modeling parameters are within a predetermined range. Server device.

13. A control unit that receives information related to photography from a plurality of photography devices and determines whether the plurality of photography devices are capable of volumetric photography based on the received information, The control unit estimates a processing time required for 3D modeling from the information of each imaging device, and determines whether volumetric imaging is possible based on whether the estimated 3D modeling processing time is within a predetermined time. Server device.

14. The camera further includes a calibration unit that acquires predetermined images captured by the plurality of imaging devices and calculates camera parameters. The server device according to any one of claims 11 to 13.

15. The image processing device further includes a modeling unit that generates 3D model data of the subject from the images captured by the plurality of imaging devices and transmits the generated 3D model data to a playback device. The server device according to any one of claims 11 to 13.

16. The image capturing device further includes a calibration unit that updates camera parameters when it is determined that the image capturing device has moved based on sensor information received on a frame-by-frame basis along with the captured image of the subject. The server device according to claim 15.

17. receiving information related to photography from a plurality of photography devices, generating setting value candidates related to volumetric photography for the photography devices based on the received information, and determining whether the plurality of photography devices are capable of volumetric photography based on the generated setting value candidates; Generate 3D model data of the subject from images captured by the plurality of imaging devices determined to be capable of volumetric imaging. 3D data generation method.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and program

    JP2019087791A

  • Image processing device, image processing method, and information processing device

    WO2014034444A1

  • Video synchronization device

    WO2017134770A1

  • Free-viewpoint image generation method and free-viewpoint image generation system

    WO2018147329A1