Information processing device, information processing method, and program
The information processing apparatus and method improve three-dimensional reconstruction accuracy by selecting and correcting data using metadata and instruction information, addressing camera parameter inconsistencies to enhance reconstruction quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY SEMICON SOLUTIONS CORP
- Filing Date
- 2025-12-25
- Publication Date
- 2026-07-23
AI Technical Summary
Existing three-dimensional reconstruction methods using multiple cameras suffer from inaccuracies due to differences in camera parameters and shooting conditions, leading to decreased reconstruction accuracy.
An information processing apparatus and method that selects and corrects shooting data using metadata and instruction information to generate a reconstruction result, incorporating data selection, correction, and reconstruction processing units to enhance accuracy.
The solution enables the generation of highly accurate 3D or 4D reconstruction results by addressing camera parameter discrepancies, improving data selection and correction processes.
Smart Images

Figure JP2025045550_23072026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Information Processing Method, and Program
[0001] The present technology relates to an information processing apparatus, an information processing method, and a program.
[0002] As a conventional technology, a technique has been proposed for performing three-dimensional reconstruction using a plurality of videos captured by a plurality of cameras and generating a multi-view interactive perspective video (Patent Document 1).
[0003] Japanese Unexamined Patent Application Publication No. 2006-12161
[0004] However, when using a plurality of videos captured by a plurality of cameras, if there are differences in camera parameters, shooting conditions, etc., differences also occur in the videos, resulting in a problem that the accuracy of the reconstruction result generated by three-dimensional reconstruction decreases.
[0005] The present technology has been made in view of such problems, and an object thereof is to provide an information processing apparatus, an information processing method, and a program capable of realizing generation of a highly accurate reconstruction result using a plurality of shooting data captured by a plurality of cameras.
[0006] In order to solve the above-described problems, a first technology is based on either or both of metadata corresponding to a plurality of shooting data generated by shooting with a plurality of cameras and instruction information indicating an instruction regarding a reconstruction result, and selects shooting data to be used for generating a reconstruction result from the plurality of shooting data. A data selection unit, and a reconstruction processing unit that generates a reconstruction result by reconstruction processing based on the shooting data selected by the data selection unit. An information processing apparatus comprising.
[0007] Further, a second technology is based on either or both of metadata corresponding to a plurality of shooting data generated by shooting with a plurality of cameras and instruction information indicating an instruction regarding a reconstruction result, and selects shooting data to be used for generating a reconstruction result from the plurality of shooting data. An information processing method that generates a reconstruction result by reconstruction processing based on the selected shooting data.
[0008] Furthermore, the third technology is a program that causes a computer to execute an information processing method that selects shooting data to be used to generate a reconstruction result from multiple shooting data based on either or both metadata corresponding to multiple shooting data generated by shooting with multiple cameras and instruction information indicating instructions regarding the reconstruction result, and generates a reconstruction result by performing a reconstruction process based on the selected shooting data.
[0009] This is an explanatory diagram of 4D data in this technology. This is a block diagram showing the configuration of the information processing system 10. This is a diagram showing an example of the structure of image data constituting video data. This is a diagram showing the configuration of the processing block of the information processing device 100 in the first embodiment. This is a diagram showing the hardware configuration of the information processing device 100. This is a flowchart showing the processing of the information processing device 100 in the first embodiment. This is an explanatory diagram of the storage of instruction information. This is an explanatory diagram of the storage of instruction information. This is an explanatory diagram of the determination of video data where the shooting viewpoint does not match. This is an explanatory diagram of the first specific example of the correction process. This is an explanatory diagram of the second specific example of the correction process. This is an explanatory diagram of the third specific example of the correction process. This is a diagram showing the configuration of the processing block of the information processing device 100 in the second embodiment. This is a flowchart showing the processing of the information processing device 100 in the second embodiment. This is a diagram showing the configuration of the processing block of the information processing device 100 in the third embodiment. This is a flowchart showing the processing of the information processing device 100 in the third embodiment. This is a diagram showing the configuration of the processing block of the information processing device 100 in the fourth embodiment. This is a flowchart showing the processing of the information processing device 100 in the fourth embodiment. This is a diagram showing the configuration of a modified version of the information processing device 100. This is a diagram showing the configuration of a modified version of the information processing device 100. This is a diagram showing the configuration of a modified version of the information processing device 100. This diagram shows a modified configuration of the information processing device 100. This diagram shows a modified configuration of the information processing system 10.
[0010] The embodiments of this technology will be described below with reference to the drawings. The description will be given in the following order: <First Embodiment> [About 4D Data] [Configuration of Information Processing System 10] [Configuration of Information Processing Device 100] [Processing in Information Processing Device 100] <Second Embodiment> [Configuration of Information Processing Device 100] [Processing in Information Processing Device 100] <Third Embodiment> [Configuration of Information Processing Device 100] [Processing in Information Processing Device 100] <Fourth Embodiment> [Configuration of Information Processing Device 100] [Processing in Information Processing Device 100] <Modifications>
[0011] <First Embodiment> [About 4D Data] The 4D data in this technology will be explained with reference to Figure 1. The information processing device 100, which constitutes the information processing system 10 in this technology, generates 3D data or 4D data as a reconstruction result from multiple shooting data captured by multiple cameras 200 at different positions and orientations through a reconstruction process. Generating 3D data from two-dimensional image data or video data, which are the shooting data, is sometimes called 3D reconstruction, and generating 4D data by adding a time axis to the 3D data is sometimes called 4D reconstruction. In this embodiment, the shooting data is assumed to be video data.
[0012] In this technology, 4D data refers to data that includes information on a 3D spatial axis that moves along the time axis (including color, material, etc.). Specifically, this includes animated 3D mesh data, 3D point cloud data, polygon data, NeRF (Neural Radiance Fields) for dynamic scenes, and 3D Gaussian Splatting, etc.
[0013] The generated 4D data can be distributed to multiple devices, and the distributed 4D data is rendered on each device, ultimately becoming a stereo video with both a time axis and a 2D axis.
[0014] It is possible to generate stereo video by pre-rendering 4D data and then deliver that stereo video to the viewer's device. In this case, viewing from a fixed viewpoint or, in the case of formats such as VR180, viewing in 3DoF (Degree of Freedom) is possible.
[0015] On the other hand, by distributing the content as 4D data rather than stereo video, it becomes possible to perform real-time rendering on the viewer's device. In this case, viewing becomes possible with 6DoF (six degrees of freedom), adding "movement in forward / backward, up / down, and left / right" to "turning the head left / right and forward / backward, and looking around." This allows viewers to view the 4D data from any viewpoint. At that time, the device estimates its own pose using SLAM (Simultaneous Localization and Mapping).
[0016] Viewing devices include HMDs (Head-Mounted Displays), glasses-type wearable devices, AR (Augmented Reality) devices such as smartphones, and spatial reproduction displays.
[0017] In the above explanation, the method of providing 4D data and stereo video was described as "distribution," but other methods of provision are also acceptable, such as saving the 4D data and stereo video to a storage medium and distributing that medium to viewers, or displaying the 4D data and stereo video in a screening.
[0018] Regarding standards for 4D data distribution, there is V-DMC (Video-based Dynamic Mesh Coding), an international standard for the compression and transmission of 3D meshes that change over time. There is also V-PCC (Video-based Point Cloud Compression), an international standard for the compression and transmission of 3D point clouds that change over time.
[0019] This technology can also be applied to volumetric capture. Volumetric capture is a technology that captures an entire space using more than 100 cameras surrounding it in a 360-degree circle, capturing the real space as three-dimensional digital data and reproducing it with high quality. The generated data can be converted into a 2D video viewed from any direction as a free-viewpoint representation, or into a 3D video that can be viewed on AR, stereoscopic monitors, HMDs, etc.
[0020] Furthermore, 4D data can also be displayed in a browser using technologies such as WebGL (Graphics Library) and WebXR, which enable 3D rendering and XR (Extended Reality) applications on a web browser.
[0021] [Configuration of Information Processing System 10] The configuration of the information processing system 10 will be described with reference to Figure 2. The information processing system 10 consists of an information processing device 100, multiple cameras 200, a terminal device 300, and a database 400.
[0022] The information processing device 100 generates 3D or 4D data as a reconstruction result by performing a reconstruction process on video data captured by multiple cameras 200 as input. The information processing device 100 is used by the person who wants to generate the reconstruction result. For example, if the subject captured by the cameras 200, i.e., the subject of the reconstruction process, is an event, the organizer of that event will use the information processing device 100 to generate the reconstruction result. The cameras 200 may capture images with their position and orientation fixed, or they may capture images while moving, i.e., while changing their position and orientation.
[0023] An event can be anything from large-scale events that attract many people, such as weddings, funerals, ceremonies, performances, recitals, concerts, festivals, and sporting events, to small-scale events like family and friends' gatherings. Furthermore, the reconstruction process is not limited to events; it can also be applied to animals, nature, landscapes, buildings, or anything else that can be photographed with Camera 200.
[0024] In this specification, the person who generates the reconstruction results using the information processing device 100, such as the event organizer described above, is referred to as the user. The user may provide the reconstruction results, and the person who receives the reconstruction results from the user and views them is referred to as the viewer. Furthermore, the person who films the event with the camera 200 is referred to as the photographer.
[0025] Assume that camera 200 includes cameras that have been calibrated and whose position, orientation, and camera parameters are known, and cameras that have not been calibrated and whose position, orientation, and camera parameters are unknown. When it is necessary to distinguish between multiple cameras 200, they will be referred to as camera 200A, camera 200B, camera 200C, camera 200D, and so on.
[0026] For example, a calibrated camera 200 is provided by the event organizer (user) of the information processing system 10 and used by the event staff. A non-calibrated camera 200 is used by the event participants. Multiple cameras 200 may differ in manufacturer, model, or type, and their sensor characteristics, f-number, white balance, exposure, and camera parameters may vary due to differences in how they are used by the photographers.
[0027] The information processing device 100, multiple cameras 200, terminal device 300, and database 400 are equipped with communication functions and are connected by a network. The network connection method may be wired or wireless. Wired connection methods include, for example, HDMI (High-Definition Multimedia Interface) and USB (Universal Serial Bus), while wireless connection methods include, for example, Wi-Fi, Bluetooth (Registered Trademark), Wireless LAN (Local Area Network), NFC (Near Field Communication), 4G (Fourth Generation Mobile Communication System), 5G (Fifth Generation Mobile Communication System), and Ethernet (Registered Trademark).
[0028] Camera 200 is equipped with an image sensor, signal processing circuitry, etc., and can obtain color or monochrome video data and image data through shooting. CCD (Charge Coupled Device), CMOS (Complementary Metal Oxide Semiconductor), etc. can be used as the image sensor. Camera 200 can be any device with camera functionality, such as a smartphone, tablet terminal, or wearable device. Due to differences in camera parameters, shooting settings, shooting position, and shooting orientation among multiple cameras 200, differences exist between the multiple video data.
[0029] When the camera 200 generates video data through shooting, it acquires and generates various information related to the shooting and adds this information to the video data as metadata corresponding to the video data. The camera 200 may also acquire and generate metadata after shooting and add it to the video data. The camera 200 then transmits the video data and metadata to the information processing device 100. The transmission of the video data and metadata may be performed automatically when the shooting mode of the camera 200 ends, or in response to a transmission instruction input from the photographer. The supply of video data and metadata from the camera 200 to the information processing device 100 may be performed via a recording medium such as a USB flash memory or SD memory card.
[0030] The reconstruction process requires position and orientation information from the camera 200 to generate the reconstruction results. Therefore, it is desirable that the camera 200 be equipped with a GPS (Global Positioning System) sensor for detecting position information, and an IMU (Inertial Measurement Unit) or inertial sensors (accelerometers, angular velocity sensors, gyroscopes for two or three axes) for detecting orientation information. If the camera 200 is equipped with a GPS sensor and an IMU, the camera 200 can add the position information detected by the GPS sensor and the orientation information detected by the IMU, etc., as metadata to the video data. The camera 200 may also be equipped with an acceleration sensor, angular velocity sensor, gyroscope, etc., as motion sensors, in which case the camera 200 can add metadata and motion detection information to the image data. Furthermore, the camera 200 may be equipped with a LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging), a ToF (Time of Flight) sensor, a stereo camera, a structured light camera, etc., as distance sensors.
[0031] Camera 200 may be equipped with known functions (such as subject detection, face detection, and scene detection) that can detect various types of information from video data obtained during shooting. If camera 200 is equipped with such information detection functions, camera 200 can add the detected information to the video data as metadata. These information detection functions can be implemented using machine learning or deep learning methods, template matching methods, matching methods based on the brightness distribution information of the subject, artificial intelligence methods, instance segmentation, etc.
[0032] Figure 3 shows an example of the structure of image data (frame images) that make up video data. The image data consists of FS, Embedded Data Lines, Image Pixels, and Meta Data. General metadata such as shooting time, exposure time, and white balance are stored in Embedded Data Lines. In addition, metadata obtained through image analysis, such as scene information, is stored in Meta Data.
[0033] In addition to the configuration shown in Figure 3, metadata can also be added by embedding it at predetermined frame intervals (for example, embedding it at fixed frame intervals of N frames), or by embedding it only in frames where a subject or action specified by the user occurs or ends.
[0034] Metadata can include camera parameters, exposure time, white balance, shooting time, shooting date and time (season), shooting location (e.g., detection result of the GPS sensor on camera 200), orientation (e.g., detection result of the IMU on camera 200), motion information of camera 200 (detection results of motion sensors such as accelerometer, angular velocity sensor, and gyroscope), camera 200 manufacturer, camera identification information (e.g., IMEI (International Mobile Equipment Identifier) number of camera 200), camera model information, photographer ID (e.g., login ID during upload), authenticity information (e.g., signature information), depth map information (e.g., PDAF (Phase Detection Auto Focus) data), polarization information, subject detection information, subject name, subject type, face detection information, subject part detection information, scene information, distance to subject information, focal length, illumination at the time of shooting, brightness of image data, distortion parameters, photographer information, etc. Metadata can be any information that can be obtained by camera 200, information that can be generated by camera 200, or information that camera 200 already holds.
[0035] Furthermore, the multiple cameras 200 may differ in their characteristics, camera parameters, signal processing, imaging settings, geometric conditions, etc.
[0036] Note that there may be only one camera 200, rather than multiple cameras. However, since multiple image data are required as input for the reconstruction process, if there is only one camera 200, that camera 200 must take pictures from multiple different positions and orientations to generate multiple image data.
[0037] The terminal device 300 is used by the user to input instructions regarding the reconfiguration process to the information processing device 100. The terminal device 300 transmits the instructions input by the user as instruction information to the information processing device 100. The terminal device 300 can be an electronic device such as a personal computer, smartphone, or tablet terminal. Note that the terminal device 300 and the information processing device 100 may be composed of the same electronic device.
[0038] The database 400 holds various types of data that the information processing device 100 uses to generate 4D data in response to user instructions. These types of data include the characteristics of the camera 200, the camera model used for 4D data generation, and the reconstruction model. The database 400 is comprised of a server, cloud, personal computer, etc. Note that the database 400 and the information processing device 100 may be comprised of the same electronic device.
[0039] [Configuration of Information Processing Device 100] Next, the configuration of the information processing device 100 will be described with reference to Figure 4.
[0040] The acquisition unit 101 receives video data and metadata transmitted from multiple cameras 200, as well as instruction information transmitted from the terminal device 300, and outputs them to the metadata addition unit 102.
[0041] The metadata addition unit 102 generates or acquires new metadata that is not already attached to the video data and adds it to the video data.
[0042] The data selection unit 103 selects video data necessary for generating a reconstruction result or excludes unnecessary video data based on either or both of the metadata and the instruction information. Note that since a plurality of video data is required to generate the reconstruction result, a plurality of video data is transmitted from the plurality of cameras 200 to the information processing apparatus 100, and the data selection unit 103 selects a plurality of video data therefrom.
[0043] The data correction unit 104 performs correction processing for reducing the differences between the plurality of video data selected by the data selection unit 103. Since there are differences in camera parameters, shooting settings, shooting positions, shooting poses, etc. among the plurality of cameras 200, there are differences between the plurality of video data.
[0044] The reconstruction processing unit 105 generates a reconstruction result from a plurality of video data by a known reconstruction process.
[0045] The output unit 106 outputs the reconstruction result generated by the reconstruction processing unit 105 to the camera 200, the terminal device 300, the database 400, other external devices, etc. via the network. Further, the output unit 106 may transfer the reconstruction result to an external device via a recording medium such as a USB flash memory or an SD memory card.
[0046] Next, referring to FIG. 5, the hardware configuration of the information processing apparatus 100 will be described.
[0047] The CPU (Central Processing Unit) 151 functions as an arithmetic processing unit that performs various processes, and controls the entire information processing apparatus 100 and each part thereof. The CPU 151 executes various processes according to a program stored in the ROM (Read Only Memory) 152 or a program loaded from the storage unit 160 into the RAM (Random Access Memory) 153. In the RAM 153, data and the like necessary for the CPU 151 to execute various processes are appropriately stored. Each processing block constituting the information processing apparatus 100 can be realized by a processor constituted by the CPU 151, the ROM 152, and the RAM 153 executing a program.
[0048] The CPU 151, ROM 152, and RAM 153 are interconnected via a bus 154, and the bus 154 is connected to a bridge 155.
[0049] An interface 157 is connected to the bridge 155 via a bus 156.
[0050] An input unit 158, a display unit 159, a storage unit 160, a drive 161, a connection port 162, and a communication unit 163 are connected to the interface 157.
[0051] The input unit 158 is various operators and operation devices such as, for example, a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, a remote controller, etc. An operation of the user is detected by the input unit 158, and a signal corresponding to the input operation is interpreted by the CPU 151.
[0052] The display unit 159 is a liquid crystal display, an organic EL display, etc. that perform display of video, images, GUI (Graphical User Interface), messages, etc.
[0053] The storage unit 160 is a large-capacity storage medium such as, for example, a hard disk, a flash memory, etc. Various applications, data, information, etc. are stored in the storage unit 160.
[0054] A removable storage medium 164 can be connected to the information processing apparatus 100 via the drive 161. The removable storage medium 164 includes a USB flash memory, an SD memory card, a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. Video data and metadata can be supplied from the camera 200 to the information processing apparatus 100 using the removable storage medium 164.
[0055] The drive 161 can read data files such as programs used for each process from the removable storage medium 164. The read data files are stored in the storage unit 160. In addition, programs and other data read from the removable storage medium 164 are installed in the storage unit 160 as needed. Furthermore, the information processing device 100 may transfer information and data to an external device via the removable storage medium 164.
[0056] External devices 165 can be connected to the information processing device 100 via the connection port 162.
[0057] The communication unit 163 includes various communication terminals and communication modules for communication processing via networks such as the Internet, and communication with various devices via wired / wireless communication, bus communication, etc. External devices can be connected to the information processing device 100 via the communication unit 163. The communication method can be either wired or wireless. Communication methods include cellular communication, 4G, 5G, Wi-Fi, Bluetooth®, NFC, Ethernet®, HDMI®, USB, etc. The information processing device 100 can receive video data and metadata transmitted from the camera 200 via communication through the communication unit 163. The information processing device 100 can also communicate with terminal devices 300 and databases 400 via communication through the communication unit 163.
[0058] Other external devices can be connected via connection port 162 or communication unit 163.
[0059] Note that the information processing device 100 does not need to have all the configurations shown in Figure 5. For example, if the information processing device 100 only performs processing and outputs the reconstruction results to the outside, the display unit 159 is not necessary.
[0060] The functions realized by the components described herein may be implemented in a circuit or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), CPUs, conventional circuits, and / or combinations thereof, programmed to realize the described functions. A processor, including transistors and other circuits, is considered a circuit or processing circuitry. A processor may be a programmed processor that executes a program stored in memory. In this specification, circuitry, unit, and means are hardware programmed to realize or execute the described functions. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to realize or execute the described functions. If such hardware is a processor that is considered a type of circuitry, then such circuitry, means, or unit is a combination of hardware and software used to constitute such hardware and / or processor.
[0061] In the information processing device 100, for example, programs and applications for processing this technology can be installed via network communication by the communication unit 163 or via a removable storage medium 164. Alternatively, programs and applications may be pre-stored in the ROM 152, storage unit 160, etc.
[0062] The information processing device 100 is configured as described above. The information processing device 100 may be configured as a standalone device, or it may be configured as an electronic device having information processing and communication functions, such as a personal computer, smartphone, or tablet terminal. Alternatively, the information processing device 100 and the information processing method may be realized by an electronic device having computer functions executing a program. The program may be pre-installed on the electronic device, or it may be distributed via download or storage media for the user to install.
[0063] The information processing device 100 may be configured on a cloud server. Furthermore, the information processing device 100 may transmit the generated reconstruction results to an external device or cloud server.
[0064] A cloud server is not limited to being composed of a single computer device; it may also be composed of a system of multiple computer devices. These multiple computer devices may be systematized, for example, by a LAN (Local Area Network). Alternatively, multiple computer devices located in remote locations may be systematized by a VPN (Virtual Private Network) using the internet, etc. These multiple computer devices may include computer devices that constitute a group of servers (cloud) available through cloud computing services.
[0065] [Processing in the Information Processing Device 100] Next, referring to Figure 6, the processing in the information processing device 100 will be described.
[0066] In step S11, the acquisition unit 101 acquires the video data and metadata transmitted from the camera 200 and the instruction information transmitted from the terminal device 300, and outputs them to the metadata addition unit 102.
[0067] The user can input instruction information via the terminal device 300, which includes priorities for generating the reconstruction results. Specifically, these include prioritizing the reproduction of a specific subject, prioritizing image quality, prioritizing processing speed, prioritizing communication volume, prioritizing privacy, and prioritizing authenticity.
[0068] Prioritizing the reproduction of a specific subject means prioritizing the generation of reconstruction results for a specific subject in video data. When specifying priority for the reproduction of a specific subject, the user must also input the identification information of that specific subject as instruction information. The identification information of a specific subject is an image of the person's face, etc., if the specific subject is a person. Alternatively, the identification information of a specific subject may be the name or ID of the specific subject. If the name or ID of the specific subject is used as the identification information, the acquisition unit 101 can acquire the face image of the specific subject from the database 400 based on the name or ID by registering face images associated with the names or IDs of multiple subjects in the database 400 in advance.
[0069] The user can freely decide which subject to photograph. Furthermore, the subject is not limited to people; it can be any object that can be photographed. In this case, as with the facial image described above, the user can input an image of the object as identification information, or the acquisition unit 101 may acquire an image of the object from the database 400 based on its name or other information.
[0070] Prioritizing image quality means prioritizing the signal-to-noise ratio (S / N), gradation, and artifact reduction of the generated reconstruction result. The system may also allow users to set specific image quality settings by inputting specific resolution values or videos representing their ideal image quality.
[0071] Prioritizing processing speed means prioritizing the reduction of processing time required to generate the reconstruction results. Prioritizing communication volume means prioritizing the reduction of communication volume in the information processing system 10 by restricting the transmission and reception of video data and metadata from the camera 200. Prioritizing privacy means prioritizing the protection of individual privacy by restricting the use of video data that shows specific individuals. Prioritizing authenticity means prioritizing the assurance of authenticity by restricting the use of video data with a history of processing or video data generated by artificial intelligence (AI).
[0072] Figure 7 shows an example of where instruction information is stored when propagating it to each block. When outputting instruction information to the metadata addition unit 102 or other processing blocks, the instruction information may be stored in the scene information directory as a file system, for example, as shown in Figure 7.
[0073] Figure 8 shows an example of where instruction information is stored when propagating it to each block. As shown in Figure 8, the reconstruction project information may be configured as a database, and the instruction information may be stored as data within it. In addition, the single or multiple algorithms used in each processing block determined based on the instruction information may be stored in the database as a list of applicable algorithms. Metadata attached to the video data may also be stored in the database.
[0074] Next, in step S12, the metadata addition unit 102 adds new metadata to the video data.
[0075] The position and orientation information of camera 200 is necessary for generating the reconstruction results. Therefore, if the position and orientation information of camera 200 is not attached as metadata to the video data, the metadata attachment unit 102 estimates the position and orientation of camera 200 by performing a known position and orientation estimation process as an analysis process on the video data, and attaches the position and orientation information to the video data as new metadata. The estimation of the camera's position and orientation based on image data can be performed using VPS (Visual Positioning System) or SfM (Structure from Motion), etc.
[0076] Furthermore, the position and orientation of multiple cameras 200 can be estimated using audio data acquired by microphones equipped on each camera 200. The distance from camera 200 to a specific sound source and the direction of the sound source relative to camera 200 can be estimated from the volume of the audio data. If there are multiple cameras 200 equipped with microphones, the position of the sound source can be estimated from the distance to the sound source and the direction of the sound source relative to each camera 200, and the position of camera 200 can also be estimated based on the position of the sound source.
[0077] Furthermore, since the audio data changes depending on the tilt of the microphone, the posture of camera 200, whose posture is unknown, can be estimated from the difference between the audio data acquired by the microphone of camera 200, whose posture is known, and the audio data acquired by the microphone whose posture is unknown.
[0078] The position and orientation of camera 200, whose position and orientation are unknown, can also be estimated based on the position and orientation information of camera 200 transmitted from a calibrated camera 200 whose position and orientation are known.
[0079] The metadata attachment unit 102 can also estimate the position and orientation of the camera 200 and attach the position and orientation information to the video data as metadata using this method.
[0080] Furthermore, the metadata addition unit 102 can perform analysis processing on the video data to generate new metadata and add it to the video data. For example, if the instruction information indicates prioritizing the reproduction of a specific subject, the metadata addition unit 102 performs face detection processing as part of the analysis process to detect the specific subject in the video data. Based on the detection results, it adds to the video data as new metadata the identification information of the video data in which the specific subject exists, the position information of the specific subject in the video, and information of the frame image in which the specific subject is captured. The identification information of the specific subject (such as a face image) is stored in the database 400 in advance, and the metadata addition unit 102 can read this identification information to compare the face detected by the face detection process with the specific subject. AI may also be used for the analysis of the video data.
[0081] The metadata addition unit 102 can perform subject detection processing and scene detection processing on video data and add the detection results to the video data as metadata.
[0082] Furthermore, the metadata addition unit 102 can also add information obtained by referring to the database 400 based on the metadata added to the video data as new metadata to the video data.
[0083] For example, camera parameters linked to the photographer ID and camera identification information are stored in the database 400 in advance. The metadata addition unit 102 can also obtain camera parameters of the camera 200 that shot the video data by referring to the database 400 based on the photographer ID and camera identification information attached to the video data as metadata, and add them to the video data as new metadata.
[0084] For example, seating information and venue layout information for events such as weddings, where seating is predetermined, are stored in the database 400 as information linked to the photographer ID. The metadata addition unit 102 can obtain the photographer's seating information by referring to the database 400 based on the photographer ID as metadata added to the video data. Once the seating information is obtained, the photographer's position can be estimated from the venue layout information, etc. Therefore, the metadata addition unit 102 can add this estimated position information to the video data as new metadata.
[0085] For example, detailed event information and schedule information are pre-stored in the database 400 as information linked to time information and location information. The metadata addition unit 102 can obtain detailed event information and schedule information by referring to the database 400 based on time information and location information as metadata attached to the video data, and can add scene information and other information that can be estimated from this information as new metadata to the video data.
[0086] Furthermore, if it is necessary to estimate the camera parameters of the camera 200 that captured the video data, in order to improve the estimation accuracy, images captured by the camera 200 may be obtained from the database 400 based on the camera identification information as metadata acquired by the acquisition unit 101.
[0087] The metadata addition unit 102 may change the algorithm used when adding new metadata to the video data according to the instruction information. For example, if camera parameters of camera 200 are required to prioritize the quality of the reconstruction result, a Brown model that takes image distortion into account is used. For example, when performing face detection or subject detection and quality is the priority, a large DNN (Deep Neural Network) such as a Vision-Transformer model is used. When processing speed is the priority, a small DNN such as MobileNet is used.
[0088] Furthermore, processing by the metadata addition unit 102 may be omitted depending on the instruction information. For example, if the instruction information indicates speed priority, based on metadata such as camera identification information acquired by the acquisition unit 101, only video data transmitted from cameras 200 that have been calibrated and whose position and orientation are known will be used, and video data transmitted from cameras 200 that have not been calibrated and whose position and orientation are unknown will not be used. As a result, there is no need for the metadata addition unit 102 to estimate the position and orientation of the cameras 200 and generate new metadata, so processing by the metadata addition unit 102 can be omitted and processing can be sped up.
[0089] Returning to the explanation of Figure 6, in step S13, the data selection unit 103 selects video data to be used to generate the reconstruction result and removes unnecessary video data based on either or both metadata and instruction information.
[0090] Unnecessary video data includes, for example, footage whose shooting time does not match the desired time or other video data. Whether or not the shooting time does not match the desired time or other video data can be determined based on the shooting time information attached as metadata to the video data.
[0091] Furthermore, whether or not the shooting time of a video data file does not match that of other video data can also be determined based on scene information, which is metadata attached to the video data. For example, if the event to be reconstructed is a wedding, and the scene information indicates that video data A is the cake-cutting scene, and the other video data files are the speech scenes, then it can be determined that video data A does not match the scene of the other data files, meaning that the shooting time does not match.
[0092] Furthermore, if the scene information indicates that video data A consists of a cake-cutting scene and a greeting scene, and that other videos besides video data A consist only of a greeting scene, the data selection unit 103 can also exclude only the cake-cutting scene from video data A. Selection and exclusion can be performed not only on a video data basis, but also on a scene-by-scene or time-based basis within the video.
[0093] Furthermore, unnecessary video data is defined as video data whose shooting viewpoint does not match the desired viewpoint or other video data. Whether or not video data has a mismatched shooting viewpoint can be determined from the position information and orientation information of camera 200 as metadata, as well as the subject detection results. For example, as shown in Figure 9, suppose there is camera 200A shooting video data A, camera 200B shooting video data B, camera 200C shooting video data C, and camera 200D shooting video data D. From the position information, orientation information, and subject detection results of camera 200, it can be seen that the specific subject S is within the field of view of cameras 200A, 200B, and 200C, but the specific subject S is not within the field of view of camera 200D. In this case, the shooting viewpoint of camera 200D is different from that of cameras 200A, 200B, and 200C, and the data selection unit 103 excludes video data D shot by camera 200D.
[0094] Additionally, video data that does not include metadata necessary for generating the reconstruction results (e.g., position and orientation information), video data that does not match the instruction information, video data that does not include identification information for a specific subject when that subject is specified in the instruction information, video data with a shallow depth of field (below a predetermined value), and video data with large motion blur (above a predetermined value) may also be excluded as unnecessary video data.
[0095] The type of video data to be excluded may be pre-configured in the information processing device 100, or it may be set by the user.
[0096] Next, in step S14, the data correction unit 104 performs a correction process on the multiple video data selected by the data selection unit 103 to reduce the differences between the video data. A specific example of the correction process will be described below.
[0097] The first specific example of the correction process is the correction of frame positions based on audio data. When frame positions do not match between multiple video data using shooting time information as metadata, or when shooting time information is not available as metadata, the frame positions between multiple video data can be corrected based on the audio contained in the video data.
[0098] The video data A, captured by camera 200A, and the video data B, captured by camera 200B, shown in Figure 10, are video data captured by different cameras, and both contain multiple consecutive frame images and audio data. It is assumed that the playback start time can be obtained in seconds from the attached shooting time information for both video data A and video data B.
[0099] The data correction unit 104 first acquires audio data A from video data A and audio data B from video data B. Next, it aligns the audio waveforms of audio data A and audio data B. For example, alignment can be performed by searching in the time direction and shifting audio data B so that the error in audio data B is minimized using audio data A as a reference. Then, all frames of video data B are also shifted by the amount of the shift in audio data B. As a result, the frame positions of video data A and video data B match. In this way, the frame positions between multiple video data can be corrected.
[0100] A second specific example of the correction process is color difference correction. Color differences in video data can be corrected based on the subject detection results in the video data. As shown in Figure 11, video data A, shot with camera 200A, and video data B, shot with camera 200B, are video data shot with different cameras, and both are assumed to have the face detection results (corresponding coordinates) of a specific subject S attached as metadata.
[0101] The data correction unit 104 first obtains the average RGB values of the face detection region from video data A and video data B, respectively. Next, it performs color conversion on video data A and video data B so that their average RGB values match. If there is only one color, it feeds it back to WBG. It estimates the color gamut of each device by referring to the database 400 and performs a predetermined color conversion process based on both color gamuts. If three or more colors are obtained, it generates and applies an RGB2RGB conversion matrix. This is a method of correcting color differences from the face region based on the premise that the color tones of people's faces should be the same, but color difference correction may also be performed based on regions other than the face. In this way, video data B' with color difference correction applied can be obtained.
[0102] A third specific example of the correction process is correction by generating complementary images. In this third example, as shown in Figure 12, the data correction unit 104 has a complementary image generation function. The data correction unit 104 generates high-quality complementary images from the frame images that make up the video data using its complementary image generation function. Techniques including deep learning-based methods such as diffusion models, Optical Flow, CNN (Convolutional Neural Network), GAN (Generative Adversarial Network), and morphing can be used to generate complementary images.
[0103] In the example shown in Figure 12, images A1 and A2 are frame images that make up video data captured by camera 200, and image B is an image read from database 400. All images show a specific subject S specified by the instruction information. In images A1 and B, the face of the specific subject S is facing forward, while in image A2, the face of the specific subject S is facing to the side.
[0104] The data correction unit 104 can generate a complementary image C in which the face of the specific subject S is turned at an angle by performing complementary image generation processing based on images A1, A2, and B. This complementary image C expands the range of images used to generate the reconstruction result.
[0105] In Figure 12, a interpolated image is generated based on the frame images that make up the video data generated by the camera 200 and the image read from the database 400. However, the interpolated image may also be generated based only on the frame images that make up the video data generated by the camera 200.
[0106] Furthermore, the data correction unit 104 may perform the following correction processing to reduce differences between multiple video data.
[0107] The data correction unit 104 may, if the color tones of a specific subject differ in multiple video data, detect the same region, for example, the skin tone of the person who is the specific subject, and perform color matching processing.
[0108] The data correction unit 104 may apply super-resolution processing to video data if, among the multiple video data, there is video data with a different resolution from the other video data.
[0109] The data correction unit 104 may generate an interpolated image to match the frame rates if there is video data among the multiple video data that has a different frame rate from the other video data.
[0110] The characteristic values of the camera 200 and training images may be stored in the database 400 in advance, and the data correction unit 104 may read them and use them for correction processing.
[0111] If the instruction information indicates that quality should be prioritized, the data correction unit 104 may acquire or estimate highly accurate camera parameters for the correction process. If camera parameters are already stored in the database 400, the data correction unit 104 can acquire camera parameters from the database 400 based on the camera identification information as metadata acquired by the acquisition unit 101.
[0112] If camera parameters are not stored in the database 400, the data correction unit 104 may estimate the camera parameters. In this case, based on the camera identification information acquired by the acquisition unit 101, images taken with the same camera 200 can be acquired from the database 400 and the data can be expanded.
[0113] Returning to the explanation of Figure 6, in step S15, the reconstruction processing unit 105 generates a reconstruction result from the multiple video data corrected by the reconstruction process.
[0114] Next, in step S16, the output unit 106 outputs the reconstruction result. After the event that was the target of the reconstruction process has finished, the reconstruction result can be screened or distributed to the event participants. Alternatively, pre-rendering can be applied to the reconstruction result to generate a stereo video, which can then be delivered to the viewer's device. Furthermore, by delivering the reconstruction result, the viewer can view it in real time on their device by performing rendering.
[0115] The processing of the first embodiment is carried out as described above. According to the first embodiment, the user's desired reconstruction result can be generated with high accuracy by selecting video data to be used for generating the reconstruction result based on metadata, or based on metadata and instruction information. Furthermore, it is possible to generate the reconstruction result with reduced computational cost.
[0116] Furthermore, even if there are differences in the multiple video data captured by multiple cameras 200, these differences can be reduced by applying correction processing, thereby generating highly accurate reconstruction results.
[0117] The metadata addition unit 102 adds new metadata to the video data, enabling the generation of highly accurate reconstruction results even when metadata cannot be obtained from the camera 200 or when there is insufficient metadata.
[0118] By using the database 400 to acquire new metadata, it is possible to generate highly accurate reconstruction results even when metadata cannot be obtained from the camera 200 or when there is insufficient metadata.
[0119] This technology can verify the actions of others by analyzing the relationship between instruction information and generated image data, as well as the relationship between instruction information and metadata. In analyzing the relationship between instruction information and metadata, it is checked whether the output results change based on the instruction information, and whether the output results change depending on the presence or absence of specific metadata.
[0120] <Second Embodiment> [Configuration of Information Processing Device 100] Referring to Figure 13, the configuration of the information processing device 100 in the second embodiment will be described. The configuration of the information processing system 10 is the same as in the first embodiment.
[0121] In the second embodiment, the information processing device 100 includes a model selection unit 107. The other components of the information processing device 100 are the same as in the first embodiment and will not be described.
[0122] In the second embodiment, the reconstruction processing unit 105 has a plurality of reconstruction models in advance for generating reconstruction results in the reconstruction process.
[0123] Based on the instruction information output from the acquisition unit 101, the model selection unit 107 adaptively selects a reconstruction model that matches the user's instructions from among multiple reconstruction models held by the reconstruction processing unit 105. The model selection unit 107 outputs the selection result to the reconstruction processing unit 105.
[0124] [Processing in the Information Processing Device 100] Next, referring to Figure 14, the processing of the information processing device 100 in the second embodiment will be described. Steps S11 to S16 are the same as in the first embodiment, so their description will be omitted.
[0125] In step S21, the model selection unit 107 adaptively selects a reconstruction model that matches the user's instructions based on the instruction information input from the acquisition unit 101. The model selection unit 107 outputs the selection result to the reconstruction processing unit 105. The selection of the reconstruction model is performed, for example, as follows.
[0126] The model selection unit 107 can select a reconstruction model according to the characteristics of the captured data. If the video data has characteristics that cannot be corrected by the data correction unit 104, it selects a reconstruction model that is robust to those characteristics. For example, if the video data has a shallow depth of field that cannot be corrected, it selects a reconstruction model that uses a thin-lens camera model that can simulate image blur. Also, if the video data has motion blur that cannot be corrected, it selects a reconstruction model that simulates and renders the blur considering the length of the exposure time.
[0127] If the instruction information specifies priorities, the model selection unit 107 can select a reconstruction model that conforms to those priorities. If the instruction information indicates speed priority, it selects a reconstruction model using a pinhole camera model that renders quickly. Alternatively, it can select a reconstruction model that has lower expressive power but is easier to optimize. If the instruction information indicates quality priority, it selects a reconstruction model with higher expressive power (for example, one with a large number of parameters such as DNN, or one with parameters that can handle time-series changes including color changes in addition to shape changes).
[0128] Then, in step S22, the reconstruction processing unit 105 generates a reconstruction result from multiple video data by performing a reconstruction process using the reconstruction model selected by the model selection unit 107.
[0129] In Figure 14, the processing by the model selection unit 107 in step S21 is performed after step S14, but this processing can be performed at any time as long as it is after the processing by the acquisition unit 101 in step S11 and before the processing by the reconstruction processing unit 105 in step S22.
[0130] The processing of the second embodiment is carried out as described above. According to the second embodiment, in addition to the effects of the first embodiment, by selecting a model to be used for generating the reconstruction results, it is possible to generate reconstruction results with higher accuracy and reconstruction results that are better suited to the user's instructions.
[0131] <Third Embodiment> [Configuration of Information Processing Device 100] Referring to Figure 15, the configuration of the information processing device 100 in the third embodiment will be described. Note that the configuration of the information processing system 10 is the same as in the first embodiment.
[0132] In the third embodiment, the information processing device 100 includes a request generation unit 108. The other components of the information processing device 100 are the same as in the first embodiment and will not be described.
[0133] The request generation unit 108 determines whether there is any missing information in the video data and metadata acquired by the acquisition unit 101. If there is missing information, it generates a request to the camera 200 to send the video data and metadata that include the missing information.
[0134] In the third embodiment, the camera 200 includes a shooting condition processing unit 201 and a transmission determination unit 202.
[0135] The shooting condition processing unit 201 performs processing to change the shooting conditions of the camera 200 based on the request generated by the request generation unit 108. Changing the shooting conditions includes switching the functions of the camera 200, turning the functions of the camera 200 on or off, and changing the shooting position or orientation of the camera 200.
[0136] The transmission determination unit 202 determines whether the video data and metadata generated by the camera 200 after receiving the request contain information corresponding to the request. If the video data and metadata contain information corresponding to the request, the camera 200 transmits the video data and metadata to the information processing device 100. On the other hand, if the video data and metadata do not contain information corresponding to the request, the camera 200 does not transmit the video data and metadata to the information processing device 100.
[0137] For example, the transmission determination unit 202 functions when the instruction information instructs a reduction in the amount of communication, and by not transmitting video data or metadata that does not contain the information corresponding to the request, the amount of communication from the camera 200 to the information processing device 100 can be reduced. Furthermore, the computation cost in the information processing device 100 and the communication cost in the information processing system 10 can also be reduced.
[0138] [Processing in the Information Processing Device 100] Next, referring to Figure 16, the processing of the information processing device 100 in the third embodiment will be described. Steps S11 to S16 are the same as in the first embodiment, so the explanation will be omitted.
[0139] In step S31, the request generation unit 108 determines whether there is any missing information in the video data and metadata. The request generation unit 108 maintains, for example, a table indicating the types of metadata that should be added to the video data in advance, and determines the presence or absence of missing information based on that table. Alternatively, the request generation unit 108 can determine the presence or absence of missing information by evaluating whether the reconstruction result generated by the reconstruction processing unit 105 satisfies the priorities indicated by the instruction information.
[0140] If there is no missing information, the process proceeds to step S16 (No. in step S31). On the other hand, if there is missing information, the process proceeds to step S32 (Yes in step S31). In step S32, the request generation unit 108 generates a request to the camera 200 to send video data and metadata containing the missing information. This request is transmitted to the camera 200 via the communication function of the information processing device 100 and the network.
[0141] If the missing information is video data of a specific subject filmed from a predetermined direction, the request will be to film the specific subject from that direction. If the missing information is specific metadata such as posture information, the request will be to retrieve the specific metadata and add that metadata to the video data. If the missing information is a specific subject, the request will be to film a composition that includes the specific subject. If the missing information is video data in a specific format, the request will be to send video data in that specific format.
[0142] Furthermore, the request may include instructions for using camera 200 functions (SLAM function, time synchronization function, signature function, etc.) to fulfill the request. The request may also include instructions for turning camera 200 functions (e.g., image stabilization functions such as EIS (Electronic Image Stabilization) and OIS (Optical Image Stabilization)) on or off to fulfill the request. For example, if the instruction information prioritizes the quality of the reconstruction result, and accurate camera 200 attitude information is required for that, the request can instruct the camera 200 to disable its image stabilization function. The request may also include instructions for turning functions other than image stabilization on or off.
[0143] The information processing device 100 uses its communication function to send a request to the camera 200. Upon receiving the request, the shooting condition processing unit 201 of the camera 200 changes the settings of the camera 200 based on the request. Furthermore, if the request requires the photographer to take a picture from a specific position or direction, the shooting condition processing unit 201 may display a message instructing the photographer to do so on the monitor of the camera 200 or output an audio message.
[0144] The processing of the third embodiment is carried out as described above. According to the third embodiment, in addition to the effects of the first embodiment, metadata can be added and a highly accurate reconstruction result can be generated by generating a request depending on whether or not there is missing information.
[0145] Furthermore, the shooting condition processing unit 201 of the camera 200 may perform processing to change the shooting conditions without a request from the request generation unit 108 for certain priorities among the priorities indicated by the instruction information. For example, if the instruction information indicates prioritizing image quality, the shooting condition processing unit 201 simply changes the camera 200 to high-quality mode, so no request is required. However, if no request is required, it is desirable to provide feedback on the quality and composition of the reconstruction result generated from the video data obtained with the shooting conditions changed without a request.
[0146] On the other hand, the shooting condition processing unit 201 needs a request to change the settings of the camera 200 when it is not possible to know whether there is missing information in the video data and metadata until after the camera 200 has finished shooting. For example, if several cameras 200 have been used to shoot and there is no video data shot at a specific position or direction, the request generation unit 108 will need to generate a request.
[0147] <Fourth Embodiment> [Configuration of Information Processing Device 100] Referring to Figure 17, the configuration of the information processing device 100 in the fourth embodiment will be described. Note that the configuration of the information processing system 10 is the same as in the first embodiment.
[0148] In the fourth embodiment, the information processing device 100 includes a rendering unit 109 and an output determination unit 110. The other components of the information processing device 100 are the same as in the first embodiment and will not be described.
[0149] The rendering unit 109 generates an output video from the reconstruction result, with a predetermined viewpoint. The output video may be a 2D video or a 3D video.
[0150] The output determination unit 110 evaluates each output video when the rendering unit 109 generates multiple output videos with different viewpoints and generation models, and selects which output video to output based on the evaluation result.
[0151] [Processing in the Information Processing Device 100] Next, referring to Figure 18, the processing of the information processing device 100 in the fourth embodiment will be described. Steps S11 to S15 are the same as in the first embodiment, so their description will be omitted.
[0152] In step S41, the rendering unit 109 generates an output video from the reconstruction result under a predetermined viewpoint. The output video may be a 2D video or a 3D video. The type of output video and the viewpoint of the output video may be specified by the user via the terminal device 300, or they may be automatically determined based on a recommended composition derived using AI. The rendering unit 109 may perform the derivation of the recommended composition using AI, or the information processing device 100 may be equipped with a processing unit for executing this process.
[0153] Next, in step S42, the output determination unit 110 evaluates multiple types of output videos with different viewpoints and generation models generated by the rendering unit 109, and selects an output video to be output externally based on the evaluation result. For example, multiple output videos may be evaluated and the one with the highest evaluation value may be output. Note that this process is unnecessary if the rendering unit 109 generates only one video.
[0154] Indicators used to evaluate video and image data include, for example, SNR (Signal to Noise Ratio), PSNR (Peak Signal to Noise Ratio), and SSIM (Structural SIMilarity).
[0155] Whether or not the output video conforms to the priorities in the instruction information can also be used as an evaluation criterion. If the output determination unit 110 determines that the evaluation result of the output video generated by the rendering unit 109 does not meet a predetermined standard, it may output only error information without outputting the output video to the outside.
[0156] Next, in step S43, the output unit 106 outputs the output video. In addition to the output video generated by the rendering unit 109, the output unit 106 may also output the reconstruction result, as in the first embodiment.
[0157] The processing of the fourth embodiment is carried out as described above. According to the fourth embodiment, in addition to the effects of the first embodiment, the information processing device 100 can generate output videos such as 2D videos and 3D videos from the reconstruction results and output them externally.
[0158] <Modifications> Although embodiments of this technology have been described in detail above, this technology is not limited to the embodiments described above, and various modifications are possible based on the technical concept of this technology.
[0159] The data used as input for the reconstruction process is not limited to captured images; it can also include image data or video data (multiple frame images that make up a video) generated by image generation AI or image generation tools.
[0160] The metadata addition unit 102 does not perform any processing. If the video data transmitted from the camera 200 has sufficient metadata necessary for generating the reconstruction result, such as the position and orientation information of the camera 200, the metadata addition unit 102 does not need to perform any processing to generate the reconstruction result from the video data.
[0161] This technology can be implemented by combining any two or more of the first to fourth embodiments.
[0162] For example, as shown in Figure 19, the information processing device 100 and the information processing system 10 may be configured by combining the second to fourth embodiments.
[0163] As shown in Figure 20, the information processing device 100 may not include a data correction unit 104. It is possible to generate reconstruction results without applying correction processing to the captured data selected by the data selection unit 103. This is also true in the second to fourth embodiments. However, by applying correction processing to the video data with the data correction unit 104, it is possible to generate reconstruction results with higher accuracy.
[0164] As shown in Figure 21, the selection result by the data selection unit 103 may be output to the metadata addition unit 102. This allows the metadata addition unit 102 to re-analyze the video data based on the selection result by the data selection unit 103 and add new metadata to the video data. For example, if the instruction information indicates speed priority, first, only video data shot at a predetermined location is selected based on the location information, which is metadata previously added to the video data. Then, the metadata addition unit 102 uses only the video data selected by the data selection unit 103 to perform computationally expensive and time-consuming analysis processing, such as accurate pose estimation of the camera 200, and adds the pose information of the camera 200 as new metadata to the video data.
[0165] The metadata addition unit 102 may also be provided by the camera 200. In that case, the camera 200 analyzes the video data itself, adds the metadata generated from the analysis results to the video data, and transmits it to the information processing device 100.
[0166] The information processing device 100 may perform the selection of video data, while other devices such as a server may perform the correction and reconstruction of the video data. Alternatively, the information processing device 100 may perform both the selection and correction of video data. In that case, the information processing device 100 transmits the selected video data to the other device.
[0167] As shown in Figure 22, the instruction information input from the terminal device 300 to the information processing device 100 may be directly input not only to the acquisition unit 101, but also to the data selection unit 103, the data correction unit 104, and the reconstruction processing unit 105, respectively.
[0168] Instructions for generating reconstruction results do not necessarily need to be input by the user via the terminal device 300. As shown in Figure 23, the information processing device 100 may include an instruction generation unit 111, which may automatically generate instruction information based on video data captured by the camera 200 and information obtained from the database 400.
[0169] For example, if the target of the reconstruction process is an event such as a wedding, the instruction generation unit 111 automatically estimates that the event is a wedding from video data and scene information obtained from the database 400, and automatically generates an instruction prioritizing the reproduction of a specific subject (for example, the bride).
[0170] Furthermore, the instruction generation unit 111 refers to the delivery date information stored in the database 400 and automatically generates an instruction prioritizing processing speed if the period until the delivery date is short (e.g., below a predetermined threshold). On the other hand, the instruction generation unit automatically generates an instruction prioritizing quality if the period until the delivery date is long (e.g., above a predetermined threshold).
[0171] The instruction information generated by the instruction generation unit 111 may receive feedback from the reconstruction processing unit 105 regarding the generation of reconstruction results. For example, the instruction generation unit 111 may evaluate the reconstruction results, and if the quality of the reconstruction results does not meet a predetermined standard, it may generate instructions that prioritize quality and generate the reconstruction results again.
[0172] Furthermore, the instruction information itself is not mandatory; for example, the reconstruction result may be generated by each processing block performing a predetermined process based on metadata. In that case, instruction information from the terminal device 300 or the instruction generation unit 111 is unnecessary.
[0173] The data stored in database 400 may be made available to different external information processing systems 20 or external devices 500 or external services, as shown in Figure 24. However, this is limited to data for which copyright permission has been obtained or data that does not pose any copyright issues. Information processing system 10 and the external information processing systems 20 or external devices 500 are connected by a cloud or network.
[0174] Methods for providing data stored in database 400 to external parties include outputting search results (whether data exists or not, and the relevant data) in response to text searches, and performing metadata searches in an external information processing system 10 when the camera 200 is taking pictures at different times, and outputting search results (whether data exists or not, and the relevant data).
[0175] If the reconstruction process is for an event with an organizer, such as a live concert or sports event, the organizer may provide an application for the information processing system 10 to multiple cameras 200. This application processes metadata such as the photographer's identification ID, shooting location, and camera parameters, which are then transmitted to the information processing device 100 and database 400. This allows for efficient aggregation of metadata.
[0176] Furthermore, the operator may allow users who film and provide video data and metadata to receive compensation (ticket discounts, provision of reconstructed results) through the application.
[0177] Furthermore, camera 200 includes both fixed cameras provided by the organizers and cameras owned by general spectators. Video data and metadata transmitted from general spectator cameras may be used to supplement the video data and metadata generated by the fixed cameras.
[0178] If multiple cameras 200 of the same model exist within a predetermined size range, and the instruction information specifies prioritizing processing speed or communication volume, one of them may be selected as the camera 200 to transmit video data and metadata. This reduces the amount of communication from multiple cameras 200 to the information processing device 100 and also increases the processing speed.
[0179] In the embodiment, it was explained that a user, such as an event organizer, uses the information processing device 100 to generate the reconstruction results. However, an individual who is not an organizer may also use the information processing device 100 to generate the reconstruction results. Furthermore, the user and the photographer may be the same person, the user and the viewer may be the same person, or the user, photographer, and viewer may be the same person. In addition, anyone can use the information processing device 100 of this technology, and the use and purpose of the information processing device 100 are not limited to events as described above, but may be used for any purpose.
[0180] The technology can also take the following configurations: (1) An information processing device comprising: a data selection unit that selects the shooting data to be used to generate the reconstruction result from the multiple shooting data based on metadata corresponding to multiple shooting data generated by shooting with multiple cameras, and instruction information indicating instructions regarding the reconstruction result, or both; and a reconstruction processing unit that generates the reconstruction result by reconstruction processing based on the shooting data selected by the data selection unit. (2) The information processing device according to (1), further comprising a metadata addition unit that adds the metadata to the shooting data. (3) The information processing device according to (2), wherein the metadata addition unit performs analysis processing on the shooting data and generates the metadata from the results of the analysis processing. (4) The information processing device according to (2) or (3), wherein the metadata addition unit obtains new metadata by referring to a database based on the metadata previously added to the shooting data by the camera. (5) The information processing device according to any one of (1) to (4), wherein the metadata is previously added to the shooting data by the camera. (6) The information processing device according to any one of (1) to (5), wherein the instruction information indicates the priorities in generating the reconstruction result. (7) The information processing device according to any one of (1) to (6), wherein the instruction information indicates a specific subject in the shooting data. (8) The information processing device according to any one of (1) to (7), further comprising a data correction unit that applies a correction process to the shooting data to reduce the differences between a plurality of the shooting data. (9) The information processing device according to (8), wherein the data correction unit applies a correction process to the shooting data based on the metadata. (10) The information processing device according to (8), wherein the data correction unit applies a correction process to the shooting data based on the subject detection result in the shooting data. (11) The information processing device according to (8), wherein the data correction unit applies a correction process to the shooting data based on instruction information relating to the generation of the reconstruction result. (12) The information processing device according to any one of (1) to (11), further comprising a model selection unit that selects a reconstruction model for generating the reconstruction result in the reconstruction processing unit.(13) The information processing apparatus according to (12), wherein the model selection unit determines the reconstruction model based on instruction information regarding the generation of the reconstruction result. (14) The information processing apparatus according to (12), wherein the model selection unit selects the reconstruction model according to the characteristics of the shooting data. (15) The information processing apparatus according to any one of (1) to (14), further comprising a request generation unit that generates a request indicating that the shooting data and metadata contain missing information. (16) The information processing apparatus according to any one of (1) to (15), further comprising a rendering unit that generates a video of a predetermined viewpoint from the reconstruction result. (17) The information processing apparatus according to (16), further comprising an output determination unit that selects the video to be output from a plurality of videos generated by the rendering unit. (18) The information processing apparatus according to any one of (1) to (17), wherein the plurality of cameras have differences in camera parameters, resulting in differences between the plurality of shooting data. (19) An information processing method for selecting the shooting data to be used to generate the reconstruction result from a plurality of shooting data based on either or both metadata corresponding to a plurality of shooting data generated by shooting with a plurality of cameras and instruction information indicating instructions regarding the reconstruction result, and generating the reconstruction result by a reconstruction process based on the selected shooting data. (20) A program for causing a computer to execute the information processing method for selecting the shooting data to be used to generate the reconstruction result from a plurality of shooting data based on either or both metadata corresponding to a plurality of shooting data generated by shooting with a plurality of cameras and instruction information indicating instructions regarding the reconstruction result, and generating the reconstruction result by a reconstruction process based on the selected shooting data.
[0181] 100... Information processing unit 102... Metadata addition unit 103... Data selection unit 104... Data correction unit 105... Reconstruction processing unit 107... Model selection unit 108... Request generation unit 109... Rendering unit 110... Output determination unit 200... Camera
Claims
A data selection unit selects the shooting data to be used to generate the reconstruction result from the multiple shooting data based on metadata corresponding to multiple shooting data generated by shooting with multiple cameras, and instruction information indicating instructions regarding the reconstruction result, or both of the above. A reconstruction processing unit generates a reconstruction result by reconstruction processing based on the captured data selected by the data selection unit. An information processing device equipped with the following features. The system includes a metadata addition unit that adds the metadata to the aforementioned shooting data. The information processing apparatus according to claim 1. The metadata addition unit performs analysis processing on the captured data and generates the metadata from the results of the analysis processing. The information processing apparatus according to claim 2. The metadata addition unit obtains new metadata by referring to a database based on the metadata previously added to the captured data by the camera. The information processing apparatus according to claim 2. The metadata is pre-added to the captured data by the camera. The information processing apparatus according to claim 1. The aforementioned instruction information indicates the priorities in generating the reconstruction results. The information processing apparatus according to claim 1. The aforementioned instruction information indicates a specific subject in the captured data. The information processing apparatus according to claim 1. The system includes a data correction unit that applies a correction process to the captured data to reduce the differences between multiple sets of captured data. The information processing apparatus according to claim 1. The data correction unit performs correction processing on the captured data based on the metadata. The information processing apparatus according to claim 8. The data correction unit performs correction processing on the shooting data based on the subject detection result in the shooting data. The information processing apparatus according to claim 8. The data correction unit performs correction processing on the captured data based on instruction information regarding the generation of the reconstruction result. The information processing apparatus according to claim 8. The reconstruction processing unit includes a model selection unit that selects a reconstruction model for generating the reconstruction result. The information processing apparatus according to claim 1. The model selection unit determines the reconstructed model based on the instruction information regarding the generation of the reconstruction result. The information processing apparatus according to claim 12. The model selection unit selects the reconstruction model according to the characteristics of the captured data. The information processing apparatus according to claim 12. The system includes a request generation unit that generates a request indicating that the missing information is included if there is missing information in the aforementioned shooting data and metadata. The information processing apparatus according to claim 1. The system includes a rendering unit that generates a video from a predetermined viewpoint based on the reconstruction results. The information processing apparatus according to claim 1. The rendering unit includes an output determination unit that selects the video to output from among the multiple videos generated by the rendering unit. The information processing apparatus according to claim 16. Due to differences in camera parameters among the aforementioned multiple cameras, differences exist between the aforementioned multiple shooting data. The information processing apparatus according to claim 1. Based on either or both metadata corresponding to multiple image data generated by shooting with multiple cameras, and instruction information indicating instructions regarding the reconstruction result, the system selects the image data to be used to generate the reconstruction result from the multiple image data, Based on the selected imaging data, a reconstruction process is performed to generate the reconstruction result. Information processing methods. Based on either or both metadata corresponding to multiple image data generated by shooting with multiple cameras, and instruction information indicating instructions regarding the reconstruction result, the system selects the image data to be used to generate the reconstruction result from the multiple image data, Based on the selected imaging data, a reconstruction process is performed to generate the reconstruction result. A program that instructs a computer to execute information processing methods.