Information processing device, information processing method, and program

The information processing apparatus and method enhance the quality of 3D or 4D data generation by strategically selecting and processing image data, addressing the issue of quality deterioration in virtual viewpoint images.

WO2026154863A1PCT designated stage Publication Date: 2026-07-23SONY SEMICON SOLUTIONS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SONY SEMICON SOLUTIONS CORP
Filing Date
2025-12-11
Publication Date
2026-07-23

Smart Images

  • Figure JP2025043257_23072026_PF_FP_ABST
    Figure JP2025043257_23072026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises: a first selection unit that selects multiple pieces of imaging data from among a plurality of pieces of imaging data captured by a camera; and a second selection unit that selects, from the multiple pieces of imaging data selected by the first selection unit, imaging data to be used as an input to reconstruction processing.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Apparatus, Information Processing Method, and Program

[0009]

[0001] The present technology relates to an information processing apparatus, an information processing method, and a program.

[0002] As a conventional technique, a technique for grasping the cause of deterioration in the quality of a virtual viewpoint image in generating a virtual viewpoint image based on a plurality of images has been proposed (Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2019-106617

[0004] However, in a virtual viewpoint image, 3D data, or 4D data which is a reconstruction result generated by reconstruction processing, there is a demand for a technique that can not only grasp the cause of quality deterioration but also generate high-quality data.

[0005] The present technology has been made in view of such problems, and an object thereof is to provide an information processing apparatus, an information processing method, and a program that can generate a high-quality reconstruction result by executing reconstruction processing with a plurality of captured data as input.

[0006] In order to solve the above-described problems, a first technique is an information processing apparatus including a first selection unit that selects from a plurality of captured data captured by a camera, and a second selection unit that selects captured data to be input for reconstruction processing from the plurality of captured data selected by the first selection unit.

[0007] A second technique is an information processing method that selects from a plurality of captured data captured by a camera and selects captured data to be input for reconstruction processing from the selected plurality of captured data.

[0008] A third technique is a program that causes a computer to execute an information processing method that selects from a plurality of captured data captured by a camera and selects captured data to be input for reconstruction processing from the selected plurality of captured data.

[0009] This is an explanatory diagram of 4D data in this technology. This is a block diagram showing the configuration of the information processing system 10. This is a diagram showing an example of the structure of image data. This is a diagram showing the configuration of the processing block of the information processing device 100 in the first embodiment. This is a diagram showing the hardware configuration of the information processing device 100. This is a flowchart showing the processing of the information processing device 100 in the first embodiment. This is a diagram showing the camera 200 and the subject that is the target of reconstruction processing. This is a schematic diagram showing the processing of the first selection unit 102 and the second selection unit 105. This is a diagram showing an example of how the second selection unit 105 selects image data. This is a diagram showing the configuration of the processing block of the information processing device 100 in the second embodiment. This is a flowchart showing the processing of the information processing device 100 in the second embodiment. This is an explanatory diagram of scoring based on the position and orientation of the camera 200. This is an explanatory diagram of scoring based on the subject. This is a diagram showing how to select image data based on the position and orientation of the camera 200. This is a diagram showing how to select image data based on the subject. This is an explanatory diagram of how to prove that another person has implemented the second embodiment. This is an explanatory diagram of the reference camera 200S and non-reference camera 200N in the third embodiment. This is a diagram showing the configuration of the processing block of the information processing device 100 in the third embodiment. This is a flowchart showing the processing of the information processing device 100 in the third embodiment. This figure illustrates the effect of the third embodiment, which is the reduction of occlusion. This figure shows the configuration of a processing block in a modified example of the third embodiment. This flowchart shows the processing in a modified example of the third embodiment. This figure shows the configuration of a processing block of the information processing device 100 in the fourth embodiment. This flowchart shows the processing of the information processing device 100 in the fourth embodiment. This figure shows a model for calculating information regarding the quality of the reconstruction result. This figure illustrates the effect of the fourth embodiment. This figure shows the configuration of a modified example of the information processing device 100 in the fourth embodiment.

[0010] The embodiments of this technology will be described below with reference to the drawings. The description will be given in the following order: <First Embodiment> [About 4D Data] [Configuration of Information Processing System 10] [Configuration of Information Processing Device 100] [Processing in Information Processing Device 100] <Second Embodiment> [Processing in Information Processing Device 100] <Third Embodiment> [Configuration of Information Processing Device 100] [Processing in Information Processing Device 100] <Fourth Embodiment> [Configuration of Information Processing Device 100] [Processing in Information Processing Device 100] <Modification>

[0011] <First Embodiment> [About 4D Data] The 4D data in this technology will be explained with reference to Figure 1. The information processing device 100, which constitutes the information processing system 10 in this technology, generates 3D data or 4D data as a reconstruction result by reconstruction processing from multiple shooting data captured by multiple cameras 200 at different positions and orientations. Generating 3D data from two-dimensional image data or video data, which are the shooting data, is sometimes called 3D reconstruction, and generating 4D data by adding a time axis to the 3D data is sometimes called 4D reconstruction. In the first to fourth embodiments, the shooting data is image data captured by the camera 200, but the shooting data may also be video data (multiple frame images that constitute the video data).

[0012] In this technology, 4D data refers to data that includes information on a 3D spatial axis that moves along the time axis (including color, material, etc.). Specifically, this includes animated 3D mesh data, 3D point cloud data, polygon data, and NeRF (Neural Radiance Fields) and 3DGS (3D Gaussian Splatting) that are compatible with dynamic scenes.

[0013] The generated 4D data can be distributed to multiple devices, and the distributed 4D data is rendered on each device, ultimately resulting in a stereo video with both a time axis and a 2D axis.

[0014] It is possible to generate stereo video by pre-rendering 4D data and then deliver that stereo video to the viewer's device. In this case, viewing from a fixed viewpoint or, in the case of formats such as VR180, viewing in 3DoF (Degree of Freedom) is possible.

[0015] On the other hand, by distributing the data in 4D format rather than stereo, it becomes possible to perform real-time rendering on the viewer's device. In this case, viewing becomes possible with 6DoF (six degrees of freedom), adding "movement in forward / backward, up / down, and left / right" to "turning the head left / right and forward / backward to look around." This allows viewers to view the 4D data from any shooting viewpoint. At that time, the device estimates its own pose using SLAM (Simultaneous Localization and Mapping).

[0016] Viewing devices include HMDs (Head-Mounted Displays), glasses-type wearable devices, AR (Augmented Reality) devices such as smartphones, and spatial reproduction displays.

[0017] In the above explanation, the method of providing 4D data and stereo video was described as "distribution," but other methods of provision are also acceptable, such as saving the 4D data and stereo video to a storage medium and distributing that medium to viewers, or displaying the 4D data and stereo video in a screening.

[0018] Regarding standards for 4D data distribution, there is V-DMC (Video-based Dynamic Mesh Coding), an international standard for the compression and transmission of 3D meshes that change over time. There is also V-PCC (Video-based Point Cloud Compression), an international standard for the compression and transmission of 3D point clouds that change over time.

[0019] This technology can also be applied to volumetric capture. Volumetric capture is a technology that captures an entire space using more than 100 cameras surrounding it in a 360-degree circle, capturing the real space as three-dimensional digital data and reproducing it with high quality. The generated data can be converted into a 2D video viewed from any direction as a free-viewpoint representation, or into a 3D video that can be viewed on AR, stereoscopic monitors, HMDs, etc.

[0020] Furthermore, 4D data can also be displayed in a browser using technologies such as WebGL (Graphics Library) and WebXR, which enable 3D rendering and XR (Extended Reality) applications on a web browser.

[0021] [Configuration of Information Processing System 10] The configuration of the information processing system 10 will be described with reference to Figure 2. The information processing system 10 consists of an information processing device 100, multiple cameras 200, a terminal device 300, a first database 400, a second database 500, and a third database 600.

[0022] The information processing device 100 generates 3D data or 4D data as a reconstruction result by performing a reconstruction process on image data captured by multiple cameras 200 as input. The information processing device 100 is used by the person who wants to generate the reconstruction result. For example, if the subject captured by the cameras 200, i.e., the subject of the reconstruction process, is an event, the operator managing the event will use the information processing device 100 to generate the reconstruction result. The cameras 200 may capture images with their position and orientation fixed, or they may capture images while moving, i.e., while changing their position and orientation.

[0023] An event can be anything from large-scale events that attract many people, such as weddings, funerals, ceremonies, performances, recitals, concerts, festivals, and sporting events, to small-scale events like family and friends' gatherings. Furthermore, the reconstruction process is not limited to events; it can also be applied to animals, nature, landscapes, buildings, or anything else that can be photographed with Camera 200.

[0024] In this specification, the person who generates the reconstruction results using the information processing device 100, such as the event organizer described above, is referred to as the user. The user may provide the reconstruction results, and the person who receives the reconstruction results from the user and views them is referred to as the viewer. Furthermore, the person who films the event with the camera 200 is referred to as the photographer.

[0025] Assume that camera 200 includes cameras that have been calibrated and whose position, orientation, and camera parameters are known, and cameras that have not been calibrated and whose position, orientation, and camera parameters are unknown. When it is necessary to distinguish between multiple cameras 200, they will be referred to as camera 200A, camera 200B, camera 200C, camera 200D, and so on.

[0026] For example, a calibrated camera 200 is provided by the event organizer (user) of the information processing system 10 and used by the event staff. A non-calibrated camera 200 is used by event participants. Multiple cameras 200 may differ in manufacturer, model, or type, and their sensor characteristics, f-stop, white balance, exposure, and camera parameters may vary due to differences in how they are used by the photographers.

[0027] The information processing device 100, multiple cameras 200, terminal device 300, first database 400, second database 500, and third database 600 are equipped with communication functions and are connected by a network. The network connection method may be wired or wireless. Wired connection methods include, for example, HDMI (High-Definition Multimedia Interface) and USB (Universal Serial Bus), while wireless connection methods include, for example, Wi-Fi, Bluetooth (Registered Trademark), Wireless LAN (Local Area Network), NFC (Near Field Communication), 4G (Fourth Generation Mobile Communication System), 5G (Fifth Generation Mobile Communication System), and Ethernet (Registered Trademark).

[0028] Camera 200 is equipped with an image sensor, signal processing circuitry, etc., and can obtain color or monochrome video data and image data through shooting. CCD (Charge Coupled Device), CMOS (Complementary Metal Oxide Semiconductor), etc., can be used as the image sensor. Camera 200 can be any device with camera functionality, such as a smartphone, tablet, or wearable device. Due to differences in camera parameters, shooting settings, shooting position, and shooting orientation, there may be differences between the multiple image data from multiple cameras 200.

[0029] When taking a picture, the camera 200 acquires or generates various information related to the shooting and adds this information to the image data as metadata corresponding to the image data. The camera 200 may also acquire or generate metadata after shooting and add it to the image data. The camera 200 then transmits the image data and metadata to the information processing device 100. The transmission of the image data and metadata may be performed automatically when the camera 200's shooting mode ends, or it may be performed in response to a transmission instruction input from the photographer. The supply of image data and metadata from the camera 200 to the information processing device 100 may be performed via a recording medium such as a USB flash memory or SD memory card.

[0030] The reconstruction process requires position and orientation information from the camera 200 to generate the reconstruction results. Therefore, it is desirable that the camera 200 be equipped with a GPS (Global Positioning System) sensor for detecting position information, and an IMU (Inertial Measurement Unit) or inertial sensors (accelerometers, angular velocity sensors, gyroscopes for two or three axes) for detecting orientation information. If the camera 200 is equipped with a GPS sensor and an IMU, the camera 200 can add the position information detected by the GPS sensor and the orientation information detected by the IMU, etc., as metadata to the image data. The camera 200 may also be equipped with motion sensors such as an accelerometer, angular velocity sensor, and gyroscope, in which case the camera 200 can add metadata and motion detection information to the image data. Furthermore, the camera 200 may be equipped with distance sensors such as a LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging), a ToF (Time of Flight) sensor, a stereo camera, or a structured light camera.

[0031] Camera 200 may be equipped with known functions (such as subject detection, face detection, and scene detection) that can detect various types of information from image data obtained through shooting. If camera 200 is equipped with such information detection functions, camera 200 can add the detected information to the image data as metadata. These information detection functions can be implemented using machine learning or deep learning methods, template matching methods, matching methods based on the brightness distribution information of the subject, artificial intelligence methods, instance segmentation, and the like.

[0032] Figure 3 shows an example of the structure of a frame image that constitutes image data or video data as captured data. The image data consists of FS, Embedded Data Lines, Image Pixels, and Meta Data. General metadata such as shooting time, exposure time, and white balance are stored in Embedded Data Lines. In addition, metadata obtained through image analysis, such as scene information, is stored in Meta Data.

[0033] In addition to the configuration shown in Figure 3, metadata can also be added by embedding it at predetermined frame intervals (for example, embedding it at fixed frame intervals of N frames), or by embedding it only in frames where a subject or action specified by the user occurs or ends.

[0034] Metadata can include camera parameters, exposure time, white balance, shooting time, shooting date and time (season), shooting location (e.g., detection result of the GPS sensor on camera 200), orientation (e.g., detection result of the IMU on camera 200), motion information of camera 200 (detection results of motion sensors such as accelerometer, angular velocity sensor, and gyroscope), camera 200 manufacturer, camera identification information (e.g., IMEI (International Mobile Equipment Identifier) ​​number of camera 200), camera model information, photographer ID (e.g., login ID during upload), authenticity information (e.g., signature information), depth map information (e.g., PDAF (Phase Detection Auto Focus) data), polarization information, subject detection information, subject name, subject type, face detection information, subject part detection information, scene information, distance to subject information, focal length, illumination at the time of shooting, brightness of image data, distortion parameters, photographer information, etc. Metadata can be any information that can be obtained by camera 200, information that can be generated by camera 200, or information that camera 200 already holds.

[0035] Furthermore, the multiple cameras 200 may differ in their characteristics, camera parameters, signal processing, imaging settings, geometric conditions, etc.

[0036] Note that there may be only one camera 200, rather than multiple cameras. However, since multiple image data are required as input for the reconstruction process, if there is only one camera 200, that camera 200 must take pictures from multiple different positions and orientations to generate multiple image data.

[0037] The terminal device 300 is used by the user to input instructions regarding the reconstruction process to the information processing device 100, and to confirm and manage the reconstruction results. The terminal device 300 transmits the instructions input by the user as instruction information to the information processing device 100. The terminal device 300 can be an electronic device such as a personal computer, smartphone, or tablet terminal. Note that the terminal device 300 and the information processing device 100 may be composed of the same electronic device. In addition, the user may input conditions for image data selection, and confirm and manage the reconstruction results on the information processing device 100.

[0038] The first database 400 is a storage device that stores various types of data used by the information processing device 100 to generate reconstruction results. These types of data include the characteristics of the camera 200, the camera 200 model used for 4D data generation, and the reconstruction processing model.

[0039] Furthermore, the first database 400 can store image data, 2D data, 3D data, etc., which can be used to correct the reconstruction results by the information processing device 100, and can also provide them in response to requests from the information processing device 100.

[0040] The second database 500 is a storage device that stores the image data generated by the camera 200 through image capture.

[0041] The third database 600 is a storage device that stores the reconstruction results generated by the information processing device 100.

[0042] The first database 400, the second database 500, and the third database 600 are each constituted by a server, a cloud, a personal computer, or the like. Note that the first database 400, the second database 500, and the third database 600 may be constituted by the same device, server, cloud, or the like. Further, any one or all of the first database 400, the second database 500, and the third database 600 may be constituted by the same electronic device as the information processing apparatus 100 or the terminal device 300.

[0043] Only the first database 400 may be constituted by a cloud, and the information processing apparatus 100, the camera 200, the terminal device 300, the second database 500, and the third database 600 may be constituted locally. Further, the processing blocks of the information processing apparatus 100, the first database 400, the second database 500, and the third database 600 may be constituted by a cloud, and the terminal device 300, the hardware configuration of the information processing apparatus 100, and the camera 200 may be constituted locally.

[0044] [Configuration of Information Processing Apparatus 100] Next, referring to FIG. 4, the configuration of the processing blocks of the information processing apparatus 100 will be described.

[0045] The acquisition unit 101 acquires a plurality of image data with metadata added thereto, which are transmitted from a plurality of cameras 200, and outputs the data to the first selection unit 102. In the following description of the first to fourth embodiments, even when simply referred to as image data, it is assumed that metadata is added to the image data. [[ID=I1]]

[0046] The first selection unit 102 selects a plurality of image data to be used in subsequent processing from the plurality of image data.

[0047] The image quality adjustment unit 103 performs image quality adjustment processing on the plurality of image data selected by the first selection unit 102.

[0048] The position and orientation estimation unit 104 estimates the position and orientation of the camera 200 that captured the plurality of image data selected by the first selection unit 102.

[0049] The second selection unit 105 selects multiple image data to be used for reconstruction processing from the multiple image data selected by the first selection unit 102.

[0050] The reconstruction processing unit 106 generates 3D data or 4D data as a reconstruction result through reconstruction processing based on multiple image data.

[0051] The evaluation unit 107 evaluates the quality of the reconstruction result and the degree of influence of the multiple image data inputs to the reconstruction process on the reconstruction result.

[0052] The information processing device 100 may include an output unit that outputs the reconstruction results to an external device. The output unit outputs the reconstruction results via a network to a camera 200, a terminal device 300, a third database 600, or other external devices. Alternatively, the output unit may output the reconstruction results via a recording medium such as a USB flash memory or an SD memory card.

[0053] Next, the hardware configuration of the information processing device 100 will be described with reference to Figure 5.

[0054] The CPU (Central Processing Unit) 151 functions as an arithmetic processing unit that performs various processing tasks and controls the entire information processing device 100 and its individual parts. The CPU 151 executes various processes according to programs stored in the ROM (Read Only Memory) 152 or programs loaded from the storage unit 160 into the RAM (Random Access Memory) 153. The RAM 153 appropriately stores data necessary for the CPU 151 to execute various processes. Each processing block constituting the information processing device 100 can be realized by the processor, which is composed of the CPU 151, ROM 152, and RAM 153, executing a program.

[0055] The CPU 151, ROM 152, and RAM 153 are interconnected via a bus 154, which is connected to a bridge 155.

[0056] Interface 157 is connected to bridge 155 via bus 156.

[0057] Interface 157 is connected to an input unit 158, a display unit 159, a storage unit 160, a drive 161, a connection port 162, and a communication unit 163.

[0058] The input unit 158 ​​is, for example, various operators and operating devices such as a keyboard, mouse, keys, dial, touch panel, touchpad, or remote controller. The user's operation is detected by the input unit 158, and the signal corresponding to the input operation is interpreted by the CPU 151.

[0059] The display unit 159 is a liquid crystal display or an organic EL display that displays video, images, the GUI (Graphical User Interface) in this technology, messages, etc.

[0060] The storage unit 160 is a large-capacity storage medium such as a hard disk or flash memory. Various applications, data, and information are stored in the storage unit 160.

[0061] The information processing device 100 can be connected to a removable storage medium 164 via a drive 161. The removable storage medium 164 includes USB flash memory, SD memory cards, magnetic disks, optical disks, magneto-optical disks, semiconductor memory, etc. Image data can be supplied from the camera 200 to the information processing device 100 using the removable storage medium 164.

[0062] The drive 161 can read data files such as programs used for each process from the removable storage medium 164. The read data files are stored in the storage unit 160. In addition, programs and other data read from the removable storage medium 164 are installed in the storage unit 160 as needed. Furthermore, the information processing device 100 may transfer information and data to an external device via the removable storage medium 164.

[0063] External devices 165 can be connected to the information processing device 100 via the connection port 162.

[0064] The communication unit 163 includes various communication terminals and communication modules for communication processing via networks such as the Internet, and communication with various devices via wired / wireless communication, bus communication, etc. External devices can be connected to the information processing device 100 via the communication unit 163. The communication method can be either wired or wireless. Communication methods include cellular communication, 4G, 5G, Wi-Fi, Bluetooth®, NFC, Ethernet®, HDMI®, USB, etc. The information processing device 100 can receive image data transmitted from the camera 200 via communication through the communication unit 163. The information processing device 100 can also communicate with terminal devices 300 and databases 400 via communication through the communication unit 163.

[0065] Other external devices can be connected via connection port 162 or communication unit 163.

[0066] Note that the information processing device 100 does not need to have all the configurations shown in Figure 5. For example, if the information processing device 100 only processes data and outputs 4D data to the outside, the display unit 159 is not necessary.

[0067] The functions realized by the components described herein may be implemented in a circuit or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), CPUs, conventional circuits, and / or combinations thereof, programmed to realize the described functions. A processor, including transistors and other circuits, is considered a circuit or processing circuitry. A processor may be a programmed processor that executes a program stored in memory. In this specification, circuitry, unit, and means are hardware programmed to realize or execute the described functions. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to realize or execute the described functions. If such hardware is a processor that is considered a type of circuitry, then such circuitry, means, or unit is a combination of hardware and software used to constitute such hardware and / or processor.

[0068] In the information processing device 100, for example, programs and applications for processing this technology can be installed via network communication by the communication unit 163 or via a removable storage medium 164. Alternatively, programs and applications may be pre-stored in the ROM 152, storage unit 160, etc.

[0069] The information processing device 100 is configured as described above. The information processing device 100 may be configured as a standalone device, or it may be configured as an electronic device having information processing and communication functions, such as a personal computer, smartphone, or tablet terminal. Alternatively, the information processing device 100 and the information processing method may be realized by an electronic device having computer functions executing a program. The program may be pre-installed on the electronic device, or it may be distributed via download or storage media for the user to install.

[0070] The information processing device 100 may be configured on a cloud server. Furthermore, the information processing device 100 may transmit the generated reconstruction results to an external device or cloud server.

[0071] A cloud server is not limited to being composed of a single computer device; it may also be composed of a system of multiple computer devices. These multiple computer devices may be systematized, for example, by a LAN (Local Area Network). Alternatively, multiple computer devices located in remote locations may be systematized by a VPN (Virtual Private Network) using the internet, etc. These multiple computer devices may include computer devices that constitute a group of servers (cloud) available through cloud computing services.

[0072] [Processing in the Information Processing Device 100] The processing in the information processing device 100 will be described with reference to Figure 6. In this embodiment, as shown in Figure 7, multiple cameras 200 capture the subject that is the target of the reconstruction result from multiple different positions and orientations to generate multiple image data. The cameras 200 may capture images with their position and orientation fixed, or they may capture images while moving, that is, while changing their position and orientation. The information processing device 100 then generates the reconstruction result by performing a reconstruction process using the multiple image data captured by the cameras 200 as input.

[0073] In step S101, the acquisition unit 101 acquires multiple image data transmitted from multiple cameras 200 and outputs them to the first selection unit 102.

[0074] Next, in step S102, as shown in Figure 8, the first selection unit 102 selects multiple image data to be used for subsequent processing. Note that Figure 8 shows the first selection unit 102, the second selection unit 105, and the reconstruction processing unit 106 separately.

[0075] The first selection unit 102 selects image data in which the subject to be reconstructed is present. For example, based on subject detection information attached as metadata to the image data, the first selection unit 102 searches all image data for the subject that is the subject to the specified reconstructed process and selects image data in which the subject to be reconstructed is present.

[0076] The target of the reconstruction process may be specified by the user through input to the terminal device 300 or the information processing device 100, or it may be automatically determined by the information processing device 100 or other devices or algorithms.

[0077] One input for specifying the target of the reconstruction process is to display a list of multiple image data acquired by the acquisition unit 101 on the display unit of the terminal device 300 or the display unit 159 of the information processing device 100, and the user selects the image data in which the subject is the target of the reconstruction process from among the multiple image data, and then further specifies the subject to be the target of the reconstruction process from among the selected image data. The input for specifying the subject to be the target of the reconstruction process may be a tap input on the subject, a click input by hovering the cursor over the subject, or an input by drawing a frame around the subject; any input method is acceptable. In addition, the target of the reconstruction process may be specified by input such as photographs, text, or time information. For example, if multiple people exist as subjects in the image data, and one of them is designated as the target of the reconstruction process, the first selection unit 102 will select only the image data in which that one person exists as the subject.

[0078] Furthermore, the first selection unit 102 may select image data based on the subject detection result and conditions related to the subject. Conditions related to the subject include, for example, the type of subject, the subject's actions or state, etc. Conditions related to the subject may be specified using an image of the subject or by a text prompt.

[0079] Furthermore, if the captured data is video data containing time information, the conditions may include the specific playback time of the video data, or a specific time or time period based on the state of the subject. For example, conditions may include "from the time a specific subject appears in the field of view until a predetermined time has elapsed" or "from the time a specific action or state of a specific subject has elapsed until a predetermined time has elapsed." The first selection unit 102 may also select image data from a plurality of consecutive frame images that constitute the video data.

[0080] Furthermore, the first selection unit 102 may select image data based on conditions relating to the quality of the image data. Conditions relating to the quality of the image data include image quality, resolution, blur, illumination, noise, and the type of camera 200 used for shooting. For example, the first selection unit 102 may select image data whose image quality meets a predetermined standard. The first selection unit 102 may also exclude image data with resolution below a predetermined value, image data with blur, image data with noise levels above a predetermined value, image data with illumination below a predetermined value, etc., and select image data other than those mentioned above. This makes it possible to exclude image data that is of poor quality and is expected to cause deterioration in the quality of the reconstruction result. Furthermore, the first selection unit 102 may also select only image data taken with a camera 200 of a specific type (manufacturer, model, model name, year of manufacture, etc.). Image data taken with a specific type of camera 200 has less variation between image data, so using it as input for the reconstruction result can improve the quality of the reconstruction result.

[0081] Furthermore, the conditions for selecting image data are not limited to these; other conditions related to image data may be set in advance, or users may be allowed to specify their own conditions.

[0082] Next, in step S103, the image quality adjustment unit 103 performs image quality adjustment processing on the multiple image data selected by the first selection unit 102. The image data transmitted from multiple cameras 200 is adjusted in advance because the image quality of the image data differs due to individual differences in the cameras 200, differences in shooting settings, and differences in shooting position and orientation. Image quality adjustment may involve adjusting, for example, resolution, color tone, white balance, gradation, and brightness, but any parameter related to image quality may be adjusted. Image quality adjustment may be performed based on image data taken by a camera 200 that has taken a large number of images, or it may be performed based on image data where specific parameters such as brightness are above or below a predetermined standard. By adjusting the image quality of multiple image data taken by multiple cameras 200 and reducing the differences in image quality, the quality of the reconstruction result can be improved.

[0083] Next, in step S104, the position and orientation estimation unit 104 estimates the position and orientation of the camera 200 that captured the multiple image data selected by the first selection unit 102. The position and orientation of the camera 200 based on the image data can be estimated using a Visual Positioning System (VPS), Structure from Motion (SfM), etc. However, the position and orientation of the camera 200 may also be estimated using other methods. Note that if position and orientation information is attached as metadata to the image data, processing by the position and orientation estimation unit 104 is unnecessary.

[0084] Next, in step S105, the second selection unit 105 selects a plurality of image data to be used as input for the reconstruction process from the plurality of image data selected by the first selection unit 102, as shown in Figure 8.

[0085] The second selection unit 105 selects image data based on the position and orientation estimation results. For example, if there are two cameras 200A and camera 200B as shown in Figure 9, and the distance between the two cameras calculated from the position and orientation estimation results is less than or equal to a predetermined value, the second selection unit 105 may exclude image data taken by either camera. This is because multiple image data taken from close distances between cameras, i.e., from close positions, will be similar, and inputting multiple such image data will not contribute to improving the quality of the reconstruction result.

[0086] Furthermore, the second selection unit 105 may exclude image data from among the multiple image data if its similarity is above a predetermined value. This is because inputting multiple image data with high similarity does not contribute to improving the quality of the reconstruction result. Image similarity can be calculated using, for example, a method using Deep Learning, but other methods may also be used. The second selection unit 105 may have a function to calculate image similarity, or the information processing device 100 may have a dedicated processing unit for calculating image similarity.

[0087] Furthermore, the second selection unit 105 may select image data based on conditions related to the quality of the image data, such as image quality, resolution, blur, illumination, noise, and the type of camera 200 used for shooting.

[0088] By removing unnecessary images that do not contribute to improving the quality of the reconstruction results, the amount of image data used as input for the reconstruction process can be reduced, thereby shortening the time required for the reconstruction process.

[0089] Next, in step S106, the reconstruction processing unit 106 takes the multiple image data selected by the second selection unit 105 as input and performs reconstruction processing to generate 3D data or 4D data as reconstruction results. The reconstruction processing unit 106 can perform reconstruction processing using the position and orientation information of the camera 200 attached as metadata to the image data, or the position and orientation estimation result by the position and orientation estimation unit 104. The reconstruction processing can be performed, for example, by NeRF (Neural Radiance Fields) or 3DGS (3D Gaussian Splatting), but it may also be performed by other already known methods or by methods that will be realized in the future.

[0090] Next, in step S107, the evaluation unit 107 evaluates the quality of the reconstruction result. Evaluation methods and indicators include, for example, SNR (Signal to Noise Ratio), PSNR (Peak Signal to Noise Ratio), SSIM (Structual SIMilarity), LPIPS (Learned Perceptual Image Patch Similarity), F-score, and evaluation methods based on mesh quality, such as aspect ratio and skewness orthogonality.

[0091] If the evaluation of the reconstruction results meets the predetermined criteria, the process is terminated (Yes in step S108).

[0092] On the other hand, if the evaluation of the reconstruction results does not meet the predetermined criteria, the process proceeds from step S108 to step S109 (No. of step S108).

[0093] Next, in step S109, the evaluation unit 107 evaluates the degree of influence of the multiple image data inputs to the reconstruction process on the quality of the reconstruction result. The evaluation unit 107 can evaluate the degree of influence based on the change in the quality of the reconstruction result when any one of the multiple image data inputs to the reconstruction process is excluded and fine-tuned. The change in quality can be determined by whether the evaluation index (PSNR, SSIM, learning loss, CLIP (Contrastive Language Image Pretraining), SDS (Score Distillation Sampling), etc.) improves or deteriorates. By performing this evaluation for all image data inputs to the reconstruction process, it is possible to identify image data that has a negative impact on the quality of the reconstruction result. The evaluation unit 107 outputs information indicating the image data that has a negative impact on the quality of the reconstruction result to the second selection unit 105.

[0094] Next, in step S105, the second selection unit 105 selects image data to be used as input for the reconstruction process by excluding image data that has been evaluated by the evaluation unit 107 as having an adverse effect on the quality of the reconstruction result. Alternatively, the second selection unit 105 may select image data based on the position, size, area, or part of the subject within the field of view, which can be identified from the position and orientation of the camera 200 that captured the image data and the subject detection result.

[0095] Next, in step S106, the reconstruction processing unit 106 takes the multiple image data selected by the second selection unit 105 as input and performs the reconstruction process again to generate a new reconstruction result.

[0096] The information processing device 100 then repeatedly executes steps S105 to S109 until the quality evaluation of the reconstruction result meets a predetermined standard.

[0097] The processing of the first embodiment is carried out as described above. According to the first embodiment, by selecting image data, it is possible to perform the reconstruction process while excluding image data that is unnecessary for the reconstruction process or image data that adversely affects the quality of the reconstruction result, thereby improving the quality of the reconstruction result. In addition, by excluding such unnecessary image data, the time required for the reconstruction process can be shortened. Furthermore, by selecting image data based on conditions, it is possible to generate a reconstruction result that satisfies the conditions.

[0098] Furthermore, by performing image quality adjustments and estimation of the camera 200's position and orientation after selecting image data in the first selection unit 102, and then selecting image data again in the second selection unit 105, the processing of unnecessary image data is reduced, the processing load is lessened, and the processing time is shortened.

[0099] To verify whether another party has implemented the first embodiment, that is, whether they generated the reconstruction result using all image data captured by camera 200 as input, or whether they selected image data to generate the reconstruction result, this can be done, for example, as follows.

[0100] Two image datasets are prepared: a first image dataset and a second image dataset. The first image dataset consists of multiple image data, while the second image dataset contains the first image dataset plus lower-quality image data. By comparing the reconstruction results using the first image dataset as input with the reconstruction results using the second image dataset as input, if the quality is equivalent, it is highly likely that this technology was used to select and reconstruct the image data, thus verifying that another party has implemented this technique.

[0101] <Second Embodiment> [Configuration of Information Processing Device 100] Next, a second embodiment of the present technology will be described. The configuration of the information processing system is the same as in the first embodiment.

[0102] Referring to Figure 10, the configuration of the information processing device 100 in the second embodiment will be described. Except for the scoring unit 121, it is the same as in the first embodiment.

[0103] The scoring unit 121 scores the multiple image data acquired by the acquisition unit 101 according to predetermined criteria.

[0104] The first selection unit 102 selects image data to be used for processing based on the scoring results from the scoring unit 121.

[0105] [Processing in the Information Processing Device 100] The processing in the information processing device 100 will be described with reference to Figure 11.

[0106] In step S101, the acquisition unit 101 acquires multiple image data transmitted from multiple cameras 200 and outputs them to the scoring unit 121.

[0107] Next, in step S201, the scoring unit 121 scores the multiple image data. In this embodiment, the image data is scored based on three criteria: image data quality, the position and orientation of the camera 200, and the subject in the image data.

[0108] Image data quality can be scored using metrics such as SNR, PSNR, SSIM, or DNN (Deep Neural Network). Alternatively, scoring may be performed based on metadata attached to the image data, such as image resolution and the shooting settings of the camera 200 (e.g., HD (High Definition) mode).

[0109] The position and orientation of camera 200 are scored by comparing the ideal position and orientation of camera 200 with the position and orientation of camera 200 at the time of shooting. The smaller the difference between the ideal position and orientation of camera 200 and the position and orientation of camera 200 at the time of shooting, the higher the score. Therefore, as shown in Figure 12, if there is an ideal position and orientation of camera 200 and actual shooting images from cameras 200A and 200B, the score of the image data captured by camera 200A will be higher than the score of the image data captured by camera 200B.

[0110] The ideal position of the camera 200 may be set in advance, and the ideal position information may be input to the information processing device 100, or it may be updated by feedback from the first selection unit 102. The position and orientation of the camera 200 at the time of shooting may be acquired by the GPS sensor or IMU built into the camera 200 itself and added as metadata to the image data, or the camera 200 or the information processing device 100 may acquire it from the image data using VPS or SfM, etc.

[0111] The ideal position and the position of camera 200 at the time of shooting can be XY coordinates, distance information from a specific object, seating information in a facility, or any other information that can identify the location.

[0112] For subjects in image data, scoring is performed based on the subject detection results. The scoring unit 121 may perform scoring by referring to the subject detection results added to the image data as metadata, the scoring unit 121 may have a subject detection function, or the information processing device 100 may have a processing unit that performs subject detection.

[0113] The scoring unit 121 scores the image based on how much of the image field of view the subject occupies. For example, as shown in Figures 13A and 13B, the larger the area occupied by the subject within the field of view, the higher the score. Also, the score is higher when the entire subject is within the field of view, and lower when part of the subject is cut off or missing. Furthermore, if the subject to be reconstructed has been predetermined and that subject is present in the image data, the score will be higher.

[0114] Next, in step S202, the first selection unit 102 selects image data to be used for subsequent processing based on the scoring results.

[0115] The first selection unit 102 selects a predetermined number of image data from the highest scores for each of the image quality, field of view, and subject, or by selecting image data with a score equal to or greater than a predetermined value.

[0116] Furthermore, regarding the position and orientation of camera 200, image data captured at a different position and orientation from the multiple image data selected based on the score may be further selected. This selection can be made based on the position and orientation information of camera 200 attached as metadata to the image data.

[0117] For example, as shown in Figure 14, there are image data A obtained from camera 200A that photographs the subject from the front, image data B obtained from camera 200B that photographs the subject from the left, image data C obtained from camera 200C that photographs the subject from the right, and image data D obtained from camera 200D that photographs the subject from the front, and image data A and image data B are selected based on the score.

[0118] In this case, since there is a lack of image data taken of the subject from the right, the first selection unit 102 prioritizes image data C, which was taken of the subject from the right, in selecting image data. On the other hand, since image data A, which was taken of the subject from the front, has already been selected, the priority of image data D, which was also taken of the subject from the front, is lowered.

[0119] Furthermore, regarding the subject, additional image data containing missing parts of the subject may be selected from the multiple image data selected based on the score.

[0120] For example, as shown in Figures 15A, 15B, and 15C, if the lower half of the subject person is not visible in any of the multiple image data selected based on the score, the first selection unit 102 will give higher priority to the image data showing the lower half of the body, as shown in Figures 15D and 15E, when selecting image data. On the other hand, it will give lower priority to the image data not showing the lower half of the body, as shown in Figure 15F. In order to select image data based on the body parts of the subject in this way, detection information of the body parts of the subject is necessary.

[0121] Steps S103 to S109 are carried out in the same manner as in the first embodiment.

[0122] The processing of the second embodiment is carried out as described above. According to the second embodiment, by selecting image data with a high score, unnecessary images can be excluded, thereby reducing the number of image data used as input for the reconstruction process. This shortens the time required for the reconstruction process. Furthermore, by selecting image data with a high score, unnecessary images with a low score that degrade the quality of the reconstruction result can be excluded, thereby improving the quality of the reconstruction result. In addition, even image data captured by an unspecified number of cameras 200 or image data captured by a smartphone can be selected based on the scoring results and used as input for the reconstruction process, thereby improving the quality of the reconstruction result. This makes the reconstruction process easy to use.

[0123] The implementation of the second embodiment by another party can be proven, for example, by the method shown in Figure 16. First, reconstruction result A is generated using multiple image data captured by multiple cameras 200 as input. Next, the lens of one of the multiple cameras 200, which are in the same position and orientation, is blocked, and the same subject is photographed, and reconstruction result B is generated using multiple image data as input. Then, reconstruction result A and reconstruction result B are compared, and if the difference in the input image data does not affect the two reconstruction results, that is, if the two reconstruction results are identical, it can be said that the other party is selecting image data using the technology of the second embodiment. Note that the implementation by another party can be proven in the same way by changing the position and orientation of one of the cameras 200 instead of blocking the lens of the camera 200.

[0124] In the second embodiment, the selection by the second selection unit 105 may be omitted, and the image data selected by the first selection unit 102 may be used as input for reconstruction processing to generate the reconstruction result.

[0125] The scoring unit 121 is not always required to score based on the three criteria of image quality, the position and orientation of the camera 200, and the subject in the image data; it may score based on one or two of these criteria.

[0126] The first selection unit 102 is not always required to select image data based on three scoring results: image quality, position and orientation of the camera 200, and subject in the image data. It may select based on one or two of these scoring results.

[0127] If image data is selected based on the score and there is insufficient image data to be captured by camera 200 at a specific position and orientation, the GUI displayed on the display unit 159 may notify the user or photographer of the optimal position and orientation of camera 200.

[0128] Alternatively, image data may be selected based on the score, and the GUI displayed on the display unit 159 may notify the user or photographer of the position and orientation in which the image data and the image data constituting the parallax image can be captured.

[0129] <Third Embodiment> [Configuration of Information Processing System 10] Next, a third embodiment of the present technology will be described. As shown in Figure 17, in the third embodiment, the plurality of cameras 200 include a reference camera 200S and non-reference cameras 200N. The reference camera 200S is a camera whose position and orientation are fixed and whose position and orientation have been determined in advance by calibration. In addition to the position and orientation, it may also be a condition for the reference camera 200S that the shooting settings are known. There may be two or three reference cameras 200S, for example, but technically there is no limit to the number of reference cameras 200S. The non-reference cameras 200N are all cameras other than the reference camera 200S among the plurality of cameras 200 that constitute the information processing system 10. In the following description, image data captured by the reference camera 200S will be referred to as reference image data (reference shooting data), and image data captured by the non-reference camera 200N will be referred to as non-reference image data (non-reference shooting data). The other configurations of the information processing system are the same as in the first embodiment.

[0130] [Configuration of Information Processing Device 100] Referring to Figure 18, the configuration of the information processing device 100 in the third embodiment will be described.

[0131] The acquisition unit 101 acquires multiple reference image data transmitted from the reference camera 200S and outputs them to the reconstruction processing unit 106. The acquisition unit 101 also acquires multiple non-reference image data transmitted from the non-reference camera 200N and outputs them to the first selection unit 102.

[0132] The reconstruction processing unit 106 takes multiple reference image data as input and generates 3D data or 4D data as a first reconstruction result through reconstruction processing. After generating the first reconstruction result, the reconstruction processing unit 106 takes non-reference image data as additional input in addition to the reference image data and generates 3D data or 4D data as a second reconstruction result through reconstruction processing.

[0133] The evaluation unit 107 performs a quality evaluation of the second reconstruction result. The evaluation method is the same as in the first embodiment.

[0134] The first selection unit 102, image quality adjustment unit 103, position and orientation estimation unit 104, and second selection unit 105 perform the same processing as in the first embodiment, but the target of processing is non-reference image data.

[0135] [Processing in the Information Processing Device 100] Referring to Figure 19, the processing of the information processing device 100 in the third embodiment will be described.

[0136] In step S301, the acquisition unit 101 acquires multiple reference image data transmitted from multiple reference cameras 200S and outputs them to the reconstruction processing unit 106.

[0137] Next, in step S302, the reconstruction processing unit 106 takes multiple reference image data as input and performs reconstruction processing to generate a first reconstruction result. Since the position and orientation of the reference camera 200S are determined in advance, the reconstruction processing unit 106 can perform reconstruction processing using position and orientation information that has been added as metadata to the image data. The reconstruction processing unit 106 outputs the first reconstruction result to the position and orientation estimation unit 104 and the evaluation unit 107. The reconstruction processing unit 106 also retains the first reconstruction result for the generation of a second reconstruction result.

[0138] Next, in step S303, the acquisition unit 101 acquires multiple non-reference image data transmitted from multiple non-reference cameras 200N and outputs them to the first selection unit 102. The acquisition unit 101 may acquire the non-reference image data transmitted from the non-reference cameras 200N before the generation of the first reconstruction result, or it may acquire the non-reference image data after the generation of the first reconstruction result.

[0139] Furthermore, the reference camera 200S may add metadata to the reference image data indicating that the image data it transmits is reference image data, so that the information processing device 100 can distinguish between reference image data and non-reference image data by referring to this metadata. Alternatively, the information processing device 100 may store camera identification information of the reference camera 200S in advance, and distinguish between reference image data and non-reference image data by referring to the camera identification information contained in the metadata. In addition, image data acquired at a predetermined timing may be designated as reference image data, and image data acquired thereafter may be designated as non-reference image data. In this case, the reference camera 200S needs to transmit image data to the information processing device 100 synchronously and simultaneously or almost simultaneously before the non-reference camera 200N.

[0140] Next, in step S304, the first selection unit 102 selects non-reference image data to be used for subsequent processing from a plurality of non-reference image data. The selection process is the same as in the first embodiment. In the third embodiment, the first selection unit 102 may make a selection based on the image quality of the reference image data. For example, the image quality of the reference image data and the non-reference image data is evaluated using evaluation criteria such as SNR, PSNR, and SSIM, and non-reference image data is selected in which there is no difference in the evaluation result or the difference is slight (the difference is less than or equal to a predetermined value).

[0141] Next, in step S305, the image quality adjustment unit 103 performs image quality adjustment processing on a plurality of non-reference image data selected by the first selection unit 102, using the reference image data as a reference. Since the non-reference image data differs in brightness, resolution, color tone, white balance, etc., due to individual differences in the camera 200 and differences in shooting settings, the image quality of the non-reference image data is adjusted in advance using the reference image data as a reference.

[0142] By adjusting the image quality of non-reference image data based on reference image data, it is possible to adjust characteristics that depend on the type and performance of the camera 200 and improve the quality of the reconstruction results. Furthermore, through image quality adjustment, reconstruction results can be generated even when image data is captured by an unspecified number of cameras 200 or a wide variety of cameras 200, as input for the reconstruction process.

[0143] Next, in step S306, the position and orientation estimation unit 104 estimates the position and orientation of each non-reference camera 200N that captured the multiple non-reference image data selected by the first selection unit 102, based on the first reconstruction result. The estimation of the position and orientation of the non-reference cameras 200N based on the first reconstruction result can be performed using VPS, SfM, or other methods.

[0144] Next, in step S307, the second selection unit 105 selects non-reference image data from a plurality of non-reference image data to be used as input for the reconstruction process. The selection method is the same as in the first embodiment.

[0145] Next, in step S308, the reconstruction processing unit 106 performs reconstruction processing using a plurality of non-reference image data selected by the second selection unit 105 as additional inputs, in addition to the reference image data, to generate a second reconstruction result. Note that the first reconstruction result and the second reconstruction result may be different types of data. For example, the first reconstruction result can be 3D point cloud data and the second reconstruction result can be 4D data.

[0146] Next, in step S309, the evaluation unit 107 evaluates the quality of the second reconstruction result. The evaluation can be performed in the same manner as in the first embodiment, or it may be performed based on reference image data using PSNR or the like.

[0147] If the evaluation of the reconstruction results meets the predetermined criteria, the process is terminated (Yes in step S310).

[0148] On the other hand, if the evaluation of the reconstruction results does not meet the predetermined criteria, the process proceeds from step S310 to step S311 (step S310 No.).

[0149] Next, in step S311, the evaluation unit 107 evaluates the degree of influence of the multiple non-reference image data inputs to the reconstruction process on the quality of the reconstruction result. The evaluation method is the same as in the first embodiment. The evaluation unit 107 outputs information to the second selection unit 105 indicating the non-reference image data that has a negative impact on the reconstruction result.

[0150] Next, in step S307, the second selection unit 105 selects non-reference image data to be used as additional input for the reconstruction process by excluding non-reference image data that the evaluation unit 107 has evaluated as having an adverse effect on the reconstruction result.

[0151] Next, in step S308, the reconstruction processing unit 106 takes a plurality of non-reference image data selected by the second selection unit 105 as additional inputs in addition to the reference image data and performs the reconstruction process again to generate a second reconstruction result.

[0152] Then, steps S307 to S311 are repeatedly executed until the quality evaluation of the reconstruction results meets a predetermined standard.

[0153] The processing of the third embodiment is carried out as described above. According to the third embodiment, by estimating the position and orientation of the non-reference camera 200N based on the first reconstruction result generated using reference image data as input, the reconstruction process can be robustly carried out even when the subject or the non-reference camera 200N is moving, by using non-reference image data as additional input. Both the subject and the non-reference camera 200N can move freely, and even non-reference image data taken by the non-reference camera 200N, whose shooting conditions such as position and field of view are unknown, can be selected to contribute to improving the quality of the reconstruction result and used as input for the reconstruction process.

[0154] There is a phenomenon called occlusion, where an object in the foreground obscures the area behind it. For example, as shown in Figure 20A, when an object OBJ is photographed from the front with the reference camera 200S, occlusion (OCC) occurs behind the object. If only the reference image data captured by the reference camera 200S is used as input in this way, a low-quality reconstruction result will be generated. In the third embodiment, based on the position and orientation of the non-reference camera 200N, as shown in Figure 20B, the reference camera 200S can select non-reference image data captured by the non-reference camera 200N that can capture the area where occlusion (OCC) occurs. Then, by performing reconstruction processing with this non-reference image data as additional input, a high-quality reconstruction result can be generated.

[0155] According to the third embodiment, the resolution of the reconstruction result can be improved by adding non-reference image data, and viewpoint-dependent components can be learned.

[0156] According to the third embodiment, for example, when filming an event such as a live concert, first, a reconstruction result is generated using reference image data captured by a professional camera or similar device, which is set up in advance as a reference camera 200S, as input. Then, a reconstruction result is generated again using non-reference image data captured by an audience member, such as a smartphone, which is a non-reference camera 200N, as additional input. This improves the quality of the reconstruction result.

[0157] The implementation of the third embodiment by another party can be verified, for example, by checking the configuration of an application or device that adds newly captured image data (such as from a smartphone) to a reconstruction result generated using image data captured by a fixed camera as input, and by confirming whether or not the reconstruction result is updated due to the additional input of image data.

[0158] In the third embodiment, the reference camera 200S is not limited to one whose position and orientation are fixed in advance and whose position and orientation are determined by calibration. The reference camera 200S may be selected from among a plurality of cameras 200 whose position and orientation are not determined.

[0159] When determining a reference camera 200S from among multiple cameras 200 whose position and orientation are not specified, the information processing device 100 includes a reference camera determination unit 131, as shown in Figure 21. The process when the reference camera determination unit 131 is included will be explained with reference to Figure 22.

[0160] In step S321, the acquisition unit 101 acquires multiple image data transmitted from multiple cameras 200 and outputs them to the first selection unit 102.

[0161] Next, in step S322, the first selection unit 102 selects image data from a plurality of image data to be used for subsequent processing. The selection process is the same as in the first embodiment. The first selection unit 102 outputs the selected image data to the reference camera determination unit 131 and the image quality adjustment unit 103.

[0162] Next, in step S323, the reference camera determination unit 131 determines the reference camera 200S from among the multiple cameras 200 that have captured the image data selected by the first selection unit 102.

[0163] The reference camera determination unit 131 may determine the camera with the most frequent model among the multiple cameras 200 as the reference camera 200S. The most recent camera model can be identified by camera model information as metadata attached to the image data. Identifying the most frequent model is not limited to completely identical models; the manufacturer, model, etc., may also be the same.

[0164] Furthermore, the reference camera determination unit 131 may determine the camera with the least movement among the multiple cameras 200 as the reference camera 200S. Whether or not the camera 200 has little movement can be determined, for example, based on the detection result of a motion sensor as metadata attached to the image data.

[0165] The reference camera determination unit 131 may determine the most frequently used camera 200 as the reference camera 200S, or it may determine the camera 200 with minimal movement as the reference camera 200S, or it may determine the most frequently used camera 200 with minimal movement as the reference camera 200S. Alternatively, the reference camera 200S may be determined based on criteria other than the most frequently used camera and movement. The reference camera determination unit 131 outputs camera identification information indicating that it is the reference camera 200S to the image quality adjustment unit 103 and the evaluation unit 107.

[0166] Next, in step S324, the image quality adjustment unit 103 performs image quality adjustment processing on the image data selected by the first selection unit 102, using the reference image data captured by the reference camera 200S as a reference.

[0167] Next, in step S325, the position and orientation estimation unit 104 estimates the position and orientation of each camera 200 that captured the image data selected by the first selection unit 102.

[0168] Next, in step S326, the second selection unit 105 selects image data from the image data selected by the first selection unit 102 to be used as input for the reconstruction process.

[0169] Next, in step S327, the reconstruction processing unit 106 takes the image data selected by the second selection unit 105 as input and performs a reconstruction process to generate a reconstruction result.

[0170] Steps S309 to S311 are the same as the process described above.

[0171] The process for determining a reference camera 200S from among multiple cameras 200 is carried out as described above. With this method, even if there is no reference camera 200S whose position and orientation have been determined in advance, a reference camera 200S can be determined from among multiple cameras 200, and the same effect as in the third embodiment can be obtained.

[0172] Furthermore, when determining a reference camera 200S from among multiple cameras 200, a first reconstruction result may be generated using reference image data captured by the determined reference camera 200S as input, and a second reconstruction result may be generated using non-reference image data captured by non-reference cameras 200N other than the reference camera 200S as additional input.

[0173] Furthermore, if static subjects make up the majority of the image data, instead of setting up a reference camera 200S whose position and orientation have been determined in advance, 3D data measurement of the subjects may be performed beforehand, and the position and orientation estimation, additional input for reconstruction processing, and evaluation of the quality of the reconstruction results may be performed using the 3D data measurement results.

[0174] For example, in live performances, plays, or presentations where multiple people are on stage, the subjects or areas of focus may differ depending on the photographer. According to the third embodiment, in such cases, the reconstruction process can be performed by adding non-reference image data, which the photographer is focusing on, captured by the non-reference camera 200N, to the reconstruction result generated using the reference image data captured by the reference camera 200S as input. In this case, it is desirable that the reference camera 200S captures the entire stage, but it is not essential that it captures the entire stage. An essential requirement is that there is an overlap between the reference image and the non-reference image; that is, each image must capture a common part with at least one other image.

[0175] <Fourth Embodiment> [Configuration of Information Processing Device 100] Next, a fourth embodiment of the present technology will be described. The configuration of the information processing system in the fourth embodiment is the same as in the first embodiment.

[0176] Referring to Figure 23, the configuration of the information processing device 100 in the fourth embodiment will be described.

[0177] The acquisition unit 101 acquires multiple image data transmitted from multiple cameras 200 and outputs them to the position and orientation estimation unit 104.

[0178] The position and orientation estimation unit 104 estimates the position and orientation of the camera 200 that has captured multiple image data.

[0179] The subject detection unit 141 detects subjects in multiple image data.

[0180] The third selection unit 142 selects a plurality of image data to be used to calculate information regarding the quality of the reconstruction result by the information calculation unit 143, based on the detection result of the subject detection unit 141.

[0181] The information calculation unit 143 calculates information regarding the quality of the reconstruction result based on the multiple image data selected by the third selection unit 142. The information regarding the quality of the reconstruction result includes either or both the contribution of each image data to the reconstruction result and the estimated quality evaluation value of the reconstruction result. The contribution of each image data to the reconstruction result is a numerical representation of how much each image data influences the quality of the reconstruction result. The contribution is calculated for each image data. The estimated quality evaluation value of the reconstruction result is a numerical representation of the estimated quality of the reconstruction result, which has not actually been generated.

[0182] [Processing in the Information Processing Device 100] The processing of the information processing device 100 in the fourth embodiment will be described with reference to Figure 24.

[0183] In step S401, the acquisition unit 101 acquires multiple image data transmitted from multiple cameras 200 and outputs them to the position and orientation estimation unit 104.

[0184] Next, in step S402, the position and orientation estimation unit 104 estimates the position and orientation of each camera 200 that captured multiple image data. Note that if position and orientation information is added as metadata to the image data, processing by the position and orientation estimation unit 104 is unnecessary.

[0185] Next, in step S403, the subject detection unit 141 performs instance segmentation using a CNN, such as MaskR-CNN (Mask Regions with Convolutional Neural Network), for each of the multiple image data to detect individual subjects (instances) contained in the image data. Instance segmentation uses deep learning to detect the boundaries and contours of individual subjects in the image data at the pixel level with more detail than conventional subject detection algorithms.

[0186] The subject detection unit 141 may perform subject detection using semantic segmentation instead of, or in combination with, instance segmentation, or it may perform subject detection using panoptic segmentation, which is a combination of instance segmentation and semantic segmentation. The subject detection unit 141 may also use known subject detection techniques other than segmentation.

[0187] Next, in step S404, the subject detection unit 141 associates an instance ID with the subject detected by instance segmentation. The instance ID is added to the image data as metadata.

[0188] When image data contains multiple subjects, all subjects may be subjected to reconstruction processing, or one or more subjects may be subjected to reconstruction processing. The degree of contribution of the image data to the reconstruction result and the estimated quality evaluation value of the reconstruction result change depending on which subjects are subjected to reconstruction processing. Therefore, before the information calculation unit 143 calculates the degree of contribution of the image data to the reconstruction result and the estimated quality evaluation value of the reconstruction result, it is necessary to detect the subjects in the image data using the subject detection unit 141.

[0189] Next, in step S405, the third selection unit 142 selects one image data from among the multiple image data for which instance segmentation has been performed by the subject detection unit 141, in which the object to be reconstructed exists as the subject.

[0190] Next, in step S406, the third selection unit 142 selects the target for reconstruction processing in one of the selected image data. The target for reconstruction processing may be specified by the user through input to the terminal device 300 or the information processing device 100, or it may be automatically determined by the information processing device 100, other devices, or algorithms. The input for specifying the target for reconstruction processing can be made in the same way as described in the first embodiment.

[0191] Next, in step S407, the third selection unit 142 searches all image data for the target of the selected reconstruction process by referring to the instance ID.

[0192] Next, in step S408, the third selection unit 142 collects multiple image data, including the target of the reconstruction process, to create an image dataset.

[0193] Next, in step S409, the third selection unit 142 generates a mask containing only the elements to be reconstructed.

[0194] The third selection unit 142 may also make selections based on the quality of the image data. Image data quality includes resolution, blur, and the type of camera 200 used for shooting. For example, the third selection unit 142 may construct an image dataset by excluding image data with a resolution below a predetermined value or image data that contains blur.

[0195] Next, in step S410, the information calculation unit 143 uses an information calculation model, as shown in Figure 25, to calculate information regarding the quality of the reconstruction result for each image data contained in the image dataset, taking the image dataset, the position and orientation information of the camera 200 attached as metadata to the image data contained in the image dataset, and the mask as input. The information regarding the quality of the reconstruction result consists of the contribution of the image data to the reconstruction result and the estimated quality evaluation value of the reconstruction result. The information calculation unit 143 may calculate both, or only one of them. The user may be allowed to select which one to calculate. A mask can be used, for example, to remove the background when an object is the target of the reconstruction process and the background of that object is not needed in the reconstruction process. Note that the use of a mask is not mandatory, so a mask does not need to be input.

[0196] Note that the configuration of the information calculation model shown in Figure 25 is just one example; any model or algorithm can be used as long as it can calculate information regarding the quality of the reconstruction results.

[0197] Next, in step S411, the evaluation unit 107 determines whether the estimated quality evaluation value of the reconstruction result is equal to or greater than a predetermined value.

[0198] If the estimated quality evaluation value of the reconstruction result is equal to or greater than a predetermined value, the process is terminated (Yes in step S412).

[0199] On the other hand, if the estimated quality evaluation value of the reconstruction result is less than or equal to a predetermined value, the process proceeds from step S412 to step S413 (No. of step S412).

[0200] Next, in step S413, the third selection unit 142 re-selects the image data to be input to the information calculation unit 143. For example, by excluding image data whose contribution is below a predetermined value, the image dataset can be reconstructed using only image data with a high contribution. This makes it possible to improve the quality estimation evaluation value of the reconstruction result.

[0201] Furthermore, the third selection unit 142 may exclude image data where the distance from the camera 200 to the object to be reconstructed is greater than or equal to a predetermined value. Also, the third selection unit 142 may exclude image data where the distance (pixels) between the center of the camera 200 and the center of the object to be reconstructed is greater than or equal to a predetermined value.

[0202] The processing of the fourth embodiment is carried out as described above. According to the fourth embodiment, both the subject and the camera 200 can move freely, and information regarding the quality of the reconstruction result can be calculated even if multiple image data are captured by multiple cameras 200 with different shooting conditions. This makes it possible to select image data that improves the quality of the reconstruction result without actually generating the reconstruction result. By performing the reconstruction process with image data that has a high contribution as input, it is possible to generate a high-quality reconstruction result. Since it is not necessary to generate the reconstruction result in order to calculate the contribution of the image data to the reconstruction result, the contribution can be calculated faster than when it is calculated after generating the reconstruction result. Because the contribution can be calculated at high speed, it is also possible to feed back the contribution in real time to the image data selection process in the third selection unit 142. By acquiring information about the subject in the image data, it is also possible to handle cases where the contribution changes when the target of the reconstruction process is different, even if the image data is the same.

[0203] By selecting image data to be used as input for the reconstruction process based on its contribution to the reconstruction result, it is possible to exclude image data with a low contribution that would degrade the quality of the reconstruction result. This allows for an improvement in the quality of the reconstruction result without the need to re-capture and add additional image data.

[0204] By pre-detecting the subject in the image data and selecting the image data according to the subject to be reconstructed, it is possible to handle cases where the same image data is used but the target of the reconstruction process differs. For example, as shown in Figure 26, suppose there is image data in which animal A and person H are subjects, and animal A is in focus but person H is out of focus. If animal A is the target of the reconstruction process, since animal A is in focus, the contribution of that image data will be high, and the estimated quality evaluation value of the reconstruction result will also be high. On the other hand, if person H is the target of the reconstruction process, since person H is out of focus, the contribution of that image data will be low, and the estimated quality evaluation value of the reconstruction result will also be low. In this way, even with the same image data, the contribution and the estimated quality evaluation value of the reconstruction result change depending on what is being reconstructed, but in this embodiment, it is possible to calculate the contribution of the image data according to the target of the reconstruction process and the estimated quality evaluation value of the reconstruction result.

[0205] In the fourth embodiment, it is not necessary to generate reconstruction results in order to calculate information regarding the quality of the reconstruction results. However, as shown in Figure 27, even in the fourth embodiment, the information processing device 100 may also include a reconstruction processing unit 106 to generate reconstruction results. In that case, high-quality reconstruction results can be generated by inputting only image data in which the contribution is greater than or equal to a predetermined value.

[0206] Furthermore, the information processing device 100 may execute the reconstruction process and generate the reconstruction result only when a user who has confirmed the contribution or the estimated quality of the reconstruction result instructs the execution of the reconstruction process.

[0207] Furthermore, when a reconstruction process is performed and a reconstruction result is generated, the reconstruction result and the image data input to the information calculation unit 143 may be evaluated using SNR, PSNR, SSIM, LPIPS, etc., and the third selection unit 142 may re-select the image data to be input to the information calculation unit 143 and the reconstruction process based on the difference between these evaluations.

[0208] The implementation of the fourth embodiment by another party can be verified in the following way. First, two types of image datasets are prepared: image dataset A and image dataset B. Image dataset B is obtained by adding low-quality image data to image dataset A. Low quality means, for example, images that are blurry, out of focus, or not suitable for reconstruction processing.

[0209] Assume there is a processing unit that takes a dataset as input and outputs a numerical value for each data point, and another processing unit that performs data removal or other operations based on the numerical values. Then, the difference between the output dataset when image dataset A is input to the processing unit and the output dataset when image dataset B is input to the processing unit is calculated. If there is no difference in the difference, or if the difference in the difference is small (less than a predetermined amount), it can be proven that the fourth embodiment has been implemented. Furthermore, if the processing unit performs reconstruction processing using the output dataset, the implementation of the fourth embodiment can be similarly proven based on the difference in the quality of the reconstruction results.

[0210] <Modifications> Although embodiments of this technology have been described in detail above, this technology is not limited to the embodiments described above, and various modifications are possible based on the technical concept of this technology.

[0211] The data used as input for the reconstruction process is not limited to captured images; it can also include image data or video data (multiple frame images that make up a video) generated by image generation AI or image generation tools.

[0212] If the background is unnecessary in the captured data, background removal processing can be applied to extract only the specific subject. This can improve the quality of the reconstruction result.

[0213] The information processing device 100 may include a subject detection unit and perform subject detection on the captured data. Subject detection can be performed using methods such as machine learning or deep learning, template matching, matching based on the brightness distribution information of the subject, artificial intelligence, or instance segmentation.

[0214] This technology can be implemented by combining any two or more of the first to fourth embodiments.

[0215] In the embodiment, it was explained that a user, such as an event organizer, uses the information processing device 100 to generate the reconstruction results. However, an individual who is not an organizer may also use the information processing device 100 to generate the reconstruction results. Furthermore, the user and the photographer may be the same person, the user and the viewer may be the same person, or the user, photographer, and viewer may be the same person. In addition, anyone can use the information processing device 100 of this technology, and the use and purpose of the information processing device 100 are not limited to events as described above, but may be used for any purpose.

[0216] The technology can also take the following configurations: (1) An information processing device comprising: a first selection unit that selects from a plurality of captured data captured by a camera; and a second selection unit that selects from the plurality of captured data selected by the first selection unit to be used as input for reconstruction processing. (2) The information processing device according to (1), wherein the first selection unit makes a selection based on the subject of the captured data that is the target of reconstruction processing. (3) The information processing device according to (1) or (2), wherein the first selection unit makes a selection based on the quality of the captured data. (4) The information processing device according to any one of (1) to (3), further comprising: a position and orientation estimation unit that estimates the position and orientation of the camera that captured the captured data selected by the first selection unit; and the second selection unit making a selection based on the estimation result by the position and orientation unit. (5) The information processing device according to any one of (1) to (4), further comprising: a reconstruction processing unit that takes the captured data selected by the second selection unit as input to perform the reconstruction processing and generates a reconstruction result. (6) The information processing device according to any one of (1) to (5), further comprising: an evaluation unit that evaluates the quality of the reconstruction result generated by the reconstruction processing. (7) An information processing device according to any one of (1) to (6), comprising a scoring unit for scoring the shooting data. (8) An information processing device according to (7), wherein the scoring unit scores based on the quality of the shooting data. (9) An information processing device according to (7), wherein the scoring unit scores based on at least one of the position and orientation of the camera that took the shooting data. (10) An information processing device according to (7), wherein the scoring unit scores based on the subject in the shooting data. (11) An information processing device according to (7), wherein the first selection unit makes a selection based on the scoring result by the scoring unit. (12) An information processing device according to (7), wherein the first selection unit selects the shooting data based on the position and orientation of the camera in the shooting data. (13) An information processing device according to (7), wherein the first selection unit selects the shooting data based on the subject in the shooting data.(14) The information processing apparatus according to (5), wherein the reconstruction processing unit performs reconstruction processing using reference image data captured by a reference camera as input. (15) The information processing apparatus according to (14), wherein the reference camera is a camera whose position and orientation have been predetermined from among a plurality of cameras. (16) The information processing apparatus according to (14), wherein the reconstruction processing unit performs reconstruction processing using non-reference image data captured by a non-reference camera, which is a camera other than the reference camera, as input, in addition to the reference image data. (17) The information processing apparatus according to (16), wherein the reconstruction processing unit performs reconstruction processing using non-reference image data captured by the non-reference camera and selected by the first selection unit and the second selection unit as input. (18) The information processing apparatus according to (17), further comprising a position and orientation estimation unit that estimates the position and orientation of the non-reference camera that captured the non-reference image data based on the reconstruction result generated by the reconstruction processing using the reference image data as input, and the second selection unit makes a selection based on the position and orientation estimation result by the position and orientation estimation unit. (19) An information processing method for selecting from multiple shooting data captured by a camera, and selecting the shooting data from the selected multiple shooting data to be used as input for a reconstruction process. (20) A program for causing a computer to execute the information processing method for selecting from multiple shooting data captured by a camera, and selecting the shooting data from the selected multiple shooting data to be used as input for a reconstruction process.

[0217] 100... Information processing unit 102... First selection unit 104... Position and orientation estimation unit 105... Second selection unit 106... Reconstruction processing unit 107... Evaluation unit 200... Camera 121... Scoring unit 200S... Reference camera 200N... Non-reference camera

Claims

1. An information processing device comprising: a first selection unit that selects from a plurality of captured data captured by a camera; and a second selection unit that selects from the plurality of captured data selected by the first selection unit to be used as input for reconstruction processing.

2. The information processing apparatus according to claim 1, wherein the first selection unit makes a selection based on the subject subject in the captured data that is the subject of reconstruction processing.

3. The information processing apparatus according to claim 1, wherein the first selection unit performs a selection based on the quality of the captured data.

4. The information processing apparatus according to claim 1, further comprising a position and orientation estimation unit that estimates the position and orientation of the camera that captured the shooting data selected by the first selection unit, wherein the second selection unit makes a selection based on the estimation result by the position and orientation unit.

5. The information processing apparatus according to claim 1, further comprising a reconstruction processing unit that performs the reconstruction process using the image data selected by the second selection unit as input and generates a reconstruction result.

6. The information processing apparatus according to claim 1, further comprising an evaluation unit for evaluating the quality of the reconstruction results generated by the reconstruction process.

7. The information processing apparatus according to claim 1, further comprising a scoring unit for scoring the aforementioned shooting data.

8. The information processing apparatus according to claim 7, wherein the scoring unit performs scoring based on the quality of the captured data.

9. The information processing apparatus according to claim 7, wherein the scoring unit performs scoring based on at least one of the position and orientation of the camera that captured the shooting data.

10. The information processing apparatus according to claim 7, wherein the scoring unit performs scoring based on the subject in the captured data.

11. The information processing apparatus according to claim 7, wherein the first selection unit makes a selection based on the scoring result by the scoring unit.

12. The information processing apparatus according to claim 7, wherein the first selection unit selects the shooting data based on the position and orientation of the camera in the shooting data.

13. The information processing apparatus according to claim 7, wherein the first selection unit selects the shooting data based on the subject in the shooting data.

14. The information processing apparatus according to claim 5, wherein the reconstruction processing unit performs reconstruction processing using reference image data captured by a reference camera as input.

15. The information processing apparatus according to claim 14, wherein the reference camera is a camera whose position and orientation have been predetermined from among a plurality of cameras.

16. The information processing apparatus according to claim 14, wherein the reconstruction processing unit performs reconstruction processing using, in addition to the reference imaging data, non-reference imaging data captured by a non-reference camera, which is a camera other than the reference camera, as input.

17. The information processing apparatus according to claim 16, wherein the reconstruction processing unit performs reconstruction processing using the non-reference image data captured by the non-reference camera and selected by the first selection unit and the second selection unit as input.

18. The information processing apparatus according to claim 17, comprising a position and orientation estimation unit that estimates the position and orientation of the non-reference camera that captured the non-reference shooting data based on the reconstruction result generated by a reconstruction process using the reference shooting data as input, wherein the second selection unit selects based on the position and orientation estimation result by the position and orientation estimation unit.

19. An information processing method that selects from multiple image data captured by a camera, and selects from the selected multiple image data to be used as input for a reconstruction process.

20. A program that causes a computer to execute an information processing method which involves selecting from multiple image data captured by a camera, and then selecting from the selected multiple image data to be used as input for a reconstruction process.