Information processing system and information processing method
Patent Information
- Application Number
- PCT/JP2026/003767
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-02-03
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026003767_01102026_PF_FP_ABST
Abstract
Description
INFORMATION PROCESSING SYSTEM AND INFORMATION PROCESSING METHOD
[0001] The present technology relates to an information processing system and an information processing method applicable to generation of viewpoint images.
[0002] Patent Document 1 discloses a learning device used for generating a 3D model. In this learning device, only a portion obtained by removing an area located farther than a predetermined background point from each captured image is used for generating the 3D model. This eliminates the need for background removal processing from the captured image.
[0003] Japanese Patent No. 7360757
[0004] In such a technical field, there is a demand for a technology that enables generation of high-quality viewpoint images based on a 3D model.
[0005] In view of the above circumstances, an object of the present technology is to provide an information processing system and an information processing method that enable generation of high-quality viewpoint images.
[0006] To achieve the above object, an information processing system according to one aspect of the present technology includes an image group acquisition unit, a correction unit, a model generation unit, and a viewpoint image generation unit. The image group acquisition unit acquires an image group including a plurality of captured images. The correction unit acquires a plurality of reference styles, and for each of the plurality of reference styles, corrects the style of the image group to the reference style, thereby generating a plurality of corrected image groups. The model generation unit generates a plurality of viewpoint image generation models based on each of the plurality of corrected image groups. The viewpoint image generation unit selects a part of the plurality of viewpoint image generation models based on a predetermined viewpoint, and generates a viewpoint image at the predetermined viewpoint based on the predetermined viewpoint and the selected viewpoint image generation model.
[0007] The correction unit may acquire the plurality of reference styles based on the plurality of captured images.
[0008] The information processing system may further include a reference image selection unit that selects a predetermined number of reference images from the image group. In this case, the correction unit may acquire a predetermined number of styles of each of the reference images selected by the reference image selection unit as the reference styles.
[0009] The reference image selection unit may select all of the captured images included in the image group as the reference image.
[0010] The reference image selection unit may select the captured image specified by the user from among the captured images included in the image group as the reference image.
[0011] The reference image selection unit may select the reference image based on the brightness distribution of the captured images included in the image group.
[0012] The reference image selection unit may select, from among the captured images included in the image group, an image that includes an object in real space as the reference image.
[0013] The reference image selection unit may determine whether or not the object is included in the captured image by using object detection.
[0014] The viewpoint image generation unit may select the viewpoint image generation model relating to the reference image whose viewpoint at the time of imaging is closest to the predetermined viewpoint.
[0015] The image acquisition unit may classify the image group into a plurality of clusters such that each cluster contains one reference image. In this case, the viewpoint image generation unit may select the viewpoint image generation model relating to the reference image of the cluster containing the captured image whose viewpoint at the time of acquisition is closest to the predetermined viewpoint.
[0016] The correction unit may acquire a reference exposure amount as the reference style.
[0017] The correction unit may acquire at least one of a reference brightness value or a reference white balance as the reference exposure amount.
[0018] The viewpoint image generation unit may select a plurality of the plurality of viewpoint image generation models and generate viewpoint images of the predetermined viewpoint based on the selected plurality of viewpoint image generation models.
[0019] The model generation unit may generate a 3DGS (3D Gaussian Splatting) model as the viewpoint image generation model.
[0020] The information processing system may further include an imaging unit that captures an image and generates the captured image.
[0021] An information processing method according to one embodiment of this technology is an information processing method executed by a computer system, which includes acquiring a group of images consisting of a plurality of captured images. A plurality of reference styles are acquired. For each of the plurality of reference styles, a plurality of corrected image groups are generated by correcting the style of the image group to the reference style. A plurality of viewpoint image generation models are generated based on each of the plurality of corrected image groups. A portion of the plurality of viewpoint image generation models is selected based on a predetermined viewpoint. A viewpoint image of the predetermined viewpoint is generated based on the predetermined viewpoint and the selected viewpoint image generation model.
[0022] Each of the plurality of captured images included in the image group acquired by the image group acquisition unit may be a developed image obtained by developing a RAW image.
[0023] The developed image may have an image format such as jpg, png, or bmp.
[0024] The correction process performed by the correction unit may be a process of correcting the style of the group of developed images, which consists of the plurality of developed images, to each of the reference styles acquired by the correction unit.
[0025] The correction unit may acquire at least one of the following as the reference style: reference exposure, reference sharpness, reference blur level, or reference contrast.
[0026] This is a schematic diagram showing an example configuration of an information processing system related to this technology. This is a schematic diagram showing the information stored in the memory unit of the information processing device. This is a schematic diagram showing an example configuration of the model generation unit. This is a schematic diagram showing a flowchart related to the processing of this technology. This is a schematic diagram showing the positional relationship between the camera pose and the input viewpoint. This is a schematic diagram showing an example configuration of an information processing system. This is a schematic diagram showing the information stored in the memory unit. This is a schematic diagram showing a flowchart related to the processing of this technology. This is a block diagram showing an example hardware configuration of a computer capable of realizing the information processing device.
[0027] <First Embodiment> Hereinafter, embodiments relating to this technology will be described with reference to the drawings. This technology is mainly used in the field of photogrammetry. Photogrammetry is a technology that measures objects using images, and by using photogrammetry, it is possible to generate a 3D model of an object that exists in real space.
[0028] Specifically, the user takes multiple images of the object, and a 3D model of the object is generated based on these images. When creating 3D models used in metaverses, etc., it is possible for the user to draw the 3D model themselves using specialized software, but this requires advanced skills and is often inefficient in terms of time and cost. On the other hand, photogrammetry makes it possible to generate 3D models using a relatively simple method of photographing objects that exist in real space.
[0029] Furthermore, this technology generates a novel viewpoint image generation model called 3DGS (3D Gaussian Splatting) as the 3D model mentioned above, and viewpoint images are generated based on the 3DGS model. 3DGS is a method of 3D reconstruction using images from multiple viewpoints as input. In 3DGS, 3D Gaussians are projected and combined from each viewpoint, and each parameter of the Gaussian is optimized to accurately reproduce the input image. As a result, color and density are continuously represented compared to other photogrammetry techniques, and high-quality rendering can be expected.
[0030] A viewpoint image is a virtual image of an object as seen from a predetermined viewpoint. By specifying a desired viewpoint and viewing the viewpoint image from that viewpoint, users can virtually view the object from various directions.
[0031] Furthermore, this technology may be applied to fields other than photogrammetry and viewpoint image generation. In other words, this technology may be applied to any field.
[0032] [Information Processing System] Figure 1 is a schematic diagram showing an example configuration of information processing system 1 related to this technology. Information processing system 1 includes a camera 2 and an information processing device 3.
[0033] Figure 2 is a schematic diagram showing the information stored in the storage unit 11 of the information processing device 3. The various types of information shown in Figure 2 are written to and read from the storage unit 11 as appropriate during processing in this embodiment.
[0034] The camera 2 in Figure 1 is typically a monocular camera, and generates multiple captured images 5 (5a, 5b, ...) shown in Figure 2 by capturing images of an object 4 in real space multiple times. Note that images 5c and beyond are omitted from the illustration.
[0035] The specific number of captured images 5 is not limited. Typically, the object 4 is photographed from multiple different directions with different exposure levels, generating a large number of different captured images 5. The shades of gray in the captured images 5 in Figure 2 represent the exposure levels of each captured image 5. In this example, for example, the exposure levels of captured image 5a and captured image 5b are different.
[0036] Figure 1 shows an image of a landscape as the object 4, but the specific type of object 4 is not limited, and any object can be the object 4, such as a small object or the interior of a room. Camera 2 corresponds to one embodiment of the imaging unit according to this technology.
[0037] The information processing device 3 is typically a smartphone, with a camera 2 built into its casing. Other devices such as a PC (Personal Computer), tablet, or AR (Augmented Reality) glasses may also be used as the information processing device 3.
[0038] The information processing device 3 includes a display unit 6, an operation unit 7, a communication unit 8, a storage unit 11, and a controller 9. These are interconnected via a bus 10. Instead of the bus 10, each block may be connected using a communication network or a non-standardized proprietary communication method.
[0039] The display unit 6 is a display device that uses, for example, liquid crystal, electroluminescence (EL), etc., and displays a GUI (Graphical User Interface), etc. The display unit 6 also displays the captured image 5 and viewpoint image 33 shown in Figure 2.
[0040] The operation unit 7 is a touch panel and is integrated with the display unit 6. When a user wants to virtually view an object 4 from a predetermined viewpoint, they input the desired viewpoint via the operation unit 7. If the information processing device 3 is a device other than a smartphone, the operation unit 7 may be a keyboard or a pointing device.
[0041] The communication unit 8 is a communication module for communicating with other devices via a network such as a LAN (Local Area Network) or WAN (Wide Area Network). It may be equipped with a wireless LAN module such as Wi-Fi or a communication module for short-range wireless communication such as Bluetooth®. Communication equipment such as a modem or router may also be used. The information processing device 3 performs communication with other devices or the cloud via the communication unit 8.
[0042] The storage unit 11 is a storage device such as non-volatile memory, and for example, an HDD (Hard Disk Drive) or SSD (Solid State Drive) may be used. In addition, any non-transient storage medium that can be read by a computer may be used. The storage unit 11 stores, for example, a control program for controlling the overall operation of the information processing device 3. The method of installing the control program in the storage unit 11 is not limited. Also, as mentioned above, the storage unit 11 stores various types of information as shown in Figure 2.
[0043] The controller 9 controls the operation of each block included in the information processing apparatus 3. The controller 9 has hardware circuits necessary for a computer, such as a CPU and memories (RAM, ROM), for example. Various processes are executed when the CPU executes a program related to the present technology stored in the storage unit 11 or the memory. As the controller 9, for example, a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array), or other devices such as an ASIC (Application Specific Integrated Circuit) may be used.
[0044] In the present embodiment, when the CPU of the controller 9 executes a program (for example, an application program) related to the present technology, as functional blocks, an image group acquisition unit 12, a reference image selection unit 13, an exposure amount acquisition unit 14, a correction unit 15, a camera pose estimation unit 16, a model generation unit 17, a model selection unit 18, and a viewpoint image generation unit 19 are implemented. The information processing method according to the present embodiment is executed by these functional blocks. Note that dedicated hardware such as an IC (integrated circuit) may be appropriately used to implement each functional block. Hereinafter, processing by these functional blocks will be described. Note that the contents of some processes may be described in detail later.
[0045] The image group acquisition unit 12 acquires a set of a plurality of captured images 5 captured by the camera 2 as an image group 5' shown in FIG. 2. The captured images 5 at this point in time are RAW images (unprocessed images) or the like. The reference image selection unit 13 selects a predetermined number of captured images 5 as reference images 5 from among the plurality of captured images 5 included in the image group 5'. In this example, three captured images 5d, 5f, and 5j are selected as the reference images 5.
[0046] The exposure amount acquisition unit 14 acquires the respective exposure amounts of the reference images 5d, 5f, and 5j selected by the reference image selection unit 13 as reference exposure amounts. That is, three types of reference exposure amounts are acquired by the exposure amount acquisition unit 14. Specifically, for example, the exposure amounts when the reference images 5d, 5f, and 5j are captured are stored in the storage unit 11, and the exposure amount acquisition unit 14 reads the respective exposure amounts from the storage unit 11. Alternatively, the exposure amount may be stored on the camera 2 side during imaging and transmitted to the exposure amount acquisition unit 14.
[0047] The correction unit 15 corrects the exposure amount of the image group 5' to the reference exposure amount for each of the plurality of reference exposure amounts, and generates a plurality of corrected image groups 20', 21', and 22'. In this example, since there are three reference images 5d, 5f, and 5j, first, a process of RAW developing all the captured images 5 included in the image group 5' is performed in accordance with the exposure amount of the reference image 5d. The exposure amount acquisition unit 14 and the correction unit 15 correspond to an embodiment of the correction unit according to the present technology.
[0048] Specifically, for example, when a luminance value is corrected, the formula "EV = log 2 {N 2 / (TS)}" is used to calculate the corrected exposure amount EV, which is applied to each captured image 5. Here, N represents the aperture value (F-number), T represents the shutter speed (in seconds), and S represents the ISO sensitivity. In addition, any parameter related to exposure amount, such as white balance, may be corrected. Accordingly, a corrected image group 20' including a plurality of corrected images 20 (20a, 20b, ...) is generated.
[0049] The corrected image 20a has the same content as the captured image 5a, but its exposure amount is corrected to the exposure amount of the reference image 5d. That is, for the corrected image 20a and the captured image 5a, the position, direction, etc. where the object 4 is captured are the same. Similarly, the corrected image 20b has the same content as the captured image 5b, but its exposure amount is corrected to the exposure amount of the reference image 5d.
[0050] Next, the captured image 5 is RAW-developed to match the exposure of the reference image 5f, thereby generating a group of correction images 21' consisting of multiple correction images 21 (21a, 21b, ...). For example, correction image 21a has the same content as captured image 5a, but its exposure has been corrected to match the exposure of the reference image 5f.
[0051] Furthermore, by developing the captured image 5 in RAW format according to the exposure amount of the reference image 5j, a group of correction images 22' consisting of multiple correction images 22 (22a, 22b, ...) is generated. For example, correction image 22a has the same content as captured image 5a, but its exposure amount is corrected to the exposure amount of the reference image 5j. Each of the correction images 20, 21, and 22 has a format such as jpg or png.
[0052] The camera pose estimation unit 16 estimates the camera pose 24 at the time of capturing the captured image 5. Specifically, it first generates a group of estimation images 23' consisting of multiple estimation images 23 (23a, 23b, ...) by developing each captured image 5 with automatic exposure settings. Estimation image 23a is the captured image 5a developed with automatic exposure settings, and estimation image 23b is the captured image 5b developed with automatic exposure settings. The specific exposure amount related to the automatic exposure settings is not particularly limited and is set to an appropriate value in advance. Each estimation image 23 has a format such as jpg or png.
[0053] Next, based on each estimation image 23, the camera pose 24 at the time of acquisition of each captured image 5 is estimated. That is, the camera pose 24 of captured image 5a is estimated based on the estimation image 23a, the camera pose 24 of captured image 5b is estimated based on the estimation image 23b, and so on, and this process is repeatedly executed.
[0054] The camera pose 24 is estimated, for example, by inputting the captured image 5 into a machine learning model. In addition, the information processing system 1 may be equipped with distance measuring sensors such as LiDAR (Light Detection And Ranging) or a stereo camera, and the sensing results from these may be used for estimation. The camera pose 24 at the time of imaging corresponds to the viewpoint at the time of imaging according to this technology.
[0055] Figure 3 is a schematic diagram showing an example of the configuration of the model generation unit 17. The model generation unit 17 generates multiple 3DGS models 25, 26, and 27 based on each of the multiple correction image groups 20', 21', and 22'. As shown in Figure 3, the model generation unit 17 includes a rendering unit 28, a loss function calculation unit 29, and a parameter update unit 30. In the example in Figure 3, the 3DGS model 25 is generated based on the correction image group 20'. Similarly, the 3DGS model 26 is generated based on the correction image group 21', and the 3DGS model 27 is generated based on the correction image group 22'.
[0056] The rendering unit 28 generates a rendered image 31 for the camera pose 24 by inputting the camera pose 24 to the currently generated 3DGS model 25. In this example, it is assumed that the camera pose 24 input is the camera pose 24 at the time the captured image 5b was captured.
[0057] The loss function calculation unit 29 compares the rendered image 31 and the corrected image 20 and calculates the loss using the loss function. The loss is typically an indicator of the difference between the two images. In this example, the camera pose 24 at the time of capturing the captured image 5b is used, so the corrected image 20b is input as the corrected image 20. Thus, the captured image 5b related to the camera pose 24 input to the rendering unit 28 and the captured image 5b that serves as the basis for the corrected image 20b input to the loss function calculation unit 29 are the same.
[0058] The parameter update unit 30 performs training of the 3DGS model 25 by updating the internal parameters of the 3DGS model 25 based on the loss calculated by the loss function calculation unit 29.
[0059] This learning process is repeated until predetermined conditions are met. The initial values for the 3DGS model 25 are random values or 3D feature points used during the estimation of the camera pose 24. Note that the method for generating the 3DGS model 25 in this example is merely one example, and the specific method is not limited to this example.
[0060] The model selection unit 18 selects a portion of the multiple 3DGS models 25 to 27 based on the input viewpoint 32 entered by the user. In this example, the model selection unit 18 has selected 3DGS model 25. The input viewpoint 32 corresponds to a predetermined viewpoint related to this technology.
[0061] The viewpoint image generation unit 19 generates a viewpoint image 33 of the input viewpoint 32 based on the input viewpoint 32 and the 3DGS model 25 selected by the model selection unit 18. The viewpoint image generation unit 19 has a rendering unit with a configuration similar to the rendering unit 28 of the model generation unit 17 shown in Figure 3, and generates the viewpoint image 33 as a rendered image by inputting the 3DGS model 25 and the input viewpoint 32 to the rendering unit. The model selection unit 18 and the viewpoint image generation unit 19 correspond to one embodiment of the viewpoint image generation unit according to this technology.
[0062] Furthermore, the specific configuration of each functional block is not limited. Some functions of the information processing device 3 may also be configured on the cloud. For example, the reference image selection unit 13, exposure amount acquisition unit 14, correction unit 15, camera pose estimation unit 16, model generation unit 17, model selection unit 18, and viewpoint image generation unit 19, which have relatively high processing loads, may be configured on the cloud.
[0063] [Processing Flow] Figure 4 is a schematic diagram showing the flowchart of the processing of this technology. Camera 2 captures an image of the object 4 (step 101). Image group acquisition unit 12 acquires an image group 5' consisting of the captured images 5 (step 102).
[0064] The reference image selection unit 13 selects the reference image 5 (step 103). In the example in Figure 2, three reference images 5d, 5f, and 5j were selected, but the number of reference images 5 is not limited; for example, all of the captured images 5 included in the image group 5' may be selected as reference images 5.
[0065] Furthermore, the reference image selection unit 13 may select a user-specified image 5 from among the captured images 5 as the reference image 5. In this case, some or all of the captured images 5 are displayed on the display unit 6, and the user selects multiple reference images 5 via the operation unit 7 while checking the display unit 6. Additionally, for example, a display or sound prompting the user to select multiple captured images 5 that include the object 4 but appear differently may be output to the user.
[0066] Alternatively, the reference image selection unit 13 may select the reference image 5 based on the brightness distribution of the captured image 5. For example, the captured image 5 may be sorted in order of brightness, and a predetermined number of reference images 5 may be selected at equal intervals.
[0067] Alternatively, the reference image selection unit 13 may select an image 5 containing the object 4 from among the captured images 5 as the reference image 5. In this case, for example, the reference image selection unit 13 determines whether or not the object 4 is included in the captured image 5 by using object detection. In addition, the presence or absence of the object 4 may be determined by any other method.
[0068] Furthermore, the method of selecting the reference image 5 by the reference image selection unit 13 is not limited, and the reference image 5 may be selected based on any arbitrary criteria. Multiple selection methods may also be used in combination.
[0069] The exposure amount acquisition unit 14 acquires a reference exposure amount (step 104). In this example, the exposure amount of each reference image 5 is acquired as the reference exposure amount. Alternatively, multiple reference exposure amounts may be acquired based on multiple captured images 5. That is, a reference image 5 may not be selected, and the reference exposure amount may be acquired based on the original captured image 5. In this case, the controller 9 does not have a reference image selection unit 13, and the processing in step 103 may be omitted.
[0070] For example, one possible method is to sort the exposure levels of all captured images 5 in descending order and select a predetermined number of exposure levels at equal intervals as reference exposure levels. Furthermore, there are no limitations on the specific methods for obtaining multiple reference exposure levels based on multiple captured images 5. It is also possible to obtain a reference exposure level that is not the exposure level of any single captured image 5, such as the average of the exposure levels of two captured images 5.
[0071] Furthermore, the reference exposure amount does not necessarily have to be obtained based on the captured image 5. For example, multiple reference exposure amounts may be stored in the storage unit 11 in advance and obtained by the exposure amount acquisition unit 14. In addition, the specific method for obtaining the reference exposure amount is not limited.
[0072] The correction unit 15 generates the corrected image group 20' to 22' (step 105). The camera pose estimation unit 16 and the model generation unit 17 generate 3DGS models 25 to 27 (step 106).
[0073] Figure 5 is a schematic diagram showing the positional relationship between the camera pose 24 and the input viewpoint 32. The model selection unit 18 selects some of the 3DGS models 25 to 27 based on the input viewpoint 32 entered by the user (step 107). For example, the model selection unit 18 selects the 3DGS models 25 to 27 that correspond to the reference image 5 in which the camera pose 24 at the time of imaging is closest to the input viewpoint 32.
[0074] In this example, the user inputs the input viewpoint 32a. The position of each reference image 5 in the figure represents the position of the camera pose 24 when the reference image 5 was captured. First, the distance between the camera pose 24 and the input viewpoint 32a when the reference image 5d was captured, the distance between the camera pose 24 and the input viewpoint 32a when the reference image 5f was captured, and the distance between the camera pose 24 and the input viewpoint 32a when the reference image 5j was captured are calculated and compared. As a result, as shown in the figure, it is determined that the distance between the camera pose 24 and the input viewpoint 32a when the reference image 5d was captured is the smallest value, and the 3DGS model 25 related to reference image 5d is selected. On the other hand, if the input viewpoint 32a is in a different position, the 3DGS model 26 or 27 may be selected.
[0075] Alternatively, for example, the image group acquisition unit 12 may classify the image group 5' into multiple clusters 34 such that each cluster 34 contains one reference image 5. Then, the model selection unit 18 may select 3DGS models 25 to 27 related to the reference image 5 of the cluster 34 that contains the captured image 5 whose camera pose 24 at the time of imaging is closest to the input viewpoint 32.
[0076] In Figure 5, clusters 34a to 34c are illustrated with dashed frames. Cluster 34a includes the reference image 5d and the captured images 5c and 5g. Cluster 34b includes the reference image 5f and the captured images 5a and 5b. Cluster 34c includes the reference image 5j and the captured images 5e, 5h, and 5i. In other words, each cluster 34 contains only one reference image 5.
[0077] In this example, the user inputs the viewpoint 32b. Next, for each of the captured images 5a to 5j, the "distance between the camera pose 24 at the time of capturing image 5 and the input viewpoint 32b" is calculated and compared. As a result, as shown in the figure, it is determined that the "distance between the camera pose 24 at the time of capturing image 5g and the input viewpoint 32b" is the smallest value. Since the cluster 34 containing the captured image 5g is cluster 34a, the reference image 5d of cluster 34a is selected, and the 3DGS model 25 related to the reference image 5d is selected.
[0078] In this example, if the camera pose 24 during imaging is closest to the input viewpoint 32b, and the 3DGS models 25 to 27 corresponding to the reference image 5 are selected, then the 3GDS model 26 corresponding to the closest reference image 5f will be selected. However, by using a method to classify into clusters 34, a different 3DGS model 25 will be acquired. On the other hand, if the input viewpoint 32b is at a different location, either 3DGS model 26 or 27 may be selected.
[0079] The specific classification method for cluster 34 is not limited, but for example, a method of dividing real space into predetermined areas, as in this example, can be used. Also, one cluster 34 may contain multiple reference images 5.
[0080] In the information processing system 1 according to this embodiment, a group of images 5' is corrected by a group of reference exposure amounts, thereby generating a group of corrected images 20' to 22'. Based on the group of corrected images 20' to 22', a group of 3DGS models 25 to 27 are generated. Furthermore, a portion of the group of 3DGS models 25 to 27 is selected based on the input viewpoint 32, and a viewpoint image 33 of the input viewpoint 32 is generated based on the selected 3DGS models 25 to 27 and the input viewpoint 32. This makes it possible to generate a high-quality viewpoint image 33.
[0081] In live-action scenes, exposure parameters are adjusted to achieve a clean image, resulting in a mix of images with different exposure levels (ISO, shutter speed, etc.). Therefore, images of the same part of an object may contain extremely bright or dark areas. Consequently, it may not be possible to properly compensate for these differences in exposure between images.
[0082] For example, an image of an indoor space taken from outdoors tends to have a low ISO because the ambient light intensity is low. Conversely, an image of the same indoor space taken from indoors tends to have a high ISO because the ambient light intensity is high. If you try to match these images, the colors will not match even though they are images of the same space.
[0083] Therefore, one might consider processing the image by, for example, placing a black haze between the camera and the subject. However, this results in an unnatural black haze-like artifact in the rendering result, which can give users who view the image a sense of unease.
[0084] Another simple method is to adjust the exposure of images with significantly different brightness levels afterward. However, changing parameters to unify the exposure during RAW development may result in overexposure or underexposure that was not noticed during shooting.
[0085] Furthermore, research has been conducted using exposure compensation as an optimization parameter, addressing the problem of floating artifacts. However, the purpose of exposure compensation is to unify the appearance from various viewpoints, not to reproduce the original captured image. Therefore, the rendering results expected at the time of imaging may not be obtained.
[0086] Another approach is to manually adjust the exposure parameters and perform RAW development to equalize the exposure conditions that differed during shooting. However, this method requires tedious work and is therefore a significant burden for the user.
[0087] In this technology, the optimal 3DGS model 25-27 is selected based on the input viewpoint 32, and a viewpoint image 33 is generated based on the selected 3DGS model 25-27 and the input viewpoint 32. This makes it possible to obtain a viewpoint image 33 that suppresses overexposure and underexposure while automatically adjusting for exposure differences that cause artifacts.
[0088] Furthermore, this technology acquires multiple reference exposure levels based on multiple captured images 5. This makes it possible to generate appropriate 3DGS models 25 to 27.
[0089] Furthermore, in this technology, a reference image 5 is selected from the image group 5', and the exposure amount of each reference image 5 is obtained as the reference exposure amount. This makes it possible to generate even more appropriate 3DGS models 25 to 27.
[0090] Furthermore, in this technology, all captured images 5 included in the image group 5' are selected as reference images 5. As a result, a large number of 3DGS models 25 to 27 are generated, making it possible to select an even more appropriate 3DGS model 25 to 27.
[0091] Furthermore, in this technology, the captured image 5 specified by the user is selected as the reference image 5. As a result, since the user's decision is involved in the selection of the reference image 5, it becomes possible to select the reference image 5 while eliminating problems that cannot be detected by the information processing system 1.
[0092] Furthermore, in this technology, a reference image 5 is selected based on the brightness distribution of the captured image 5. This makes it possible to select a reference exposure amount without bias, for example, and generate a high-quality viewpoint image 33.
[0093] Furthermore, in this technology, the image 5 containing the object 4 is selected as the reference image 5. This allows for the selection of an even more appropriate reference exposure amount.
[0094] Furthermore, this technology uses object detection to determine whether or not the captured image 5 contains the target object 4. This makes it possible to determine with even greater accuracy whether or not the target object 4 is included.
[0095] Furthermore, in this technology, the 3DGS models 25 to 27 corresponding to the reference image 5 that is closest to the input viewpoint 32 when the camera pose 24 is captured are selected. This generates an even higher quality viewpoint image 33.
[0096] In this technology, the image group 5' is classified into multiple clusters 34 such that each cluster 34 contains one reference image 5. Furthermore, 3DGS models 25-27 related to the reference image 5 are selected from the cluster 34 containing the image capture image 5 whose camera pose 24 at the time of acquisition is closest to the input viewpoint 32. This results in the generation of an even higher quality viewpoint image 33.
[0097] <Second Embodiment> A more detailed embodiment of the information processing system 1 related to this technology will be described as a second embodiment. In the following description, parts that are similar to the configuration, operation, and effects of the information processing system 1 described in the above embodiment may be omitted or simplified.
[0098] [Style Correction] Figure 6 is a schematic diagram showing an example of the configuration of the information processing system 1. Figure 7 is a schematic diagram showing the information stored in the storage unit 11. The differences between the processing content of this embodiment and the processing content of embodiments such as Figure 2 will be explained. First, in Figure 2, the image group acquisition unit 12 acquired the unprocessed RAW images captured by the camera 2.
[0099] On the other hand, in this embodiment, as shown in Figure 7, the image group acquisition unit 12 acquires a group of developed images 5' consisting of multiple developed images 5. The developed images 5 are images developed from RAW images and have, for example, a jpg or png image format. The developed images 5 may also have other image formats such as bmp.
[0100] For example, images are captured by multiple different types of cameras 2, and developed images 5 are acquired by the image group acquisition unit 12. In this case, as shown in Figure 6, multiple types of cameras 2 (2a, 2b, 2c) are configured within the information processing system 1. Although three types of cameras 2 are shown in this example, the specific number of camera types 2 is not limited. Possible combinations of multiple different types of cameras 2 include, for example, single-lens reflex cameras and 360-degree cameras. Furthermore, the specific types of cameras 2 are not limited.
[0101] Alternatively, the developed images 5 may be collected from multiple different sources on the Internet. In this case, the camera 2 is not configured within the information processing system 1. Furthermore, the specific method of acquiring the developed images 5 by the image acquisition unit 12 is not limited.
[0102] Each developed image 5 in the developed image group 5' is typically captured using different imaging equipment and techniques, and therefore each has a different style. Style includes exposure, sharpness, bokeh, contrast, etc. Depth of field and super-resolution may also be included in the style. Exposure includes exposure, hue, white balance, etc.
[0103] Let me explain another difference. In Figure 2, in order to estimate the camera pose 24, the image group 5', which is a RAW image, was developed with automatic exposure settings to generate the estimation image group 23'. On the other hand, in this embodiment, the estimation image group 23' is not generated. This is because in this embodiment, the developed image group 5' is obtained, and it is possible to directly estimate the camera pose 24 based on that developed image group 5'.
[0104] [Processing Flow] Figure 8 is a schematic diagram showing the flowchart of the processing of this technology. The image group acquisition unit 12 acquires the developed image group 5' (step 201). Next, it is determined whether the user has specified a reference image 5 or not (step 202). In this example, the system does not select a reference image 5, and the user specifies the reference image 5 themselves.
[0105] For example, the user views the developed image group 5' and designates several developed images 5 that appear pleasing as reference images 5. For example, consider the case where a single-lens reflex camera 2a and a 360-degree camera 2b are used as camera 2. Since the single-lens reflex camera can capture relatively clearer images than the 360-degree camera, the developed images 5 captured by the single-lens reflex camera 2a are more likely to be designated as reference images 5 by the user.
[0106] The reference image selection unit 13 accepts the reference image 5 specified by the user. Alternatively, the system may select the reference image 5 using the same method as described above. For example, a developed image 5 containing the object 4 may be selected as the reference image 5.
[0107] The determination is repeated until the user specifies a reference image 5 (No in step 202). If a reference image 5 is specified (Yes in step 202), the style acquisition unit 40 in Figure 6 acquires the reference style for each reference image 5 (step 203). The reference style acquired by the style acquisition unit 40 includes at least one of the reference exposure, reference sharpness, reference blur, or reference contrast of the reference image 5. The reference style may also include reference depth of field or reference super-resolution. The reference exposure includes reference exposure, reference hue, reference white balance, etc.
[0108] Furthermore, the style of an image other than the user-prepared developed image 5 (hereinafter referred to as a reference image) may be acquired as the base style. For example, if the styles of all the developed images 5 are of low quality (do not look good to the user), the user may prepare a reference image with a high-quality style. Alternatively, the user may select one of the developed images 5, edit the selected developed image 5 to achieve the desired color tone, etc., and use the edited image as the reference image.
[0109] In these cases, while the developed image 5 is always a monochrome or grayscale image, a color reference image may also be used.
[0110] Correction image groups 20' to 22' are generated (step 204). The correction unit 15 generates correction image groups 20' to 22' by performing a process to correct the style of the developed image group 5' to the respective reference styles acquired by the style acquisition unit 40. For example, the correction image 20a in Figure 7 has the same content as the developed image 5a, but its style has been corrected to the style of the reference image 5d.
[0111] Correction methods include, for example, classical color consistency optimization methods and style correction using machine learning models, but the specific method is not limited. The style acquisition unit 40 and the correction unit 15 correspond to one embodiment of the correction unit according to this technology.
[0112] The processing steps 205 to 208 are the same as those in steps 106 to 109, so their explanation will be omitted.
[0113] The effects of this embodiment will now be explained. In order to generate a 3DGS model, it may be desirable to use a group of images captured by multiple types of cameras. For example, when using both a single-lens reflex camera and a 360-degree camera, it may be desirable to appropriately change the camera used depending on the part of the object being studied.
[0114] While 360-degree cameras allow for easy and comprehensive shooting, the resulting images tend to be of low quality. In other words, users often perceive the captured images as being of poor quality.
[0115] On the other hand, images captured with a single-lens reflex (SLR) camera tend to have a high-quality style. In other words, it is possible to capture images that look very close to the real-world object. However, SLR cameras make it difficult to capture a wide range of subjects, and the cost of shooting tends to be high.
[0116] Furthermore, there are cases where it is desirable to use a collection of images gathered from the internet to generate a 3DGS model. For example, when generating a 3DGS model of a large building, it is difficult for the user to take photographs of the building themselves. In such cases, multiple images of the building are collected from various sources on the internet. However, since these images were taken with different imaging equipment, the style of each image will differ.
[0117] As described above, when using different types of cameras or when collecting captured images from the internet, if no processing is performed, artifacts will appear in the rendered image due to differences in the style of each captured image.
[0118] In this embodiment, the correction unit 15 performs a process to correct the style of the developed image group 5' to its respective reference style. This prevents the occurrence of artifacts in the rendered image. In addition, since the optimal 3DGS models 25 to 27 are selected based on the input viewpoint 32, blown-out highlights and crushed blacks are suppressed.
[0119] In addition, since it becomes possible to use methods such as imaging with multiple types of cameras 2 or collecting developed images 5 from the internet, the convenience of generating 3DGS models 25 to 27 is improved.
[0120] For example, it is possible to comprehensively photograph the object 4 with a 360-degree camera while using the style of a single-lens reflex camera for generating the 3DGS model. Furthermore, even if the developed image 5 collected from the internet is a monochrome image, it is possible to generate color 3DGS models 25-27.
[0121] In this embodiment, the developed image 5 is acquired by the image group acquisition unit 12. Furthermore, the developed image 5 has an image format such as jpg or png. While RAW images function even in environments requiring high dynamic range, such as dark places, their large file size results in a high processing load on the system. In this embodiment, since the developed image 5 is acquired, the processing load is reduced.
[0122] Furthermore, while it is impossible to correct styles such as depth of field and super-resolution during the development of a RAW image, in this example, since the source for correction is the developed image 5, it is possible to correct depth of field and super-resolution.
[0123] In this embodiment, the style acquisition unit 40 acquires at least one of the following: reference exposure, reference sharpness, reference blur level, or reference contrast. This further enhances the accuracy of suppressing artifacts and other issues.
[0124] <Other Embodiments> This technology is not limited to the embodiments described above, and various other embodiments can be realized.
[0125] In this embodiment, a 3DGS model was used as the viewpoint image generation model, but other viewpoint image generation models such as NeRF (Neural Radiance Fields) may also be used.
[0126] In the example in Figure 2, a group of estimation images 23' is generated separately from the corrected image group 20' to 22', and each camera pose 24 is estimated based on the estimation image group 23'. Since the estimation image group 23' was developed with automatic exposure settings, the camera pose 24 can be estimated efficiently by using the estimation image group 23'. Alternatively, the camera pose 24 may be estimated based on the corrected image group 20' to 22'.
[0127] In this embodiment, three reference images 5d, 5f, and 5j were selected, and three 3DGS models 25 to 27 were generated. However, the number of reference images 5 selected can be arbitrary. For example, if you want to generate a high-quality viewpoint image 33, a large number of reference images 5 (up to the number of captured images 5) can be selected. On the other hand, if you want to prioritize processing speed, a small number of reference images 5 can be selected.
[0128] The model selection unit 18 may select multiple 3DGS models, and the viewpoint image generation unit 19 may generate a viewpoint image 33 of the input viewpoint 32 based on the multiple 3DGS models. In other words, the multiple 3DGS models may be blended to generate the viewpoint image 33. This makes it possible to generate an even higher quality viewpoint image 33.
[0129] The information processing system 1 may be implemented by multiple computers or by a single computer.
[0130] Figure 9 is a block diagram showing an example of the hardware configuration of a computer 500 capable of realizing the information processing device 3. The computer 500 includes a CPU 501, ROM 502, RAM 503, an input / output interface 505, and a bus 504 connecting these to each other. A display unit 506, an input unit 507, a storage unit 508, a communication unit 509, and a drive unit 510 are connected to the input / output interface 505.
[0131] The display unit 506 is a display device using, for example, liquid crystal, EL, etc. The input unit 507 is, for example, a keyboard, pointing device, touch panel, or other operating device. If the input unit 507 includes a touch panel, the touch panel may be integrated with the display unit 506. The storage unit 508 is a non-volatile storage device, for example, an HDD, flash memory, or other solid memory. The drive unit 510 is a device capable of driving the removable recording medium 511, for example, an optical recording medium or magnetic recording tape. The communication unit 509 is a modem, router, or other communication device for communicating with other devices, which can be connected to a LAN, WAN, etc. The communication unit 509 may communicate using either wired or wireless methods. The communication unit 509 is often used separately from the computer 500.
[0132] Information processing by the computer 500 having the hardware configuration described above is realized through the cooperation of software stored in the memory unit 508 or ROM 502, etc., and the hardware resources of the computer 500. Specifically, the information processing method related to this technology is realized by loading the programs that constitute the software, stored in the ROM 502, etc., into the RAM 503 and executing them.
[0133] The program is installed on the computer 500, for example, via a removable recording medium 511. Alternatively, the program may be installed on the computer 500 via a global network or the like. In addition, any non-transient storage medium that the computer 500 can read may be used.
[0134] In this disclosure, "system" means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules in one enclosure, are both considered systems.
[0135] The execution of the information processing method related to this technology by a computer system includes both cases where the acquisition of image groups and developed image groups, selection of reference images, acquisition of reference styles, generation of corrected image groups, generation and selection of 3DGS models, generation and display of viewpoint images, etc., are performed by a single computer, and cases where each process is performed by different computers. Furthermore, the execution of each process by a predetermined computer includes having another computer perform part or all of the process and obtaining the results. In other words, the information processing method related to this technology can also be applied to cloud computing configurations in which a single function is shared and processed jointly by multiple devices via a network.
[0136] The information processing system, image group, developed image group, reference image, corrected image group, estimation image group, camera pose, 3DGS model, input viewpoint, viewpoint image, and processing flows described with reference to each drawing are merely one embodiment and can be arbitrarily modified without departing from the spirit of this technology. In other words, other arbitrary configurations and algorithms may be adopted to implement this technology.
[0137] Where the word "abbreviated" is used in this disclosure, it is used solely to facilitate understanding of the explanation, and there is no special meaning in the use or non-use of the word "abbreviated." In other words, in this disclosure, concepts that define shape, size, positional relationships, states, etc., such as "center," "central," "uniform," "equal," "same," "orthogonal," "parallel," "symmetric," "extending," "axial," "cylindrical," "cylindrical shape," "ring shape," and "annular shape," include concepts such as "substantially centered," "substantially central," "substantially uniform," "substantially equal," "substantially the same," "substantially orthogonal," "substantially parallel," "substantially symmetric," "substantially extending," "substantially axial," "substantially cylindrical," "substantially cylindrical shape," "substantially ring shape," and "substantially annular shape." For example, states that fall within a predetermined range (e.g., a range of ±10%) based on "perfectly centered," "perfectly central," "perfectly uniform," "perfectly equal," "perfectly the same," "perfectly orthogonal," "perfectly parallel," "perfectly symmetric," "perfectly extending," "perfectly axial," "perfectly cylindrical," "perfectly cylindrical shape," "perfectly ring shape," and "perfectly annular shape" are also included. Therefore, even if the word "abbreviated" is not added, the concept may still be included in what is typically expressed with "abbreviated" added. Conversely, the state expressed with "abbreviated" does not mean that the complete state is excluded.
[0138] In this disclosure, expressions using "greater than A" such as "greater than A" and "less than A" are expressions that comprehensively include both concepts that include cases where something is equivalent to A and concepts that do not include cases where something is equivalent to A. For example, "greater than A" is not limited to cases where something is not equivalent to A, but also includes "greater than or equal to A". Similarly, "less than A" is not limited to "less than A", but also includes "less than or equal to A". When implementing this technology, you may appropriately adopt specific settings from the concepts included in "greater than A" and "less than A" so that the effects described above are achieved.
[0139] It is also possible to combine at least two of the feature features of the present technology described above. In other words, the various feature features described in each embodiment may be combined arbitrarily, regardless of the specific embodiment. Furthermore, the various effects described above are merely examples and not limiting, and other effects may also be exhibited.
[0140] Furthermore, this technology can also be configured as follows: (1) An information processing system comprising: an image group acquisition unit that acquires an image group consisting of multiple captured images; a correction unit that acquires multiple reference styles and generates multiple corrected image groups by correcting the style of the image group to the reference style for each of the multiple reference styles; a model generation unit that generates multiple viewpoint image generation models based on each of the multiple corrected image groups; and a viewpoint image generation unit that selects a part of the multiple viewpoint image generation models based on a predetermined viewpoint and generates a viewpoint image of the predetermined viewpoint based on the predetermined viewpoint and the selected viewpoint image generation model. (2) The information processing system according to (1), wherein the correction unit acquires the multiple reference styles based on the multiple captured images. (3) The information processing system according to (2), further comprising: a reference image selection unit that selects a predetermined number of reference images from the image group, wherein the correction unit acquires the predetermined number of styles of each of the reference images selected by the reference image selection unit as the reference styles. (4) An information processing system according to (3), wherein the reference image selection unit selects all of the captured images included in the image group as the reference image. (5) An information processing system according to (3) or (4), wherein the reference image selection unit selects from among the captured images included in the image group a captured image specified by the user as the reference image. (6) An information processing system according to any one of (3) to (5), wherein the reference image selection unit selects the reference image based on the brightness distribution of the captured images included in the image group. (7) An information processing system according to any one of (3) to (6), wherein the reference image selection unit selects from among the captured images included in the image group a captured image that includes an object in real space as the reference image. (8) An information processing system according to (7), wherein the reference image selection unit determines whether or not the captured image includes an object by using object detection.(9) An information processing system according to any one of (3) to (8), wherein the viewpoint image generation unit selects the viewpoint image generation model relating to the reference image whose viewpoint at the time of imaging is closest to the predetermined viewpoint. (10) An information processing system according to any one of (3) to (9), wherein the image group acquisition unit classifies the image group into a plurality of clusters such that one reference image is included in one cluster, and the viewpoint image generation unit selects the viewpoint image generation model relating to the reference image of the cluster containing the captured image whose viewpoint at the time of imaging is closest to the predetermined viewpoint. (11) An information processing system according to any one of (1) to (10), wherein the correction unit acquires a reference exposure amount as the reference style. (12) An information processing system according to (11), wherein the correction unit acquires at least one of a reference brightness value or a reference white balance as the reference exposure amount. (13) An information processing system according to any one of (1) to (12), wherein the viewpoint image generation unit selects a plurality of viewpoint image generation models and generates a viewpoint image of a predetermined viewpoint based on the selected plurality of viewpoint image generation models. (14) An information processing system according to any one of (1) to (13), wherein the model generation unit generates a 3DGS (3D Gaussian Splatting) model as the viewpoint image generation model. (15) An information processing system according to any one of (1) to (14), further comprising an imaging unit that captures an image and generates the captured image.(16) An information processing method in which a computer system performs the following: acquiring a group of images consisting of multiple captured images, acquiring multiple reference styles, generating a group of corrected images by correcting the style of the image group to the reference style for each of the multiple reference styles, generating a group of viewpoint image generation models based on each of the multiple corrected image groups, selecting a part of the multiple viewpoint image generation models based on a predetermined viewpoint, and generating a viewpoint image for the predetermined viewpoint based on the predetermined viewpoint and the selected viewpoint image generation model. (17) An information processing system according to any one of (1) to (15), wherein each of the multiple captured images included in the image group acquired by the image group acquisition unit is a developed image obtained by developing a RAW image. (18) An information processing system according to (17), wherein the developed image has an image format of jpg, png, or bmp. An information processing system according to (19), (17), or (18), wherein the correction processing by the correction unit is a process of correcting the style of the group of developed images consisting of the plurality of developed images to each of the reference styles acquired by the correction unit. An information processing system according to (20), (19), wherein the correction unit acquires at least one of a reference exposure, reference sharpness, reference blur, or reference contrast as the reference style.
[0141] 1... Information processing system 2... Camera 3... Information processing device 4... Object 5... Captured image, reference image, developed image 5'... Image group, developed image group 9... Controller 12... Image group acquisition unit 13... Reference image selection unit 14... Exposure amount acquisition unit 15... Correction unit 16... Camera pose estimation unit 17... Model generation unit 18... Model selection unit 19... Viewpoint image generation unit 20, 21, 22... Corrected images 20', 21', 22'... Corrected image group 23... Estimation image 23'... Estimation image group 24... Camera pose 25, 26, 27... 3GDS model 28... Rendering unit 29... Loss function calculation unit 30... Parameter update unit 31... Rendered image 32... Input viewpoint 33... Viewpoint image 34... Cluster 40... Style acquisition unit
Claims
1. An information processing system comprising: an image group acquisition unit that acquires an image group consisting of multiple captured images; a correction unit that acquires multiple reference styles and generates multiple corrected image groups by correcting the style of the image group to the reference style for each of the multiple reference styles; a model generation unit that generates multiple viewpoint image generation models based on each of the multiple corrected image groups; and a viewpoint image generation unit that selects a portion of the multiple viewpoint image generation models based on a predetermined viewpoint and generates a viewpoint image of the predetermined viewpoint based on the predetermined viewpoint and the selected viewpoint image generation model.
2. An information processing system according to claim 1, wherein the correction unit acquires the plurality of reference styles based on the plurality of captured images.
3. An information processing system according to claim 2, further comprising a reference image selection unit that selects a predetermined number of reference images from the group of images, wherein the correction unit acquires a predetermined number of styles of each of the reference images selected by the reference image selection unit as the reference styles.
4. An information processing system according to claim 3, wherein the reference image selection unit selects all of the captured images included in the image group as the reference image.
5. An information processing system according to claim 3, wherein the reference image selection unit selects an image captured by a user from among the captured images included in the image group as the reference image.
6. An information processing system according to claim 3, wherein the reference image selection unit selects the reference image based on the brightness distribution of the captured images included in the image group.
7. An information processing system according to claim 3, wherein the reference image selection unit selects from the captured images included in the image group an image containing an object in real space as the reference image.
8. An information processing system according to claim 7, wherein the reference image selection unit determines whether or not the captured image contains the object by using object detection.
9. An information processing system according to claim 3, wherein the viewpoint image generation unit selects the viewpoint image generation model relating to the reference image whose viewpoint at the time of imaging is closest to the predetermined viewpoint.
10. An information processing system according to claim 3, wherein the image group acquisition unit classifies the image group into a plurality of clusters such that one reference image is included in one cluster, and the viewpoint image generation unit selects the viewpoint image generation model relating to the reference image of the cluster that includes the captured image whose viewpoint at the time of imaging is closest to the predetermined viewpoint.
11. An information processing system according to claim 1, wherein the correction unit acquires a reference exposure amount as the reference style.
12. An information processing system according to claim 11, wherein the correction unit acquires at least one of a reference luminance value or a reference white balance as the reference exposure amount.
13. An information processing system according to claim 1, wherein the viewpoint image generation unit selects a plurality of viewpoint image generation models and generates viewpoint images of a predetermined viewpoint based on the selected plurality of viewpoint image generation models.
14. An information processing system according to claim 1, wherein the model generation unit generates a 3DGS (3D Gaussian Splatting) model as the viewpoint image generation model.
15. An information processing system according to claim 1, further comprising an imaging unit for capturing an image and generating the captured image.
16. An information processing method in which a computer system performs the following steps: acquire a group of images consisting of multiple captured images; acquire multiple reference styles; generate a group of corrected images by correcting the style of the image group to the reference style for each of the multiple reference styles; generate a group of viewpoint image generation models based on each of the multiple corrected image groups; select a portion of the multiple viewpoint image generation models based on a predetermined viewpoint; and generate a viewpoint image for the predetermined viewpoint based on the predetermined viewpoint and the selected viewpoint image generation model.
17. An information processing system according to claim 1, wherein each of the plurality of captured images included in the image group acquired by the image group acquisition unit is a developed image obtained by developing a RAW image.
18. An information processing system according to claim 17, wherein the developed image is in the image format jpg, png, or bmp.
19. An information processing system according to claim 17, wherein the correction process performed by the correction unit is a process of correcting the style of the group of developed images consisting of the plurality of developed images to each of the reference styles acquired by the correction unit.
20. An information processing system according to claim 19, wherein the correction unit acquires at least one of a reference exposure, reference sharpness, reference blur, or reference contrast as the reference style.