Image processing device, image processing method, and program

By selectively processing specific images based on the virtual camera's position and orientation, the image processing apparatus reduces costs and enhances image quality in generating virtual viewpoint images.

JP2026076516APending Publication Date: 2026-05-12CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-10-24
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The processing cost increases when improving the image quality of virtual viewpoint images using specific processes on all captured images.

Method used

An image processing apparatus that selectively determines specific images from among a plurality of captured images based on the position and orientation of a virtual camera, performing specific processing such as super-resolution on these images to generate a virtual viewpoint image, thereby reducing processing costs.

Benefits of technology

This approach reduces processing costs while maintaining or improving image quality by efficiently applying specific processing only to necessary images, allowing for the generation of high-quality virtual viewpoint images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026076516000001_ABST
    Figure 2026076516000001_ABST
Patent Text Reader

Abstract

Applying a specific process to all of the multiple captured images used to generate a virtual viewpoint image increases processing costs. [Solution] The rendering apparatus 40 includes acquisition means for acquiring first information indicating the position and orientation of a virtual camera, determination means for determining a specific image to be used for generating color information of a virtual viewpoint image from among a plurality of captured images used for generating a virtual viewpoint image corresponding to the virtual camera, based on the first information, and processing means for performing specific processing on at least a part of the specific image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing apparatus, an image processing method, and a program, and particularly to a technique for generating a virtual viewpoint image.

Background Art

[0002] Techniques for installing a plurality of cameras at different positions and synchronously capturing images from multiple viewpoints, and generating a virtual viewpoint image using the plurality of captured images obtained by the capturing have attracted attention. According to the technique of generating a virtual viewpoint image from a plurality of captured images, it is possible to produce content with a compelling viewpoint, including viewpoints that were not previously accessible by a camera.

[0003] In generating a virtual viewpoint image using this technique, it is required to generate a virtual viewpoint image with higher image quality. As a countermeasure, Patent Document 1 describes a method of improving the image quality of a virtual viewpoint image by performing super-resolution processing on the captured image used for generating the virtual viewpoint image.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, if a specific process for improving the image quality is performed on all of the plurality of captured images used for generating the virtual viewpoint image, the processing cost increases.

[0006] An object of the present disclosure is to reduce the processing cost in improving the image quality of a virtual viewpoint image using a specific process.

Means for Solving the Problems

[0007] An image processing apparatus according to one embodiment of the present disclosure has the following configuration: an acquisition means for acquiring first information indicating the position and orientation of a virtual camera; a determination means for determining, based on the first information, a specific image to be used for generating color information of a virtual viewpoint image from among a plurality of captured images used for generating a virtual viewpoint image corresponding to the virtual camera; and a processing means for performing a specific process on at least a part of the specific image. [Effects of the Invention]

[0008] According to this disclosure, processing costs can be reduced when improving the image quality of virtual viewpoint images using specific processing. [Brief explanation of the drawing]

[0009] [Figure 1] This figure shows an example of the schematic configuration of the virtual viewpoint image generation system according to Example 1. [Figure 2] This figure shows an example of the hardware configuration of the rendering apparatus according to Example 1. [Figure 3] This figure shows an example of the functional configuration of the rendering apparatus according to Example 1. [Figure 4] This figure shows an example of a flowchart of the information processing method according to Example 1. [Figure 5] This figure shows the camera determination flow in the information processing method according to Example 1. [Figure 6] This diagram is intended to assist in explaining the determination of whether or not super-resolution processing is necessary in Example 1. [Modes for carrying out the invention]

[0010] <Embodiment> According to a preferred embodiment of the present disclosure, the image processing apparatus has acquisition means for acquiring first information indicating the position and orientation of a virtual camera. The image processing apparatus also has determination means for determining, based on the first information, a specific image from among a plurality of captured images used to generate a virtual viewpoint image corresponding to the virtual camera, to be used for generating color information of the virtual viewpoint image. The image processing apparatus also has processing means for performing specific processing on at least a portion of the specific image. The virtual camera is a camera placed in a virtual space and is also called a virtual viewpoint. The specific processing is processing to improve the image quality of the virtual viewpoint image, for example, super-resolution processing. Super-resolution processing is processing to increase the resolution of an image. In other words, it is processing to output an image with a higher resolution than the resolution of the input image. The specific processing is not limited to super-resolution processing and may be processing to remove noise or adjusting brightness.

[0011] In this embodiment, specific processing can be performed efficiently by selecting the captured images necessary for generating color information of the virtual viewpoint image and performing specific processing on them, thereby reducing processing costs.

[0012] Furthermore, the determination means can determine the specific captured image corresponding to the first region included in the virtual viewpoint image based on the first information. The first region is the region containing the subject included in the virtual viewpoint image. The first region may be a region showing only the subject included in the virtual viewpoint image, or it may be a region surrounding the subject. For example, by projecting a 3D model of the subject placed in virtual space onto a virtual camera, the region showing only the subject in the virtual viewpoint image can be extracted. For example, by utilizing existing subject recognition processing, the region surrounding the subject in the virtual viewpoint image can be determined. The specific captured image is the captured image used to generate the color information of the first region.

[0013] Furthermore, if the virtual viewpoint image includes multiple subjects, a first region can be set for each subject. In this configuration, an appropriate specific captured image to be used for generating color information of the virtual viewpoint image can be determined for each subject. Multiple first regions may be set for a single subject. In that case, an appropriate specific captured image can be determined according to the part of the subject. For example, even if a subject is not obscured by other subjects in the virtual viewpoint image, a part of the subject may be obscured by other subjects in the captured image. Therefore, if a specific captured image is determined for each subject, an captured image in which a part of the subject is obscured by other subjects may be determined as the specific captured image, and as a result, the image quality of the virtual viewpoint image may not be improved. Therefore, by determining an appropriate specific captured image according to the part of the subject, it is possible to appropriately improve the image quality of the virtual viewpoint image while reducing the processing cost associated with specific processing when photographing multiple subjects or when there are obstructions such as tools in the shooting area.

[0014] Furthermore, the determination means determines a second region of the specific captured image used to generate color information for a first region of a virtual viewpoint image corresponding to the virtual camera, based on the first information, and the processing means can perform specific processing on the second region. The second region is a region containing a subject included in the specific captured image. The second region may be a region showing only the subject included in the specific captured image, or it may be a region surrounding the subject.

[0015] In this embodiment, specific processing is performed on the region used to generate color information for a virtual viewpoint image in a specific captured image, thereby enabling more efficient processing of the specific processing and reducing the processing cost associated with that processing.

[0016] In addition, the acquisition means can acquire second information indicating the positions and orientations of a plurality of real cameras corresponding to a plurality of captured images including the specific captured image, and a 3D model of the subject generated based on the plurality of captured images. Further, the determination means can determine a component of the 3D model corresponding to the first region based on the first information, and determine the second region based on the component and the second information. Here, a real camera is a camera arranged in the real space. Note that the components of the 3D model differ depending on the type of the shape data of the 3D model. For example, when the 3D model is a mesh model, the components are polygons constituting the mesh model. For example, when the 3D model is a point cloud, the components are points constituting the point cloud. As described above, they differ depending on the type of the shape data of the 3D model. Note that the type of the shape data of the 3D model is not particularly limited.

[0017] According to this aspect, since the second region corresponding to the component of the 3D model can be determined, pixels to which specific processing is applied in a specific captured image can be specified more accurately.

[0018] In addition, the specific captured image is acquired from a real camera having a position and orientation close to the position and orientation of the virtual camera among the plurality of real cameras.

[0019] According to this aspect, since a captured image acquired from a real camera having a position and orientation close to the position and orientation of the virtual camera can be used for generating the color information of the virtual viewpoint image, a virtual viewpoint image that reproduces a more realistic appearance can be generated.

[0020] In addition, when the second region is smaller than the first region, the processing means can perform specific processing on the second region.

[0021] According to this aspect, while suppressing a change in the image quality of the virtual viewpoint image, the captured images to be subjected to specific processing can be reduced, and the processing cost can be reduced.

[0022] Furthermore, the processing means can perform specific processing on the second region when the number of pixels in the second region is less than the number of pixels in the first region.

[0023] Furthermore, the processing means can perform specific processing on the second region when the ratio of the number of pixels in the second region to the number of pixels in the first region is less than a threshold.

[0024] Furthermore, the image processing apparatus has a generation means for generating the virtual viewpoint image based on at least a portion of the specific captured image that has undergone specific processing by the processing means.

[0025] Furthermore, according to other preferred embodiments of the present disclosure, the image processing method includes an acquisition step of acquiring first information indicating the position and orientation of a virtual camera. The image processing method also includes a determination step of determining, based on the first information, a specific image from among a plurality of captured images used to generate a virtual viewpoint image corresponding to the virtual camera, to be used to generate color information for the virtual viewpoint image. The image processing method also includes a processing step of performing a specific process on at least a portion of the specific image.

[0026] Furthermore, according to other preferred embodiments of the present disclosure, the program causes the computer to function as the image processing apparatus described above.

[0027] <Examples> The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted.

[0028] <Example 1> In this embodiment, the rendering device 40 of the virtual viewpoint image generation system 1 achieves the aforementioned objective by selecting an captured image (corresponding to a specific captured image) to which super-resolution processing (corresponding to a specific processing) will be applied to each pixel or region of the virtual viewpoint image. The super-resolution processing in this embodiment uses a technique that trains a deep learning model and generates an image with improved resolution by interpolating the details of the image. However, the super-resolution processing is not limited to this; any method that processes a single image to achieve super-resolution is acceptable.

[0029] Figure 1 shows an example of a virtual viewpoint image generation system 1 according to Embodiment 1. The virtual viewpoint image generation system 1 consists of an imaging area 2, a subject 3, an imaging device 10, a modeling device 20, a database 30, a rendering device 40, an input device 50, and a display device 60. The imaging device 10 may be multiple imaging devices 10 or a single imaging device 10. In this embodiment, multiple imaging devices 10 (hereinafter, multiple imaging devices 10) will be used for explanation. In the virtual viewpoint image generation system 1, the subject 3, i.e., the foreground, located within the target area 2 is imaged by the multiple imaging devices 10. A virtual viewpoint image is then generated based on the images obtained by the multiple imaging devices 10.

[0030] The multiple imaging devices 10 are arranged to surround the subject 3, as shown in Figure 1, for example, and each captures the target area 2 from a different imaging position. Note that the multiple imaging devices 10 are cameras placed in real space (corresponding to actual cameras).

[0031] The modeling device 20 uses the captured image obtained by the imaging device 10 to generate three-dimensional shape information showing the three-dimensional shape of the subject 3, and stores it in the database 30 together with the captured image.

[0032] In addition to the three-dimensional shape information generated by the modeling device 20, the database 30 also records rendering information such as the field of view information from the imaging device 10 and the captured images.

[0033] The rendering device 40 (image processing device) uses data in the database 30 to generate (render) a virtual viewpoint image corresponding to the virtual viewpoint input from the input device 50. The rendering device 40 employs a viewpoint-dependent rendering method, primarily using images captured by the nearest real camera based on the position and orientation of the input virtual viewpoint.

[0034] The input device 50 is a device that inputs the position and orientation of a virtual viewpoint (equivalent to a virtual camera), and outputs the virtual viewpoint information to the rendering device 40.

[0035] The display device 60 displays the virtual viewpoint image generated by the rendering device 40. The display device 60 is, for example, a display, a tablet device, or a smartphone.

[0036] Viewpoint-dependent rendering is a method of generating color information for a subject model based on the position and orientation of a virtual viewpoint. For example, color information is determined based on an image captured by an imaging device with a line of sight close to that of the virtual viewpoint. For areas invisible from the imaging device closest to the virtual viewpoint's line of sight, color information is determined from the next closest imaging device. In this case, the color may be determined from a single imaging device, or it may be combined using some kind of weighting. Since the imaging device for determining the color information is selected depending on the position and orientation of the virtual viewpoint, the color information of the subject model changes as the virtual viewpoint moves. Therefore, in this embodiment, this method is called viewpoint-dependent rendering.

[0037] Furthermore, multiple imaging devices 10 can perform synchronized imaging continuously. In this case, the virtual viewpoint image generation system 1 can generate three-dimensional shape information of the subject in time series, and further generate a virtual viewpoint image in time series, i.e., a virtual viewpoint image.

[0038] Figure 2 is a block diagram showing an example of a computer hardware configuration applicable to the information processing device 100 according to this embodiment. The information processing device 100 includes a CPU 201, ROM 202, RAM 203, auxiliary storage device 204, communication I / F 205, and bus 206. The modeling device 20 can also be realized using similar hardware. Furthermore, the rendering device 40 may be composed of multiple information processing devices connected via a network, for example.

[0039] The CPU 201 controls the entire information processing device 100 using computer programs or data stored in the ROM 202 or RAM 203, thereby realizing the functions of each processing unit of the information processing device 100 shown in Figure 1. The information processing device 100 may also have one or more dedicated hardware components separate from the CPU 201, in which case at least a portion of the processing performed by the CPU 201 can be executed by the dedicated hardware. Examples of dedicated hardware include ASICs (Application-Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), DSPs (Digital Signal Processors), and GPUs (Graphics Processing Units).

[0040] ROM202 is memory that stores programs and other data that do not require modification. RAM203 is memory that temporarily stores programs or data supplied from the auxiliary storage device 204, and data supplied from external sources via the communication interface 205. The auxiliary storage device 204 is composed of storage such as a hard disk drive and stores various types of data such as image data or audio data. The communication interface 205 is used for communication between the information processing device 100 and external devices. For example, if the information processing device 100 is connected to an external device by wire, a communication cable is connected to the communication interface 205. If the information processing device 100 communicates wirelessly with an external device, the communication interface 205 is equipped with an antenna. The bus 206 connects the various parts of the information processing device 100 and transmits information.

[0041] Figure 3 shows an example of the configuration of the rendering apparatus 40 according to this embodiment 1. The rendering apparatus 40 includes a viewpoint information acquisition unit 310, a data acquisition unit 320, a camera determination unit 330, a conversion unit 340, and a synthesis unit 350.

[0042] The viewpoint information acquisition unit 310 acquires field of view information, including the position and orientation of the virtual viewpoint for generating a virtual viewpoint image, and the time code of the object to be drawn from the input device 50, and supplies the field of view information to the camera determination unit 330 and the conversion unit 340, and the time code to the data acquisition unit 320.

[0043] Angle of view information includes the focal length f and the center of the image (c). x , c y It consists of internal parameters such as ) and external parameters that indicate the position t and orientation R of the imaging device (including the actual camera and virtual camera). By using the internal and external parameters, the coordinates (u,v) on the image of the imaging device 10 where the three-dimensional coordinate X within the target area 2 is captured can be calculated using the following formula.

number

[0044] The data acquisition unit 320 acquires multiple captured images taken by the imaging device 10 with a time code supplied by the viewpoint information acquisition unit 310, along with three-dimensional shape information of the subject and field of view information of the imaging device 10, from the database 30 of the virtual viewpoint image generation system 1. The acquired field of view information and three-dimensional shape information are supplied to the camera determination unit 330 and the conversion unit 340, and the multiple captured images are supplied to the conversion unit 340. If the position and orientation of the imaging device do not change, the field of view information of the imaging device may be provided to the camera determination unit 330 and the conversion unit 340 in advance and stored by each unit, without going through the data acquisition unit.

[0045] The camera determination unit 330 uses the field of view information of the virtual viewpoint supplied from the viewpoint information acquisition unit 310, the field of view information of the imaging device 10 supplied from the data acquisition unit 320, and the three-dimensional shape information of the subject to determine the optimal imaging device for acquiring the color of each pixel in the virtual viewpoint image. It identifies the three-dimensional shape components of the subject corresponding to the pixels in the virtual viewpoint image, and among the real cameras that show those components, it determines the real camera with the closest angle difference from the virtual viewpoint as the assigned camera (corresponding to a specific imaging device).

[0046] The conversion unit 340 generates an image by converting the supplied image to match the field of view information of the virtual viewpoint supplied by the viewpoint information acquisition unit 310, based on the field of view information of the imaging device 10 supplied by the data acquisition unit 320, for each camera determined by the camera determination unit 330. In this case, if the virtual viewpoint image has a higher resolution, super-resolution processing is applied to generate a high-quality image. Since the conversion unit 340 includes super-resolution processing using machine learning, a pre-trained model is loaded. By inputting the captured image acquired from the assigned camera into this pre-trained model, a super-resolution processed captured image is obtained. Multiple such trained models may be maintained according to the magnification, and a trained model may be selected according to the magnification.

[0047] The image synthesis unit 350 synthesizes the captured images generated by the conversion unit 340 for each assigned camera to generate a single image, and supplies the generated image to the display device 60.

[0048] Figure 4 is a flowchart of the processing performed by the rendering device 40 in this embodiment. A high-resolution virtual viewpoint image is generated by the processing shown in Figure 4. The processing shown in Figure 4 can be achieved by the CPU 201 executing a control program that is stored in ROM 202 and read into RAM 203.

[0049] In S401, the viewpoint information acquisition unit 310 acquires the field of view information of the virtual viewpoint and the time code of the object to be drawn, in response to the input from the input device 50.

[0050] In S402, the data acquisition unit 320 acquires information necessary for drawing corresponding to the time code acquired in S401 from the database 30. The information necessary for drawing includes the field of view information of the imaging device 10, multiple captured images, and the three-dimensional shape information of the subject.

[0051] In S403, the camera determination unit 330 determines the assigned camera for each pixel of the virtual viewpoint image. The flow for determining the assigned camera is shown in Figure 5 and will be described later.

[0052] In S404, the conversion unit 340 converts the image for each assigned camera determined in S403 to match the virtual viewpoint image, and determines whether the resolution after conversion exceeds the resolution of the captured image. Here, the resolution can be determined by the distance from the actual camera or virtual viewpoint to the three-dimensional coordinates of the target and the focal length. If the resolution on the virtual viewpoint image is greater than the resolution at the time of capture, the process proceeds to S405; otherwise, it proceeds to S406. The processing in S404 will be explained in detail using Figure 6.

[0053] Figure 6 is a diagram illustrating the super-resolution processing according to Embodiment 1. Captured images 610 and 620 are used as captured images when generating the virtual viewpoint image 600. In S403, the assigned cameras were determined as follows: region 601 is the real camera that captured captured image 610, and regions 602 and 603 are the real cameras that captured captured image 620. Although shown as regions for illustrative purposes, each pixel may be assigned a corresponding camera. Region 601 will acquire color from region 611 of captured image 610. Here, comparing the resolution of region 601 on the virtual viewpoint image 600 with the resolution of region 611 on the corresponding captured image 610, region 611 is larger, therefore super-resolution processing is not performed on region 601. Similarly, for region 602, the resolution on the virtual viewpoint image 600 is greater than the resolution of region 622 on the corresponding captured image 620, so super-resolution processing is not performed. On the other hand, region 603 on the virtual viewpoint image 600 has a higher resolution than region 623 on the corresponding captured image 620. Therefore, super-resolution processing will be performed on region 623, which corresponds to region 603.

[0054] Furthermore, if the resolution of the virtual viewpoint image 600 is only slightly higher than the resolution of the captured image, it is conceivable that super-resolution will have little effect. Therefore, a condition may be set that the ratio of the resolution of the virtual viewpoint image 600 to the resolution of the captured image is greater than or equal to a predetermined threshold. However, the condition is not limited to this, and the ratio of the resolution of the virtual viewpoint image 600 to the resolution of the captured image may also be less than the threshold.

[0055] In S405, the conversion unit 340 applies super-resolution processing to the area around the target pixel of the assigned camera. As mentioned above, super-resolution processing is performed on the pixel (region 623) that was determined to be subject to super-resolution processing in S404. Since super-resolution processing is difficult for a single pixel, the surrounding area is processed together. The area to be processed together is the minimum area necessary for super-resolution processing. The minimum necessary area is set according to the trained model that performs super-resolution processing.

[0056] In the S406, for each camera, the corresponding pixels in the captured image or an image obtained by super-resolution of the captured image are projected onto a virtual viewpoint for the pixels that the camera is responsible for.

[0057] In S407, the synthesis unit 350 generates a virtual viewpoint image by combining the projection images generated in S406.

[0058] Figure 5 is a flowchart showing the process of determining which camera is responsible for each pixel in the virtual viewpoint image.

[0059] Select the pixels on the virtual viewpoint image to be processed by S501. This can be done by processing the image sequentially using a raster scan, or by dividing it into blocks.

[0060] In S502, the three-dimensional coordinates corresponding to the pixels selected in S501 are obtained by calculating the position where a line passing through the virtual viewpoint's position and the three-dimensional coordinates corresponding to the selected pixels intersects with the three-dimensional shape of the subject. If the line does not intersect with the three-dimensional shape, that pixel is excluded from processing because it has no foreground.

[0061] In S503, it selects a real camera for which calculations have not yet been completed for the selected pixels.

[0062] In S504, visibility is determined by checking whether the position obtained by projecting the three-dimensional coordinates determined in S502 onto the actual camera selected in S503 is within the image and not hidden by other parts of the three-dimensional shape of the subject. For actual cameras that are not visible, the pixels are excluded from subsequent processing or their evaluation value is set to 0.

[0063] In S505, an evaluation value is calculated for the selected pixel from the camera selected in S503 to extract color. The evaluation value E is determined by the following equation, based on the vector connecting the virtual viewpoint and the object's three-dimensional coordinates, the angle θ between the vector connecting the selected camera and the object's three-dimensional coordinates, and the resolution R of the selected camera in the object's three-dimensional coordinates. E = αθ + be - βR α and β are weights that determine how much priority each of them should be given.

[0064] In S506, it is determined whether all cameras have been processed for the selected pixels. If there are still cameras that have not been processed, the process returns to S503 and proceeds to process the next camera.

[0065] In S507, it is determined whether all pixels of the virtual viewpoint image have been processed. If there are pixels that have not yet been processed, the process returns to S501 and the next pixel is processed.

[0066] In S508, the camera responsible for each pixel is determined based on the calculated evaluation value. The camera with the highest evaluation value for each pixel is selected as the responsible camera. At this time, since the camera selected for each pixel may vary, color differences due to the assigned camera may become noticeable in the generated virtual viewpoint image. Therefore, filtering may be applied to suppress the variation in assigned cameras.

[0067] Furthermore, since the three-dimensional shape is a point cloud, the visibility of each point cloud can be calculated in advance. In that case, S503 selects from cameras with good visibility, and S504 becomes unnecessary.

[0068] As described above, by determining the camera responsible for each pixel in the virtual viewpoint image and then performing super-resolution processing as needed, it is possible to generate high-definition virtual viewpoint images while suppressing the processing costs associated with super-resolution processing.

[0069] <Example 2> In Example 1, one camera was assigned to each pixel. However, if only one camera is assigned to each pixel, the color may change where the assigned camera changes, resulting in a mottled pattern. Therefore, in this example, this problem is solved by assigning multiple cameras to each pixel. The configuration of the image processing device is the same as in Figure 3.

[0070] In Example 1, one camera was selected in step S403 in Figure 4 or step S507 in Figure 5, but in this example, multiple cameras are selected based on evaluation values. The top n cameras based on evaluation values ​​may be selected, or a threshold of 1 / n of the highest evaluation value may be used.

[0071] Furthermore, when compositing with S407, a weighted average can be used to generate smooth images where the switching between cameras is less noticeable. The weights used in this process are the evaluation values ​​mentioned above.

[0072] When you have multiple cameras to manage, you can perform super-resolution processing on all of them, but you can reduce processing costs by performing it on only one representative camera. In this case, by increasing the weight of the camera that underwent super-resolution processing and combining the images, you can combine them without losing the effect of the super-resolution processing. The representative camera can be the camera with the highest evaluation score, or the visible light camera with the highest resolution for the relevant area.

[0073] Furthermore, in the S404 assessment, if it is determined that super-resolution processing is not necessary for even one of the multiple cameras in charge, the processing cost can be further reduced by not performing super-resolution processing on the image corresponding to that pixel.

[0074] Up to this point, we have explained how to determine the camera to be used and how to perform super-resolution processing on a pixel-by-pixel basis. However, since super-resolution processing requires an image of a certain size, processing it as a region rather than on a pixel-by-pixel basis can sometimes further reduce computational costs.

[0075] As described above, by determining which camera is responsible for which image area rather than individual pixels and performing super-resolution processing accordingly, it becomes unnecessary to perform super-resolution processing multiple times on a single camera, thereby reducing processing costs. Furthermore, since it eliminates the possibility of a single pixel within an area being handled by a different camera, it also prevents color differences between cameras from appearing in the generated image.

[0076] Although this disclosure has been described above based on several embodiments, this disclosure is not limited to the above embodiments, and various modifications are possible in accordance with the spirit of this disclosure, and these modifications are not excluded from the scope of this disclosure.

[0077] Furthermore, in this embodiment, some or all of the control may be provided to an image processing system, etc., via a network or various storage media, by supplying a computer program that realizes the functions of the embodiment described above. The computer (or CPU, MPU, etc.) in the image processing system, etc., may then read and execute the program. In that case, the program and the storage medium storing the program constitute the present disclosure.

[0078] Furthermore, the disclosure of this embodiment includes the following configuration, method, and program.

[0079] (Composition 1) An acquisition means for acquiring first information indicating the position and orientation of a virtual camera, A determination means for determining, based on the first information, a specific image to be used for generating color information of the virtual viewpoint image from among a plurality of captured images used for generating a virtual viewpoint image corresponding to the virtual camera, Processing means for performing specific processing on at least a portion of the specific captured image, An image processing apparatus characterized by having

[0080] (Configuration 2) The image processing apparatus according to configuration 1, characterized in that the determination means determines a specific captured image corresponding to a first region included in the virtual viewpoint image based on the first information.

[0081] (Composition 3) The image processing apparatus according to configuration 2, characterized in that the specific captured image is an captured image used to generate color information for the first region.

[0082] (Composition 4) The image processing apparatus according to configuration 2, characterized in that the first region is a region that includes a subject included in the virtual viewpoint image.

[0083] (Composition 5) Based on the first information, the determination means determines the second region of the specific captured image used to generate color information for the first region of the virtual viewpoint image corresponding to the virtual camera. The image processing apparatus according to configuration 2, characterized in that the processing means performs a specific process on the second region.

[0084] (Composition 6) The image processing apparatus according to configuration 5, characterized in that the second region is a region containing a subject included in the specific captured image.

[0085] (Composition 7) The acquisition means acquires second information indicating the position and orientation of a plurality of real cameras corresponding to a plurality of captured images including the specific captured image, and a 3D model of the subject generated based on the plurality of captured images. The image processing apparatus according to configuration 6, characterized in that the determination means determines the components of the 3D model corresponding to the first region based on the first information, and determines the second region based on the components and the second information.

[0086] (Composition 8) The image processing apparatus according to configuration 7, characterized in that the specific captured image is acquired from one of the plurality of real cameras, which is located at a position and orientation close to the position and orientation of the virtual camera.

[0087] (Composition 9) The image processing apparatus according to configuration 5, characterized in that the processing means performs a specific processing on the second region when the second region is smaller than the first region.

[0088] (Composition 10) The image processing apparatus according to configuration 5, characterized in that the processing means performs a specific process on the second region when the number of pixels in the second region is less than the number of pixels in the first region.

[0089] (Composition 11) The image processing apparatus according to configuration 5, characterized in that the processing means performs a specific process on the second region when the ratio of the number of pixels in the second region to the number of pixels in the first region is less than a threshold.

[0090] (Composition 12) The image processing apparatus according to configuration 1, characterized in that the aforementioned specific processing is super-resolution processing.

[0091] (Composition 13) The image processing apparatus according to configuration 1, characterized in that it has a generation means for generating a virtual viewpoint image based on at least a portion of the specific captured image that has undergone specific processing by the processing means.

[0092] (method) An acquisition step to acquire first information indicating the position and orientation of a virtual camera, A determination step in which, among a plurality of captured images used to generate a virtual viewpoint image corresponding to the virtual camera, a specific captured image used to generate the color information of the virtual viewpoint image is determined based on the first information, A processing step which involves performing a specific process on at least a portion of the aforementioned specific captured image, An image processing method characterized by having the following features.

[0093] (program) A program for causing a computer to function as an image processing device as described in any one of the configurations 1 to 13. [Explanation of Symbols]

[0094] 40 Rendering equipment 320 Viewpoint Information Acquisition Unit 330 Camera Determination Unit 340 Conversion Unit 350 Synthesis part

Claims

1. An acquisition means for acquiring first information indicating the position and orientation of a virtual camera, A determination means for determining, based on the first information, a specific image from among a plurality of images used to generate a virtual viewpoint image corresponding to the virtual camera, to be used to generate the color information of the virtual viewpoint image, Processing means for performing specific processing on at least a portion of the specific captured image, An image processing apparatus characterized by having

2. The image processing apparatus according to claim 1, characterized in that the determination means determines a specific captured image corresponding to a first region included in the virtual viewpoint image based on the first information.

3. The image processing apparatus according to claim 2, characterized in that the specific captured image is an captured image used for generating color information of the first region.

4. The image processing apparatus according to claim 2, characterized in that the first region is a region that includes a subject included in the virtual viewpoint image.

5. The determination means determines, based on the first information, a second region of the specific captured image used to generate color information of a first region of a virtual viewpoint image corresponding to the virtual camera, The image processing apparatus according to claim 2, characterized in that the processing means performs a specific processing on the second region.

6. The image processing apparatus according to claim 5, characterized in that the second region is a region containing a subject included in the specific captured image.

7. The acquisition means acquires second information indicating the position and orientation of a plurality of real cameras corresponding to a plurality of captured images including the specific captured image, and a 3D model of the subject generated based on the plurality of captured images. The image processing apparatus according to claim 6, characterized in that the determination means determines the components of the 3D model corresponding to the first region based on the first information, and determines the second region based on the components and the second information.

8. The image processing apparatus according to claim 7, characterized in that the specific captured image is acquired from one of the plurality of real cameras, which is located at a position and orientation close to the position and orientation of the virtual camera.

9. The image processing apparatus according to claim 5, characterized in that the processing means performs a specific processing on the second region when the second region is smaller than the first region.

10. The image processing apparatus according to claim 5, characterized in that the processing means performs a specific process on the second region when the number of pixels in the second region is less than the number of pixels in the first region.

11. The image processing apparatus according to claim 5, wherein the processing means performs a specific process on the second region when the ratio of the number of pixels in the second region to the number of pixels in the first region is less than a threshold.

12. The image processing apparatus according to claim 1, characterized in that the aforementioned specific processing is super-resolution processing.

13. The image processing apparatus according to claim 1, further comprising a generation means for generating a virtual viewpoint image based on at least a portion of the specific captured image that has undergone specific processing by the processing means.

14. An acquisition step to acquire first information indicating the position and orientation of a virtual camera, A determination step in which, among a plurality of captured images used to generate a virtual viewpoint image corresponding to the virtual camera, a specific captured image used to generate color information for the virtual viewpoint image is determined based on the first information, A processing step which involves performing a specific process on at least a portion of the aforementioned specific captured image, An image processing method characterized by having the following features.

15. A program for causing a computer to function as an image processing device according to any one of claims 1 to 13.