Image processing device, control method and program for image processing device

The image processing device adjusts virtual object attributes to match real-world lighting and resolution, addressing unnaturalness in composite images by synchronizing virtual and real object characteristics.

JP7775012B2Active Publication Date: 2025-11-25CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021166499
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-11
Filing Date
2021-10-08
Publication Date
2025-11-25
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

Existing technologies fail to account for changes in lighting conditions and resolution when combining virtual and real objects in composite images, leading to unnaturalness in the composite image.

Method used

An image processing device adjusts the saturation, brightness, and resolution of virtual objects based on texture data from real objects captured by multiple physical cameras to match the lighting conditions and resolution of the real environment, generating a more natural composite image.

Benefits of technology

The device effectively generates a virtual viewpoint image with virtual objects that seamlessly integrate with real objects, reducing unnaturalness and enhancing the overall image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007775012000001
    Figure 0007775012000001
  • Figure 0007775012000002
    Figure 0007775012000002
  • Figure 0007775012000003
    Figure 0007775012000003
Patent Text Reader

Abstract

To appropriately create a virtual viewpoint image including a virtual object.SOLUTION: An image processing apparatus acquires one or more images based on photographing performed by one or more photographing devices, acquires information on a virtual object, and on the basis of the one or more images and the information on the virtual object, creates a two-dimensional image including the virtual object. In the creation of the two-dimensional image, the image processing apparatus determines color information on the virtual object on the basis of color information on a real object included in the one or more images.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device that generates a virtual viewpoint image, a control method for the image processing device, and a program. [Background technology]

[0002] There is a technology that uses images captured by multiple physical cameras (hereinafter referred to as multi-viewpoint images) to reproduce an image (hereinafter referred to as a virtual viewpoint image) from a camera (hereinafter referred to as a virtual camera) virtually placed in a three-dimensional space. There is also a technology that superimposes a computer graphics (hereinafter referred to as a CG) image, which is generated by capturing a virtual object with a virtual camera, on such a virtual viewpoint image. For example, a virtual object, such as a virtual signboard for advertising purposes, is placed in a three-dimensional space (hereinafter referred to as a virtual space) obtained by reconstructing the captured space based on the multi-viewpoint images. Then, by capturing an image of this virtual space with a virtual camera, it becomes possible to superimpose and draw a virtual advertisement (CG image) on the virtual viewpoint image.

[0003] When a virtual viewpoint image based on a photographed image is superimposed on a CG image generated independently of the photographed image and displayed, unnaturalness may occur, such as the CG image appearing to float in the virtual viewpoint image. Patent Document 1 discloses a configuration for generating a more natural composite image of a photographed image and a CG image, in which noise processing is performed to estimate and add noise generated in a photographed image of a virtual object, and the noise-processed CG image is superimposed on the photographed image to generate the composite image. According to Patent Document 1, unnaturalness in the composite image is reduced by matching the noise appearance of the photographed image and the CG image. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-203326 Summary of the Invention [Problem to be solved by the invention]

[0005] However, Patent Document 1 does not take into consideration matching the color and resolution of a virtual object with the color and resolution of a real object, which change depending on the lighting conditions in the shooting space, the shooting conditions of a physical camera, etc. As a result, unnaturalness still occurs in a composite image of a virtual viewpoint image generated based on a shot image and a CG image of a virtual object.

[0006] An object of the present disclosure is to appropriately generate a virtual viewpoint image including a virtual object. [Means for solving the problem]

[0007] An image processing device according to an aspect of the present disclosure has the following configuration. an acquisition means for acquiring a two-dimensional image including a virtual object; a processing means for processing the saturation and / or brightness of the virtual object in the two-dimensional image so that it matches or approaches the saturation and / or brightness included in color information corresponding to the three-dimensional shape of a real object generated based on a plurality of images; With death, The processing means processes the resolution of the virtual object in the two-dimensional image when a difference between the resolution of the virtual object in the two-dimensional image and the resolution of the real object is greater than a predetermined value. It is characterized by: [Effects of the Invention]

[0008] According to the present disclosure, a virtual viewpoint image including a virtual object can be appropriately generated. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 2A is a block diagram showing an example of the configuration of an image processing system, and FIG. 2B is a block diagram showing an example of the hardware configuration of an image processing device. [Figure 2] FIG. 1 is a block diagram showing an example of the functional configuration of an image processing apparatus. [Figure 3] 3A to 3C are diagrams illustrating generation of a composite image according to the first embodiment. [Figure 4]5 is a flowchart showing an example of a synthetic image generation process according to the first embodiment. [Figure 5] 5 is a flowchart showing an example of a CG image processing process according to the first embodiment. [Figure 6] 10A to 10C are diagrams illustrating generation of a composite image according to a second embodiment. [Figure 7] 10 is a flowchart showing an example of a CG image processing process according to the second embodiment. [Figure 8] 10A to 10C are diagrams illustrating generation of a composite image according to a third embodiment. [Figure 9] 10 is a flowchart showing an example of a CG image processing process according to the third embodiment. [Figure 10] FIG. 10 is a block diagram showing an example of the arrangement of an image processing system according to a fourth embodiment. [Figure 11] FIG. 10 is a block diagram showing an example of the functional arrangement of an image processing apparatus according to a fourth embodiment. [Figure 12] 10A to 10C are diagrams illustrating generation of a composite image according to the fourth embodiment. [Figure 13] 10 is a flowchart showing an example of a synthetic image generation process according to the fourth embodiment. [Figure 14] 10 is a flowchart showing an example of a synthetic image processing process according to the fourth embodiment. [Figure 15] FIG. 11 is a block diagram showing an example of the arrangement of an image processing system according to a fifth embodiment. [Figure 16] FIG. 11 is a block diagram showing an example of the functional arrangement of an image processing apparatus according to a fifth embodiment. [Figure 17] 13 is a flowchart showing an example of a CG image generation process according to the fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0011] In the following embodiments, a virtual viewpoint image is an image generated by a user and / or a dedicated operator freely manipulating the position and orientation of a virtual camera, and is also called a free viewpoint image, an arbitrary viewpoint image, etc. Unless otherwise specified, the term "image" will be explained as including the concepts of both moving images and still images.

[0012] (First embodiment) Real objects depicted in virtual viewpoint images, such as signs that exist in the real world, become brighter or darker depending on factors such as weather, lighting, and shadows from obstructions. On the other hand, virtual objects depicted in CG images, such as virtual signs, cannot be brightened in response to changes in the surrounding environment unless they are rendered by simulating such lighting conditions. However, faithfully simulating constantly changing lighting conditions using computer graphics is extremely difficult. A composite image obtained by combining a virtual viewpoint image, whose brightness varies depending on the lighting conditions of the shooting space, with a CG image that does not experience such changes may appear unnatural to the viewer.

[0013] In the first embodiment, a configuration will be described in which, when generating a composite image in which a CG image is superimposed on a virtual viewpoint image generated from a plurality of images from a plurality of viewpoints (hereinafter, multi-viewpoint images), a more natural composite image is generated by adjusting the color of the CG image. The color of the CG image can be adjusted, for example, by processing the saturation and / or brightness of the CG image based on texture data (i.e., color information) of a subject (real object) generated using the multi-viewpoint images used to generate the virtual viewpoint image.

[0014] <Image processing system hardware configuration> 1(a) is a diagram showing an example of the configuration of an image processing system 10 according to the first embodiment. The image processing system 10 includes an imaging system 101, an image processing device 102, and an information processing device 103, and generates a virtual viewpoint image, a CG image, and a composite image in which the virtual viewpoint image and the CG image are superimposed.

[0015] The imaging system 101 has multiple imaging devices (hereinafter also referred to as physical cameras) arranged at different positions, and performs synchronized imaging using the multiple physical cameras (in this embodiment, simultaneous imaging using the multiple physical cameras). The imaging system 101 simultaneously captures images of a three-dimensional space using the multiple physical cameras, thereby obtaining multi-viewpoint images made up of multiple captured images, and transmits these images to the image processing device 102.

[0016] The image processing device 102 generates a virtual viewpoint image captured by a virtual camera based on the multiple viewpoint images. The image processing device 102 also generates a CG image obtained by capturing an image of a virtual object placed in a three-dimensional space with the virtual camera. The image processing device 102 then processes the CG image based on information (texture) used to render the real object. More specifically, the image processing device 102 processes the CG image based on color information of the real object that is close to the virtual object to match the colors of the real object and the virtual object, and generates a composite image by superimposing the processed CG image and the virtual viewpoint image. By processing the CG image in this way, a CG image whose color changes depending on the capturing conditions of the physical camera included in the capturing system 101 and the lighting conditions of the capturing space can be generated, resulting in a more natural composite image. The viewpoint of the virtual camera is expressed by camera parameters determined by the information processing device 103, which will be described later. The image processing device 102 transmits the generated composite image to the information processing device 103.

[0017] The information processing device 103 includes a controller for controlling the virtual camera (viewpoint) and a display unit for displaying composite images, etc. The controller includes common devices for user input operations, such as a keyboard and a mouse, as well as a joystick, knob, jog dial, etc. for controlling the position and orientation of the virtual camera. The display unit includes one or more display devices (hereinafter referred to as "monitors") and displays information required by the user. If a touch panel display is used as the display device, the touch panel may also serve as part or all of the controller. A UI screen for controlling the virtual camera is displayed on the monitor. The user can specify the amount of operation of the virtual camera, such as the movement direction, orientation (orientation), rotation, movement distance, and movement speed, while viewing the monitor display. The information processing device 103 determines camera parameters indicating the position, orientation, or zoom of the virtual camera from the amount of operation specified by the user via the controller and transmits them to the image processing device 102. The determined parameters may be displayed on the monitor as the state of the virtual camera. The information processing device 103 also receives a composite image generated by the image processing device 102 and displays it on the monitor.

[0018] FIG. 1B is a diagram illustrating an example of the hardware configuration of the image processing device 102. The image processing device 102 includes a CPU 111, a RAM 112, a ROM 113, and a communication unit 114. The CPU 111 is a processor that uses the RAM 112 as a work memory, executes programs stored in the ROM 113, and comprehensively controls each component of the image processing device 102. The CPU 111 executes various programs to realize each functional unit shown in FIG. 2 (described below). The RAM 112 temporarily stores computer programs read from the ROM 113 and intermediate calculation results. The ROM 113 holds computer programs and data that do not require modification. Data held in the ROM 113 includes camera parameters of a physical camera, 3D data of a background model and virtual objects (described below), and the like. The communication unit 114 has communication means such as Ethernet or USB, and communicates with the image capture system 101 and the information processing device 103.

[0019] <Functional configuration of image processing device 102> Fig. 2 is a block diagram showing an example of the functional configuration of the image processing device 102. Fig. 2 shows an example of the functional configuration for realizing a process of processing a CG image to become a CG image according to the shooting conditions of the physical camera and the lighting conditions of the shooting space, and generating a composite image by superimposing the processed CG image on a virtual viewpoint image. The image processing device 102 of this embodiment has, as its functional configuration, a communication control unit 201, a virtual viewpoint image generation unit 202, a CG image generation unit 203, a CG image processing unit 204, and a composite image generation unit 205.

[0020] The communication control unit 201 receives multiple viewpoint images from the imaging system 101 and receives information on camera parameters of the virtual camera from the information processing device 103, using the communication unit 114. The communication control unit 201 outputs the received multiple viewpoint images to the virtual viewpoint image generation unit 202 and outputs the received camera parameters to the virtual viewpoint image generation unit 202 and the CG image generation unit 203. The communication control unit 201 also inputs a composite image from the composite image generation unit 205 and transmits it to the information processing device 103 via the communication unit 114. The virtual viewpoint image generation unit 202 generates a virtual viewpoint image based on the multiple viewpoint images and camera parameters of the virtual camera received from the communication control unit 201, and the camera parameters of the physical camera stored in advance in ROM 113.

[0021] The virtual viewpoint image generation unit 202 generates a virtual viewpoint image, for example, using the following method. First, the virtual viewpoint image generation unit 202 acquires a foreground image in which a foreground region corresponding to a predetermined object, such as a person or a ball, is extracted from the multiple viewpoint images, and a background image in which a background region other than the foreground region is extracted from the multiple viewpoint images. Based on the foreground image and the camera parameters of the physical camera, the virtual viewpoint image generation unit 202 generates a foreground model representing the three-dimensional shape of the predetermined object, information about the position of the foreground model in the virtual space, and texture data for rendering the surface of the foreground model. The virtual viewpoint image generation unit 202 also generates texture data for rendering the surface of a background model representing the three-dimensional shape of a background object, such as a stadium, pre-stored in ROM 113, based on the background image. The virtual viewpoint image generation unit 202 generates a virtual viewpoint image by mapping the texture data to the foreground model and the background model and rendering the virtual space according to the camera parameters of the virtual camera. However, the method for generating a virtual viewpoint image is not limited to this, and various methods can be used, such as generating a virtual viewpoint image by projective transformation of a captured image without using a three-dimensional model.

[0022] A foreground image is an image in which a region of a predetermined object (foreground region) is extracted from a multi-viewpoint image. A predetermined object is, for example, a dynamic object (moving body) that moves (its absolute position and shape may change) when photographed from a fixed position. Examples of such objects include people such as players and umpires on a field where a sport is played, a ball used in a ball game, and singers, musicians, performers, and presenters in concerts and entertainment events. A background image is an image of at least a region (background region) different from the foreground object. More specifically, a background image refers to an object (immobile body) that remains stationary or nearly stationary when photographed from a fixed position. Examples of such objects include a stage for a concert, a stadium where an event such as a competition is held, a structure such as a goal used in a ball game, or a field.

[0023] The virtual viewpoint image generation unit 202 sends the virtual viewpoint image generated as described above to the CG image processing unit 204 and the composite image generation unit 205. The virtual viewpoint image generation unit 202 also sends intermediate data (for example, the positions of each foreground object in the virtual space, their texture data, etc.) generated when generating the multiple viewpoint images and the virtual viewpoint image to the CG image processing unit 204.

[0024] Based on the camera parameters of the virtual camera received from the communication control unit 201, the CG image generation unit 203 performs processes such as photographing the virtual object (projecting the virtual object onto the image), rasterizing, and color determination to generate a CG image. A virtual object is an object that has 3D data (mesh data, texture data, material data, etc.) and exists only in virtual space. The CG image generation unit 203 acquires the 3D data of the virtual object stored in advance in the ROM 113, places the virtual object at an arbitrary position in the virtual space, and generates an image viewed from the virtual camera. The virtual object is, for example, a virtual signboard that does not exist in the shooting space of the physical camera, and a virtual advertising image (CG image) is generated by photographing it with the virtual camera. The CG image generation unit 203 sends the generated CG image and the 3D data of the virtual object used to generate the CG image to the CG image processing unit 204.

[0025] The CG image processing unit 204 processes the CG image received from the CG image generation unit 203, using the virtual viewpoint image, multiple viewpoint images, and intermediate data sent from the virtual viewpoint image generation unit 202, and the 3D data of the virtual object sent from the CG image generation unit 203. The CG image processing unit 204 sends the processed CG image to the composite image generation unit 205. A method of processing CG images will be described later with reference to FIG. 3. The composite image generation unit 205 generates a composite image in which the foreground, background, and CG are rendered by combining the virtual viewpoint image input from the virtual viewpoint image generation unit 202 and the CG image input from the CG image processing unit 204. The composite image generation unit 205 outputs the generated composite image to the communication control unit 201.

[0026] <Explanation of the method for processing CG images and generating composite images> Using FIG. 3, the procedure for processing CG images and generating composite images according to the first embodiment will be described. Here, as an example of the processing of CG images by the CG image processing unit 204, a mode of changing the color tone of a CG image based on the texture (color information) of a real object in a virtual viewpoint image will be described. FIG. 3(a) is a schematic diagram showing an overhead view of the state of photographing a foreground object 301, a background object 302, and a virtual object 303 that does not exist in the photographing space by a virtual camera 304 on the photographing space of the photographing system 101. FIG. 3(b) is a schematic diagram explaining a method of synthesizing the virtual viewpoint image 312 and the CG image 313 generated in FIG. 3(a).

[0027] The physical cameras 305 and 306 included in the photographing system 101 photograph the foreground object 301 and the background object 302, and generate photographed images 307 and 308. Since the foreground object 301 is illuminated by a light source 309 such as the sun or a light, the photographed image �07 is a dark image due to backlighting, and the photographed image 308 is a bright image due to frontlighting. The image processing apparatus 102 generates a foreground model 310 and a background model 311 in which the foreground object 301 and the background object 302 are reconstructed in a virtual space based on a plurality of viewpoint images photographed by the photographing system 101. Then, based on the camera parameters of the virtual camera 304 received from the information processing apparatus 103, the image processing apparatus 102 generates a virtual viewpoint image 312 obtained by photographing the foreground model 310 and the background model 311, and a CG image 313 obtained by photographing the virtual object 303.

[0028] When generating the texture data of the foreground model 310, in order to make the image closer to the image captured by the virtual camera, the image processing apparatus 102 preferentially uses the texture of the captured image of the physical camera 305 close to the virtual camera 304. Then, the foreground model 310 reflected in the virtual viewpoint image 312 becomes a darker image than the captured image 308. On the other hand, the CG image 313 is not affected by the lighting conditions in the shooting space of the shooting system 101 unless the light source 309 is simulated in the virtual space, and the color of the virtual object 303 does not change even if the virtual camera 304 is moved. Therefore, if the virtual viewpoint image 312 and the CG image 3 are directly synthesized, the colors of the foreground model 310 and the virtual object 303 may be different, resulting in an unnatural synthesized image. It is possible to simulate the lighting in the shooting space of the shooting system 101 when generating the CG image 313. However, when there are a large number of illuminations such as the sun or the lights in the stadium, the image processing apparatus 102 requires a large amount of processing time to simulate them. In addition, since the sun, the lights on the stage, etc. change in color and intensity every moment, it is difficult to faithfully simulate them. Therefore, the image processing apparatus 102 according to the first embodiment generates a CG image 314 in which the saturation and brightness of the CG image 313 are changed based on the texture data of the foreground model 310 in the virtual viewpoint image 312 in order to generate a natural synthesized image. By synthesizing the CG image 314 obtained by processing the CG image 313 as described above with the virtual viewpoint image 312, it is possible to generate a more natural synthesized image 315 of the virtual viewpoint image and the CG image. <CG image processing and control of generating synthesized image>

[0029] Fig. 4 is a flowchart showing the generation of a CG image and a virtual viewpoint image and the synthesis process thereof according to this embodiment. The process shown in Fig. 4 is realized, for example, by reading a control program stored in the ROM 113 into the RAM 112 and executing it with the CPU 111. The communication control unit 201 receives multi-viewpoint images from the imaging system 101 and receives camera parameters of the virtual camera from the information processing device 103. When the virtual viewpoint image generation unit 202 and the CG image generation unit 203 receive this received data from the communication control unit 201, the process shown in Fig. 4 starts.

[0030] In S401, the virtual viewpoint image generation unit 202 acquires multiple images from multiple image capture devices (multiple physical cameras). In this embodiment, multiple viewpoint images (multiple captured images captured by multiple image capture devices) are used as the multiple images, but the present invention is not limited to this. For example, the multiple images may include texture images of foreground objects (partial images in which the foreground object area is extracted from the captured image) obtained from each of the multiple captured images. Alternatively, the multiple images may include texture images of background objects (partial images in which the background object area is extracted from the captured image) obtained from each of the multiple captured images. In these cases, the image capture system 101 extracts foreground objects and background objects from each of the captured images, and the extracted images are sent to the image processing device 102 as multiple images. When the image processing device 102 receives multiple viewpoint images, the extraction of foreground objects and background objects from each of the multiple captured images constituting the multiple viewpoint images may be performed, for example, by the virtual viewpoint image generation unit 202. In S402, the CG image generation unit 203 acquires information about virtual objects. The information about the virtual object includes at least one of the three-dimensional shape data of the virtual object (for example, polygon data or voxel data) and the placement position of the virtual object in the virtual space (for example, the position set by the user (operator)).

[0031] In S403, the virtual viewpoint image generation unit 202 generates a virtual viewpoint image captured by a virtual camera based on the plurality of viewpoint images acquired in S401 and the camera parameters of the virtual camera (the position and orientation of the virtual viewpoint). The camera parameters of the virtual camera are input from the communication control unit 201. That is, an image representing the real space observed from the specified virtual viewpoint is generated as the virtual viewpoint image. The generated virtual viewpoint image is sent to the composite image generation unit 205. Also, intermediate data during the generation of the virtual viewpoint image (for example, the position of each foreground object in the virtual space, their texture data, etc.) is sent to the CG image processing unit 204. In S404, the CG image generation unit 203 generates a CG image obtained by capturing the virtual object from the virtual camera based on the information on the virtual object acquired in S402 and the camera parameters of the virtual camera. The camera parameters of the virtual camera are those used in S403 as well and are input from the communication control unit 201. That is, an image representing the virtual object observed from the virtual viewpoint used for generating the above virtual viewpoint image is generated as the CG image. The generated CG image is sent to the CG image processing unit 204.

[0032] In S405, the CG image processing unit 204 processes the CG image input from the CG image generation unit 203 based on the intermediate data input from the virtual viewpoint image generation unit 202. The details of the CG image processing in S405 will be described later by referring to FIG. 5. The processed CG image is output to the composite image generation unit 205. In S406, the composite image generation unit 205 synthesizes the virtual viewpoint image input from the virtual viewpoint image generation unit 202 and the processed CG image input from the CG image processing unit 204 to generate a composite image. Then, this process ends. After the end of this process, the generated composite image is transmitted from the composite image generation unit 205 to the information processing device 103 via the communication control unit 201. The above is the composite image processing of the virtual viewpoint image and the CG image according to the first embodiment.

[0033] <Explanation of CG Image Processing> FIG. 5 is a flowchart showing an example of CG image processing according to this embodiment, illustrating details of the CG image processing in S405 of FIG. 4. In the CG image processing according to the first embodiment, color information of the CG image is determined and processed based on texture data (saturation and brightness) of a foreground model close to the virtual object in three-dimensional space. More specifically, the CG image processing unit 204 acquires the saturation and brightness of a foreground model close to the position where the virtual object is placed in virtual space, and changes the saturation and brightness of the virtual object to match those values. The saturation and brightness of the CG image before processing are calculated by the CG image generation unit 203 based on the texture data and material data of the virtual object. The processing shown in FIG. 5 is executed by the CG image processing unit 204. The processing shown in FIG. 5 is triggered by the acquisition of intermediate data from the virtual viewpoint image generation unit 202 and the acquisition of a CG image from the CG image processing unit 204.

[0034] In S501, the CG image processing unit 204 selects a foreground model that is closest to the virtual object from the intermediate data input from the virtual viewpoint image generation unit 202 and acquires its texture data. The foreground model that is closest to the virtual object is, for example, the foreground model of the real object that is closest to the virtual object. Note that the foreground model selected in S501 is not limited to this. For example, foreground models of a predetermined number of real objects may be selected in order of proximity to the virtual object, or foreground models of all real objects that exist within a predetermined distance from the virtual object may be selected. That is, when processing a CG image, if there are multiple foreground models located near the virtual object, the CG image may be processed based on not only one foreground model but also multiple foreground models. Furthermore, the distance between the real object and the virtual object may be, for example, the distance between the centers of gravity of both objects or the shortest distance between the surfaces of both objects. Furthermore, the entire texture data used to render the selected foreground model does not necessarily need to be used. That is, texture data of a portion of the foreground object (e.g., texture data corresponding to a portion of the foreground object within a predetermined range centered on the point of the foreground object closest to the virtual object) may be acquired.

[0035] In S502, the CG image processing unit 204 converts the texture data of the foreground model acquired in S502 from RGB space to HSI space. In S503, the CG image processing unit 204 converts the color information of the CG image generated by the CG image generation unit 203 from RGB space to HSI space. In S504, the CG image processing unit 204 processes the saturation and lightness of the CG image based on the HSI space model of the texture data of the foreground model. For example, the CG image processing unit 204 calculates the average values ​​of the saturation and lightness of the texture data of the foreground model. Then, the CG image processing unit 204 changes the saturation and lightness of the CG image so that the average values ​​of the saturation and lightness of the CG image match or approach the average values ​​calculated for the texture data of the foreground model, while leaving the hue of the CG image unchanged. In S505, the CG image processing unit 204 inversely converts the CG image from HSI space to RGB space. The above is an example of CG image processing according to the first embodiment.

[0036] As described above, in the first embodiment, the saturation and brightness of a CG image are modified using texture data of a foreground model generated based on multiple viewpoint images. As a result, when a virtual viewpoint image and a CG image are superimposed, the colors of the real object in the virtual viewpoint image and the CG image become closer, making it possible to generate a more natural composite image.

[0037] In the first embodiment, an example in which saturation and brightness are processed using texture data of a foreground model has been described. However, this is not limiting. For example, texture data of a background model may be used. Furthermore, if the background model is large, such as a stadium or a field, only a portion of the background model may be used. For example, texture data of a portion of the background model's texture data close to the virtual object may be extracted, and the extracted texture data may be used. Furthermore, an image captured by a physical camera selected based on the position and surface orientation of the virtual object may be selected from multiple viewpoint images, and the saturation and brightness of the selected captured image may be used. The surface orientation of the virtual object may be determined based on, for example, the surface normal or vertex normal of a polygon. In this case, the entire color information of the selected captured image may be used, or a portion of the selected captured image may be extracted and the color information of the extracted portion may be used. Examples of the image of the portion extracted from the captured image include an image of a predetermined range near the virtual object, an image of a region of a real object close to the virtual object, and the like.

[0038] Furthermore, when processing a CG image, the CG image may be changed gradually to prevent abrupt changes. Instead of processing the CG image for each frame, a predetermined unit time may be set and the CG image may be processed for each unit time. Furthermore, even in this case, the CG image may be changed gradually. Processing of a CG image based on a multi-viewpoint image may not be performed continuously, and the user may switch processing on and off via the information processing device 103. Furthermore, processing of a CG image may not only be based on a multi-viewpoint image, but may also be performed using computer graphics by setting a light source or material. Although the present embodiment describes the CG image generation unit 203 and the CG image processing unit 204 separately, this is not limiting. For example, the CG image processing unit 204 may combine some or all of the functions of the CG image generation unit 203. In this case, for example, the CG image processing unit 204 may generate a CG image while processing an image of a virtual object based on a multi-viewpoint image (or texture data obtained from the multi-viewpoint image).

[0039] (Second embodiment) In the first embodiment, a configuration was described in which the saturation and brightness of a CG image superimposed on a virtual viewpoint image are adjusted based on the texture of a foreground model to match the colors of the real object in the virtual viewpoint image and the virtual object in the CG image, thereby reducing the sense of incongruity in the composite image. In the second embodiment, a configuration is described in which the sense of incongruity caused by differences in resolution between the real object and the virtual object is reduced. For example, an image of a real signboard captured close to a physical camera or with a telephoto lens has a high resolution, while an image of a real signboard captured far from a physical camera or with a wide-angle lens has a low resolution. While the resolution of a real signboard image varies depending on the shooting conditions, the resolution of a CG image of a virtual signboard, which is a virtual object, is always constant. Therefore, the resolution of the image of a virtual signboard placed in a virtual space differs from that of the surrounding real signboards, resulting in uneven resolution between the signs, potentially resulting in an unnatural composite image that feels unnatural.

[0040] In the second embodiment, the resolution of the texture data of a virtual model is processed according to the resolution of the texture data (i.e., color information) of a background model that is close to the virtual object. Note that explanations of parts common to the first embodiment, such as the hardware configuration and functional configuration of the image processing device 102, will be omitted or simplified, and the following explanation will focus on the difference, namely, the processing control of CG images.

[0041] Fig. 6 is a diagram illustrating a synthesis process of a virtual viewpoint image and a CG image according to the second embodiment. Fig. 7 is a flowchart showing a CG image processing process according to the second embodiment. In the CG image processing process according to the second embodiment, the resolution of the CG image is changed according to the resolution of the texture data for rendering the surface of the background model.

[0042] In FIG. 6 , virtual viewpoint image 601 is an image obtained by capturing an image of a real signboard 602 that actually exists with a virtual camera, and CG image 603 is an image obtained by capturing an image of a virtual signboard 604 with a virtual camera. The resolution of the texture data of the virtual signboard 604 shown in CG image 603 is higher than the resolution of the texture data of the real signboard 602 in virtual viewpoint image 601. If the resolution of the texture data of the virtual signboard 604 is excessively high compared to that of the real signboard 602, and CG image 603 is superimposed on virtual viewpoint image 601, an unnatural composite image may result. Therefore, the CG image processing unit 204 of the second embodiment processes CG image 603 so that the resolution of the texture data of the virtual signboard 604 approaches the resolution of the texture data of the real signboard 602 in virtual viewpoint image 601. By superimposing the processed CG image 605 on virtual viewpoint image 601, the composite image generation unit 205 generates composite image 606 in which the unnaturalness caused by the difference in resolution is reduced or eliminated.

[0043] Fig. 7 is a flowchart showing in detail the process (CG image processing process) of S405 in Fig. 4. When the texture data of the background model is input from the virtual viewpoint image generation unit 202 and the texture data of the virtual object is input from the CG image generation unit 203, the CG image processing unit 204 starts the processing process shown in Fig. 7.

[0044] In S701, the CG image processing unit 204 acquires texture data to be applied to a background model close to the virtual object. The texture data to be applied to the background model close to the virtual object is, for example, texture data obtained by extracting a portion close to the virtual object (a range of a predetermined size centered on the position closest to the virtual object) from the texture data of the background object. Of course, this is not limited to this, and for example, the entire texture data of the background model corresponding to the real object closest to the virtual object may be selected. In S702, the CG image processing unit 204 acquires texture data to be applied to the virtual object.

[0045] In S703, the CG image processing unit 204 compares the resolution of the texture data of the background model and the virtual object. In this embodiment, the resolution of the texture data of the background model is obtained, for example, from the number of pixels in the area to which the texture data is applied in the captured image that is the source of the texture data, and the actual size of that area in three-dimensional space. Also, in this embodiment, the resolution of the texture data of the virtual object is obtained from the size of the area to which the texture data is applied in the virtual object, and the number of pixels in the CG image corresponding to that area. In S704, the CG image processing unit 204 branches the process based on the result of the resolution comparison in S703. If the difference in resolution between the texture data of the background model and the virtual object is equal to or less than a predetermined value (NO in S704), this process ends. In this case, the CG image generated in S404 is used as is in generating the composite image in S406.

[0046] If the difference in resolution between the texture data of the background model and the virtual object is greater than the predetermined value (YES in S704), the process proceeds from S704 to S705. In S705, the CG image processing unit 204 changes the resolution of the texture data of the virtual object in the CG image based on the resolution of the texture data of the background model. For example, suppose the virtual object and background model are the same size and the resolution of the texture data of the virtual object is twice the resolution of the texture data of the background model. In this case, the CG image processing unit 204 processes the CG image generated in S404 to reduce the size of the texture data of the virtual object by 50%. In this way, the resolution of the texture data of the images of the real signboard and the virtual signboard used in the composite image can be made to match, or the resolution of the texture data can be made closer.

[0047] As described above, according to the second embodiment, the CG image is processed by changing the resolution of the virtual object according to the resolution of the texture data of the background model, and a more natural composite image is generated by superimposing the processed CG image and the virtual viewpoint image. Note that the texture data used to process the CG image is not limited to the background model, and texture data of a foreground model existing near the virtual object may be used as in the first embodiment. Also, in the second embodiment, the CG image is processed based on the texture data of a signboard (real signboard) that actually exists near the virtual signboard, but this is not limited to this. For example, when replacing a real signboard with a virtual signboard, the texture data of the virtual signboard may be changed based on the texture data of the real signboard that is the target of replacement.

[0048] (Third embodiment) In the first and second embodiments, texture data was selected based on a real object selected based on its positional relationship with the virtual object, and the CG image was processed based on the selected texture data. In the third embodiment, the imaging system 101 identifies the type of real object (player, ball, sign, etc.) present in the space being imaged, and the CG image is processed based on the texture data of the real object selected based on the identified type and the type of virtual object. Note that explanations of parts common to the first embodiment, such as the hardware configuration and functional configuration of the image processing device 102, will be omitted or simplified, and the following explanation will focus on the difference, namely, the processing of the CG image.

[0049] Fig. 8 is a diagram illustrating the synthesis process of a virtual viewpoint image and a CG image according to the third embodiment. Fig. 9 is a flowchart showing the processing of a CG image according to the third embodiment. As described above, in the third embodiment, the type of object is identified, and the CG image is processed according to the texture data of a predetermined object from among the objects whose type has been identified.

[0050] In FIG. 8, a virtual viewpoint image 801 is an image of a signboard 802 that actually exists and an object other than a signboard 803, photographed by a virtual camera. A CG image 804 is an image of a virtual signboard 805 photographed by the virtual camera. The texture data of the virtual signboard 805 has a higher resolution than the texture data of the signboard 802 that actually exists. In the second embodiment, a method of processing the resolution of a CG image using texture data of a foreground or background model that is similar to the virtual object was described. In contrast, the CG image processing unit 204 of the third embodiment identifies the type of object in the foreground or background model, selects an object of a predetermined type, and processes the CG image using the texture data of the selected object.

[0051] In the example of FIG. 8, the type of virtual object is a signboard, and a signboard 802, which is a real object of the same type, is selected, and the texture data of a virtual signboard 805 is changed based on the texture data of the signboard. The composite image generation unit 205 generates a composite image 807 by superimposing the processed CG image 806 on the virtual viewpoint image 801. For example, the color or resolution of the virtual object is changed based on the color or resolution of the texture data of the real object. The color change is as described in the first embodiment, and the resolution change is as described in the second embodiment. As a result, when a signboard that actually exists and a virtual signboard are captured by the virtual camera, the color and resolution of the real signboard and the virtual signboard are similar, making it possible to generate a composite image that does not look unnatural.

[0052] The processing shown in the flowchart of Fig. 9 is processing executed by the CG image processing unit 204 of the third embodiment. The CG image processing unit 204 starts the processing of Fig. 9 when the positions and texture data of the foreground model or background model are input from the virtual viewpoint image generation unit 202, and the texture data of the virtual object is input from the CG image generation unit 203. Fig. 9 is a flowchart illustrating in detail the processing of S405 in Fig. 4 (processing of processing a CG image).

[0053] In S901, the CG image processing unit 204 identifies the type of real object in the foreground or background model and selects a real object of the same type as the virtual object. For example, an object detection algorithm may be used to identify the type of object. For a background model, information indicating the type may be embedded in mesh data, for example. In S902, the CG image processing unit 204 processes the CG image of the virtual object based on the texture data of the real object selected in S901. The processing of the CG image is as described in the first and second embodiments.

[0054] As described above, according to the third embodiment, a CG image is processed based on texture data for a predetermined type of object from among real objects whose types have been identified. According to the third embodiment, for example, when generating a composite image in which one of a plurality of real signs is replaced with a virtual sign, the CG image can be processed more reliably using the texture data of the sign object. In this way, since the CG image is processed using the texture data of a real object of the same type as the virtual object, a more natural composite image can be obtained. It goes without saying that the first to third embodiments can be used in combination as appropriate.

[0055] (Fourth embodiment) In the first to third embodiments, a process for processing a CG image based on texture data of an object generated from multiple viewpoint images when superimposing the CG image on a virtual viewpoint image, which is a two-dimensional image, has been described. In the fourth embodiment, an augmented reality image is used as the two-dimensional image. That is, in the fourth embodiment, a process for processing a CG image based on an image captured by a camera in augmented reality (AR) will be described. Recently, various services using augmented reality technology have been provided. Using AR technology, a CG image can be superimposed on an image (real image) captured of real space. For example, a virtual advertisement (CG image) can be displayed on the real image. However, when a real image and a CG image are superimposed and displayed, an unnatural effect may occur in which the CG image appears to float in the real image. This is similar to the unnatural effect described in the first to third embodiments. In augmented reality, it is also necessary to match the color and resolution of an object (real object) existing in real space that is captured in the real image with a virtual object that is captured in the CG image, depending on the lighting conditions of the shooting space, the shooting conditions of the camera, etc. Therefore, in the fourth embodiment, as an example of matching the colors of a real object and a virtual object, a method of processing the saturation and / or brightness of a CG image based on a real image captured by a camera will be described. Note that parts common to the first to third embodiments will be omitted or simplified in the description.

[0056] <Image processing system hardware configuration> FIG. 10 is a diagram showing an example of the overall configuration of an image processing system according to this embodiment. The image processing system 1000 acquires an image captured by a camera and determines color information of a CG image according to color information of a real object captured in the captured image. The image processing system 1000 then generates and outputs a composite image based on the determined color information of the CG image. The image processing system 1000 includes a camera 1010, an image processing device 1020, and a display device 1030. An example of the hardware configuration of the image processing device 1020 of the fourth embodiment is as described in the first embodiment (FIG. 1(b)).

[0057] The camera 1010 captures an image of real space. The captured image, pre-calculated internal parameters of the camera 1010, and the like are transmitted to the image processing device 1020. The internal parameters of the camera are internal parameters specific to the camera, such as focal length, image center, and lens distortion parameters. The number of cameras is not limited to one, and multiple cameras may be used. When multiple cameras are used, multiple captured images and internal parameters are transmitted to the image processing device 1020. In this embodiment, a process of processing a CG image based on an image captured by one camera is described, but the present invention is not limited to this, and one CG image may be processed based on images captured by multiple cameras.

[0058] The image processing device 1020 estimates three-dimensional information of the real space and the position and orientation of the camera based on the captured image input from the camera 1010 and the internal parameters of the camera. The estimation of the three-dimensional information of the real space and the position and orientation of the camera is performed using a technique such as Visual SLAM. This technique makes it possible to calculate three-dimensional information about the surroundings of the camera and parameters indicating the position and orientation of the camera (external parameters of the camera). For example, with Visual SLAM, the three-dimensional information of the real space is output as a set of three-dimensional coordinates of a large number of feature points by recognizing feature points of surrounding objects from the captured image of the camera. In other words, three-dimensional information (three-dimensional coordinates) of a group of real objects in the real space can be acquired. Note that the estimation of the three-dimensional information of the real space is not limited to this method, and sensors such as laser sensors typified by LiDAR may also be used or combined, and the image processing device 1020 may include various sensors (depth sensors and acceleration sensors). The three-dimensional information of the real objects may include not only three-dimensional coordinates but also information such as the type of real object (signboard, floor, wall, etc.) corresponding to the three-dimensional coordinates, using object recognition technology. Estimation of the camera's external parameters may be geometrically calculated, for example, by detecting markers with unique identification information placed in real space, and is not limited to the above method. The camera's external parameters are expressed, for example, by a rotation matrix and a position vector. The camera's internal parameters and external parameters are collectively referred to as camera parameters. In this embodiment, it is assumed that the position and orientation of the real camera (real camera) match the position and orientation of the virtual camera, and the camera parameters of the virtual camera are matched to the camera parameters of the real camera. However, this is not limiting, and the camera parameters of the real camera and the virtual camera may be different. It is also assumed that the coordinate systems of the real space where the three-dimensional information is estimated and the virtual space where the virtual advertisement is placed match.

[0059] The image processing device 1020 captures a virtual space in which virtual objects such as virtual advertisements are placed, based on the acquired camera parameters, and generates a CG image. 3D data of the virtual objects is stored in advance in the ROM 113 (FIG. 1(b)), and reads and uses the data. The 3D data of the virtual objects includes placement information (posture information and position information), and the virtual objects are placed in the virtual space based on the placement information. Note that the placement position of the virtual objects is not limited to this, and the image processing device 1020 may change the position and posture of the virtual objects, or place them in any position and posture in the virtual space. The image processing device 1020 then processes the CG image based on the captured image and the three-dimensional information of the generated real objects. More specifically, the image processing device 1020 processes the CG image based on color information of real objects close to the virtual object so that the colors of the real objects and the virtual objects match. First, the image processing device 1020 selects a real object (or a point) that is close to the virtual object using the three-dimensional information of the real objects, and projects the real object onto the captured image based on the camera parameters. The image processing device 1020 then acquires color information of the projection destination and changes and processes the color information of the virtual object based on that color information. Finally, the image processing device 1020 generates a composite image by superimposing the processed CG image on the captured image. By processing the CG image in this way, a CG image is generated whose color changes depending on the shooting conditions of the camera 1010 and the lighting conditions of the shooting space, allowing for a more natural composite image to be obtained.

[0060] The image processing device 1020 outputs the generated composite image to the display device 1030.

[0061] The display device 1030 displays the composite image generated and output by the image processing device 1020. The display device 1030 realizes augmented reality by successively updating and continuously displaying the composite image. The user views the composite image displayed on the display device 1030 and operates the camera 1010 to specify the zoom and position and orientation of the camera. The display device 1030 is, for example, a monitor, a tablet, or a video see-through head-mounted display (HMD).

[0062] In this embodiment, the camera 1010, the image processing device 1020, and the display device 1030 are assumed to be separate devices, but this is not limiting, and the image processing device 1020 may also function as both the camera 1010 and the display device 1030. For example, when a composite image is displayed on a tablet employing a touch panel display built into the camera, the tablet can also function as the camera 1010, the image processing device 1020, and the display device 1030. The composite image is displayed on the touch panel display of the tablet, and the user operates the tablet by touching the touch panel display, and the composite image is generated by cooperation between a CPU, RAM, and ROM, which will be described later.

[0063] <Functional configuration of image processing device> FIG. 11 is a diagram showing an example of the functional configuration of the image processing device 1020 for processing a CG image based on an image captured by the camera 1010.

[0064] The image processing device 1020 includes a communication control unit 1101 , a camera information estimation unit 1102 , a CG image generation unit 1103 , a CG image processing unit 1104 , and a composite image generation unit 1105 .

[0065] The communication control unit 1101 receives the captured image and internal parameters of the camera 1010 from the camera 1010 using the communication unit 114. The communication control unit 1101 outputs the received captured image and internal parameters to the camera information estimation unit 1102. The communication control unit 1101 also outputs the captured image to the CG image processing unit 1104. Furthermore, the communication control unit 1101 transmits the composite image received from the composite image generation unit 1105 to the display device 1030 using the communication unit 114.

[0066] Based on the captured image of camera 1010 and the internal parameters acquired from the communication control unit 1101, the camera information estimation unit 1102 estimates the three-dimensional information of the real space and the external parameters of camera 1010. The camera information estimation unit 1102 outputs the internal parameters input from the communication control unit 1101, the generated three-dimensional information of the real space, and the external parameters of camera 1010 to the CG image generation unit 1103.

[0067] Based on the camera parameters input from the camera information estimation unit 1102, the CG image generation unit 1103 captures a virtual object and generates a CG image. Since the camera parameters are those of camera 1010, the viewing angle, position, and orientation of the generated CG image match those of the captured image of camera 1010. The CG image generation unit 1103 outputs the generated CG image, the arrangement information of the virtual object, the three-dimensional information of the real space input from the camera information estimation unit 1102, and the camera parameters to the CG image processing unit 1104.

[0068] Using the captured image of camera 1010 sent from the communication control unit 1101, the arrangement information of the virtual object, the three-dimensional information of the real space, and the camera parameters sent from the CG image generation unit 1103, the CG image processing unit 1104 processes the CG image received from the CG image generation unit 1103. The CG image processing unit 1104 sends the processed CG image and the captured image of camera 1010 to the composite image generation unit 1105. The method for processing the CG image will be described later using FIG. 12.

[0069] By synthesizing the captured image of camera 1010 and the CG image input from the CG image processing unit 1104, the composite image generation unit 1105 generates a composite image with CG drawn on the captured image of camera 1010. The composite image generation unit 1105 outputs the generated composite image to the communication control unit 1101.

[0070] <Explanation of the method for processing the CG image and generating the composite image> The procedure for processing a CG image and generating a composite image according to the fourth embodiment will be described with reference to FIG. 12. Here, as an example of processing a CG image by the CG image processing unit 1104, a mode of changing the color of a CG image based on a captured image of a real object will be described. FIG. 12(a) is a schematic diagram showing, from a bird's-eye view, a state in which a real object 1201 in the capture space of the camera 1010 and a virtual object 1202 that does not actually exist in the capture space are captured by a real camera (virtual camera) 1203. Note that the positions and orientations of the real camera and the virtual camera are the same. FIG. 12(b) is a schematic diagram illustrating a method for combining the captured image 1204 of FIG. 12(a) with a CG image 1206.

[0071] A real camera 1203 captures an image of a real object 1201 and generates a captured image 1204. The real camera 1203 corresponds to the camera 1010. The real object 1201 is illuminated by a light source 1205 such as the sun or a light, and is therefore backlit, resulting in a darker color than in front-lit conditions. The image processing device 1020 estimates three-dimensional information including the real object 1201 and external parameters of the camera based on the captured image 1204 captured by the real camera 1203. The image processing device 1020 then generates a CG image 1206 based on the camera parameters of the camera 1010. Furthermore, the image processing device 1020 processes the color saturation and brightness of the CG image 1206 based on the estimated three-dimensional information and the captured image 1204 so that the color saturation and brightness are closer to those of the real object 1201, thereby generating a CG image 1207. Here, the real object used to process the color saturation and brightness of the CG image 1206 is selected to be, for example, one that is close to the virtual object. Finally, the image processing device 1020 superimposes the processed CG image 1207 and the captured image 1204 to generate a composite image 1208.

[0072] As described above, by processing the color of the virtual object 1202 appearing in the CG image 1206 in accordance with the color information of the real object 1201 appearing in the photographed image 1204, it is possible to generate a natural composite image 1208 without imitating the light source 1205 in the virtual space.

[0073] <Control of CG Image Processing and Generation of Composite Image> FIG. 13 is a flowchart showing the composite processing of a photographed image and a CG image according to the present embodiment. The processing shown in FIG. 13 is realized by, for example, a control program stored in the ROM 113 being read into the RAM 112 and the CPU 111 executing them. The communication control unit 1101 receives a photographed image and internal parameters from the camera 1010. When the camera information estimation unit 1102 receives these received data from the communication control unit 1101, the processing shown in FIG. 13 is started.

[0074] In S1301, the camera information estimation unit 1102 acquires a photographed image and internal parameters of the camera 1010. Then, based on the acquired data, it estimates the three-dimensional information of the real space and the external parameters of the camera 1010. The camera information estimation unit 1102 outputs the camera parameters and the three-dimensional information of the real space to the CG image generation unit 1103.

[0075] [[ID=⑨]] In S1302, the CG image generation unit 1103 acquires the camera parameters of the camera 1010 and the 3D data of the virtual object, and generates a CG image showing the virtual object. The CG image generation unit 1103 outputs the generated CG image, the arrangement information of the virtual object, the three-dimensional information of the real space, and the camera parameters to the CG image processing unit 1104.

[0076] In S1303, the CG image processing unit 1104 processes the CG image input from the CG image generation unit 1103 based on the arrangement information of the virtual object, the three-dimensional information of the real space, the camera parameters input from the CG image generation unit 1103, and the photographed image input from the communication control unit 11O1. Details of the CG image processing in S13O3 will be described later with reference to FIG. 14. The CG image processing unit 1104 outputs the processed CG image and the photographed image to the composite image generation unit 1105.

[0077] In S1304, the composite image generation unit 1105 synthesizes the captured image input from the CG image processing unit 1104 and the CG image to generate a composite image. Then, in S1305, the composite image generation unit 1105 outputs the generated composite image. In the present embodiment, the composite image generation unit 1105 transmits the composite image to the display device 1030 via the communication control unit 1101. The above is the composite process of the captured image of the camera 1010 and the CG image according to the fourth embodiment.

[0078] <Explanation of the processing of the CG image> FIG. 14 is a flowchart showing an example of the processing of the CG image according to the present embodiment, and shows the details of the processing of the CG image in S1303 of FIG. 13. In the processing of the CG image in the fourth embodiment, the color information of the CG image is determined and processed based on the saturation and brightness of the captured image of the real object close to the virtual object in the three-dimensional space. More specifically, the CG image processing unit 1104 acquires the saturation and brightness of the captured image of the real object whose arrangement position is close to the virtual object in the real (virtual) space, and changes the saturation and brightness of the virtual object so as to approach them. The saturation and brightness of the CG image before processing are calculated by the CG image generation unit 1103 based on the texture data and material data of the virtual object. The process shown in FIG. 14 is executed by the CG image processing unit 104. Further, the process shown in FIG. 14 is started triggered by the acquisition of the captured image from the communication control unit 1101 and the acquisition of the three-dimensional information of the real space, the CG image, the arrangement information of the virtual object, and the camera parameters from the CG image generation unit. However, the three-dimensional information of the real space assumes a set (point group) of the three-dimensional coordinates of the real object group.

[0079] In S1401, the CG image processing unit 1104 acquires a point cloud of a real object that is close to the placement position of the virtual object from a set of three-dimensional coordinates (point cloud) of a group of real objects input from the CG image generation unit. For example, the CG image processing unit 1104 acquires a point cloud of a real object that exists within a predetermined distance from the placement position of the virtual object. Note that the point cloud may include coordinates of multiple points, or may be the coordinate of a single point. If a point cloud and type information corresponding to that point cloud exist, the selection of the point cloud of the real object may be determined based on the type information. For example, when processing a CG image of a virtual billboard advertisement, the point cloud of a signboard that exists in real space may be selected preferentially.

[0080] In S1402, the CG image processing unit 1104 projects the point cloud of the selected real object onto the captured image input from the communication control unit 1101 based on the camera parameters input from the CG image generation unit.

[0081] In S1403, the CG image processing unit 1104 acquires color information of the captured image onto which the point cloud of the real object is projected.

[0082] In S1404, the CG image processing unit 1104 converts the acquired color information from the RGB space to the HSI space.

[0083] In S1405, the CG image processing unit 1104 converts the color information of the CG image generated by the CG image generation unit 1103 from the RGB space to the HSI space.

[0084] In S1406, the CG image processing unit 1104 processes the saturation and lightness of the CG image based on the HSI space model of the color information of the captured image. For example, the CG image processing unit 1104 calculates the average values ​​of the saturation and lightness of the color information of the captured image acquired in S1404. Then, the CG image processing unit 1104 changes the saturation and lightness of the CG image so that the average values ​​of the saturation and lightness of the CG image match or approach the average values ​​calculated for the color information of the captured image, while leaving the hue of the CG image unchanged.

[0085] In S1407, the CG image processing unit 1104 inversely converts the CG image from the HSI space to the RGB space. The above is an example of the CG image processing according to the fourth embodiment.

[0086] As described above, in the fourth embodiment, a process of processing the saturation and brightness of a CG image is performed using an image captured by the camera 1010. As a result, when a captured image and a CG image are superimposed in augmented reality, the colors of the real object in the captured image and the virtual object in the CG image become closer, making it possible to generate a more natural composite image. Note that the various methods described in the first to third embodiments, such as the color processing method of the first embodiment, can also be used in appropriate combination in the fourth embodiment.

[0087] (Fifth embodiment) In the fourth embodiment, a process for displaying a composite image in which a CG image is superimposed on a real image on the display device 1030 is described. In the fifth embodiment, an aspect is described in which a CG image processed based on a captured real image is displayed on an optical see-through display device, for example, an optical see-through HMD for augmented / mixed reality. Note that parts common to the fourth embodiment will be omitted or simplified in the description.

[0088] FIG. 15 is a diagram showing an example of the overall configuration of an image processing system according to this embodiment. The image processing system 1500 acquires an image captured by a camera and determines color information of a CG image according to color information of a real object captured in the captured image. The image processing system 1500 then displays the CG image for which color information has been determined. The image processing system 1500 includes a camera 1510, an image processing device 1520, and a display device 1530. However, the camera 1510 is similar to the camera 1010. An example of the hardware configuration of the image processing device 1520 of the fifth embodiment is as shown in the first embodiment (FIG. 1(b)).

[0089] Unlike the image processing device 1020, the image processing device 1520 outputs a CG image processed based on an image captured by the camera 1510 directly to the display device 1530 without superimposing the image on the captured image.

[0090] The display device 1530 displays the CG image generated and output by the image processing device 1520. The display device 1530 is, for example, an optical see-through HMD, which displays the CG image on a transparent screen and superimposes it on the real scenery to realize augmented / mixed reality.

[0091] In this embodiment, the camera 1510, the image processing device 1520, and the display device 1530 are assumed to be separate devices, but as in the fourth embodiment, the image processing device 1520 may also have the functions of the camera 1510 and the display device 1530.

[0092] FIG. 16 is a diagram showing an example of the functional configuration of the image processing device 1520 for processing a CG image based on an image captured by the camera 1510.

[0093] The image processing device 1520 includes a communication control unit 1601, a camera information estimation unit 1602, a CG image generation unit 1603, and a CG image processing unit 1604. However, the camera information estimation unit 1602 and the CG image generation unit 1603 are similar to the camera information estimation unit 1102 and the CG image generation unit 1103, respectively.

[0094] The communication control unit 1601 receives the captured image and internal parameters of the camera 1510 from the camera 1510 using the communication unit 114. The communication control unit 1601 outputs the received captured image and internal parameters to the camera information estimation unit 1602, and outputs the captured image to the CG image processing unit 1604. In addition, the communication control unit 1601 transmits the CG image received from the CG image processing unit 1604 to the display device 1530 using the communication unit 114.

[0095] The CG image processing unit 1604 processes the CG image received from the CG image generation unit 1603, using the image captured by the camera 1510 sent from the communication control unit 1601, and the virtual object placement information, three-dimensional information of the real space, and camera parameters sent from the CG image generation unit 1103. The CG image processing unit 1604 sends the processed CG image to the communication control unit 1601. The method of processing the CG image is the same as the method of processing the CG image in the fourth embodiment.

[0096] Fig. 17 is a flowchart showing a CG image generation process according to the fifth embodiment. As described above, in the fifth embodiment, a CG image generated based on an image captured by the camera 1510 is output to the display device 1530 as is without being superimposed on the captured image. The process shown in Fig. 17 is realized, for example, by reading a control program stored in the ROM 113 into the RAM 112 and having the CPU 111 execute the program. The communication control unit 1601 receives the captured image and internal parameters from the camera 1510. When the camera information estimation unit 1602 receives the received data from the communication control unit 1601, the process shown in Fig. 17 starts. However, S1701 and S1702 are the same as S1301 and S1302, respectively.

[0097] In S1703, the CG image processing unit 1604 processes the CG image input from the CG image generation unit 1603 based on the virtual object placement information input from the CG image generation unit 1603, the three-dimensional information of the real space, the camera parameters, and the captured image input from the communication control unit 1601. The processing of the CG image is as described in the fourth embodiment (FIG. 14). In S1704, the CG image processing unit 1604 outputs the processed CG image to the communication control unit 1601. Then, this processing ends. The processed CG image is transmitted from the communication control unit 1601 to the display device 1530. This completes the CG image generation processing according to the fifth embodiment.

[0098] As described above, in the fifth embodiment, a CG image is processed using an image captured by the camera 1510, and the CG image is displayed on a transparent screen such as an optical see-through HMD. This makes the colors of the real scenery and the CG image closer to each other, making it possible to achieve a more natural augmented / mixed reality. Note that the various methods described in the first to third embodiments, such as the CG image processing method of the fourth embodiment, can also be used in combination as appropriate in the fifth embodiment.

[0099] (Other embodiments) The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0100] The present disclosure is not limited to the above-described embodiments, and various modifications and variations can be made without departing from the spirit and scope of the present invention. Therefore, to apprise the public of the scope of the present invention, the following claims are appended. [Explanation of symbols]

[0101] 101: imaging system, 102: image processing device, 103: information processing device, 201: communication control unit, 202: virtual viewpoint image generation unit, 203: CG image generation unit, 204: CG image processing unit, 205: composite image generation unit

Claims

1. an acquisition means for acquiring a two-dimensional image including a virtual object; a processing means for processing the saturation and / or brightness of the virtual object in the two-dimensional image so that it matches or approaches the saturation and / or brightness included in color information corresponding to the three-dimensional shape of a real object generated based on a plurality of images; and The image processing device is characterized in that the processing means processes the resolution of the virtual object in the two-dimensional image when a difference between the resolution of the virtual object in the two-dimensional image and the resolution of the real object is larger than a predetermined value.

2. An acquisition means for acquiring a two-dimensional image including a virtual object; a processing means for processing the saturation and / or brightness of the virtual object in the two-dimensional image so that the saturation and / or brightness coincides with or approaches the saturation and / or brightness included in color information corresponding to the three-dimensional shape of a real object generated based on a plurality of images; an identification means for identifying the type of the real object; and The image processing device is characterized in that the processing means processes the saturation and / or brightness of the virtual object based on the saturation and / or brightness contained in color information corresponding to the three-dimensional shape of the identified type of real object.

3. An acquisition means for acquiring a two-dimensional image including a virtual object; a processing means for processing the saturation and / or brightness of the virtual object in the two-dimensional image so that it matches or approaches the saturation and / or brightness included in color information corresponding to the three-dimensional shape of a real object generated based on a plurality of images; and the processing means determines the saturation and / or brightness of the virtual object based on the saturation and / or brightness of an image selected from the plurality of images based on the position of the virtual object and the direction of a surface of the virtual object.

4. moreover, a means for generating a virtual viewpoint image based on the plurality of images and viewpoint information indicating a position of a virtual viewpoint and a line of sight direction from the virtual viewpoint; a means for generating a composite image based on the virtual viewpoint image and the two-dimensional image processed by the processing means; 4. The image processing device according to claim 1, further comprising:

5. The image processing device according to claim 4 , wherein the composite image is generated by superimposing the virtual viewpoint image on the two-dimensional image.

6. the acquisition means acquires a position of the virtual object in a virtual space; 6. The image processing device according to claim 1, wherein the processing means processes the saturation and / or brightness of the virtual object in the two-dimensional image based on the saturation and / or brightness included in color information corresponding to the three-dimensional shape of a real object selected based on its positional relationship with the virtual object.

7. 7. The image processing apparatus according to claim 6, wherein the processing means selects a real object that is closest to the virtual object or that is within a predetermined distance from the virtual object.

8. 8. The image processing apparatus according to claim 1, wherein the real object is a person.

9. 8. The image processing apparatus according to claim 1, wherein the real object is a structure.

10. 10. The image processing device according to claim 1, wherein the two-dimensional image is an image generated based on the plurality of images.

11. 11. The image processing device according to claim 1, wherein the processing means processes the saturation and brightness of the virtual object in the two-dimensional image so that average values ​​of the saturation and brightness included in color information corresponding to the three-dimensional shape of the real object and average values ​​of the saturation and brightness of the virtual object in the two-dimensional image match or approach each other.

12. 12. The image processing device according to claim 1, wherein the processing means processes the saturation and / or brightness of the virtual object while maintaining the hue of the virtual object.

13. an acquisition step of acquiring a two-dimensional image including the virtual object; a processing step of processing the saturation and / or brightness of the virtual object in the two-dimensional image so that it matches or approaches the saturation and / or brightness included in color information corresponding to the three-dimensional shape of a real object generated based on a plurality of images; and a control method for an image processing device, characterized in that, in the processing step, if a difference between the resolution of the virtual object in the two-dimensional image and the resolution of the real object is larger than a predetermined value, the resolution of the virtual object in the two-dimensional image is processed.

14. An acquisition step of acquiring a two-dimensional image including a virtual object; a processing step of processing the saturation and / or brightness of the virtual object in the two-dimensional image so that it matches or approaches the saturation and / or brightness included in color information corresponding to the three-dimensional shape of a real object generated based on a plurality of images; an identification step of identifying the type of the real object; and A control method for an image processing device, characterized in that in the processing step, the saturation and / or brightness of the virtual object is processed based on the saturation and / or brightness contained in color information corresponding to the three-dimensional shape of the identified type of real object.

15. An acquisition step of acquiring a two-dimensional image including a virtual object; a processing step of processing the saturation and / or brightness of the virtual object in the two-dimensional image so that it matches or approaches the saturation and / or brightness included in color information corresponding to the three-dimensional shape of a real object generated based on a plurality of images; and a control method for an image processing device, characterized in that, in the processing step, the saturation and / or brightness of the virtual object is determined based on the saturation and / or brightness of an image selected from the plurality of images based on the position of the virtual object and the direction of a surface of the virtual object.

16. A program for causing a computer to function as each of the means of the image processing device according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Image processing program, image processing device, image processing system, and image processing method

    JP2014203326A

  • Head mounted display device, method for controlling the same and computer program

    JP2016218547A

  • Imaging apparatus and control method of the same

    JP2017192106A

  • Information processing device, image generation method, and program

    JP2018180654A

  • Information processing device, information processing method, and program

    WO2019087513A1