Image processing system, image processing method, and program
The image processing system addresses the issue of unnatural blending by generating separate virtual viewpoint images for transparent and opaque subject portions, ensuring accurate color representation and natural appearance in virtual viewpoint images.
Patent Information
- Application Number
- JP2024088440
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-11
AI Technical Summary
Existing image processing systems generate unnatural virtual viewpoint images when highly transparent subjects are captured, as they blend with the background, and fail to accurately represent the colors of subjects behind the transparent objects.
An image processing system that acquires viewpoint information and generates separate virtual viewpoint images for transparent and opaque portions of a subject, combining them to create a natural-looking image by removing background color from the transparent areas.
The system effectively generates virtual viewpoint images with appropriate color representation for highly transparent objects, ensuring a natural appearance even when the virtual viewpoint differs from the imaging device.
Smart Images

Figure 2025180832000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing system, an image processing method, and a program. [Background technology]
[0002] A technology that captures images using multiple image capture devices positioned at different locations and generates a virtual viewpoint image that recreates the view from a virtual viewpoint using the captured images has attracted attention. This technology is expected to be applicable to a wide range of fields, such as sports broadcasts on television and film production, and is expected to capture images of subjects with various characteristics. For example, it is expected to capture images of performers wearing thin clothing or performers together with props and background sets made of glass, acrylic resin, etc. In other words, it is expected to capture images of subjects with different transmittances depending on the body part or multiple subjects with different transmittances.
[0003] Conventionally, when generating a virtual viewpoint image, the texture of the subject area in the virtual viewpoint image is determined using the texture of the subject area from multiple captured images. Therefore, when a highly transparent subject is captured, the captured image contains the background behind the subject, and therefore, the generated virtual viewpoint image also contains a texture that blends in with the background as seen from the imaging device. Here, if the virtual viewpoint for generating the virtual viewpoint image is set at a position different from the imaging device, the background seen from the imaging device through the highly transparent subject and the background seen from the virtual viewpoint through the highly transparent subject should be different. However, the generated virtual viewpoint image contains a texture that blends in with the background of the real space as seen from the imaging device, resulting in an unnatural-looking virtual viewpoint image.
[0004] Patent Document 1 describes a technique for generating a captured image in which the background is not reflected in the area of a highly transmittance subject by removing background color information from the area of the subject in a captured image of the highly transmittance subject. By applying this technique, a virtual viewpoint image can be generated using a captured image in which the background of real space is not reflected in the area of the highly transmittance subject, and a natural virtual viewpoint image can be generated. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 6-225329 Summary of the Invention [Problem to be solved by the invention]
[0006] However, if another subject appears behind a subject with high transmittance in the captured image, the color of the other subject cannot be removed from the subject's area, resulting in an unnatural-looking virtual viewpoint image being generated.
[0007] The present disclosure can generate a virtual viewpoint image including a highly transparent object that is expressed in appropriate colors. [Means for solving the problem]
[0008] The image processing system of the present disclosure has the following configuration: an acquisition means for acquiring viewpoint information indicating a position of a virtual viewpoint and a line of sight direction from the virtual viewpoint; a first generation means for generating a first virtual viewpoint image including a transparent or semi-transparent first portion of a subject, using the viewpoint information and a first captured image including a plurality of pixels corresponding to the first portion that do not correspond to an opaque second portion of the subject captured through the first portion; and for generating a second virtual viewpoint image including the second portion, using a second captured image including the second portion and the viewpoint information; a second generating means for generating a third virtual viewpoint image including the first portion and the second portion based on the first virtual viewpoint image and the second virtual viewpoint image; It has. [Effects of the Invention]
[0009] According to the present disclosure, it is possible to generate a virtual viewpoint image including a subject with high transmittance that is expressed in appropriate colors. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram showing the overall configuration of an image processing system according to a first embodiment. [Figure 2] FIG. 4 is a flowchart illustrating a process for generating a virtual viewpoint image according to the first embodiment. [Figure 3] FIG. 1 is a diagram showing an outline of a scene assumed in a first embodiment. [Figure 4] FIG. 10 is a flowchart showing a process of selecting a texture by a texture selection unit 105 according to the first embodiment. [Figure 5] 10 is an image diagram showing a determination result of the texture selection unit 105 in the first embodiment. FIG. [Figure 6] 10 is an image diagram showing a determination result of the texture selection unit 105 in the first embodiment. FIG. [Figure 7] FIG. 10 is a flowchart showing a process in which a virtual visual point image generating unit 106 according to the first embodiment generates a virtual visual point image. [Figure 8] 10 is an image diagram showing a rendering result of a virtual viewpoint image generating unit 106 according to the first embodiment. FIG. [Figure 9] FIG. 2 is a hardware configuration diagram of an image processing device 108 according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating the overall configuration of an image processing system according to a second embodiment. [Figure 11] FIG. 10 is a flowchart showing a process of correcting a captured image by a captured image correcting unit 109 in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] According to a preferred embodiment of the present invention, an image processing system includes an acquisition unit that acquires viewpoint information indicating the position of a virtual viewpoint and a line of sight direction from the virtual viewpoint. The image processing system acquires a first captured image that includes a transparent or semi-transparent first portion of a subject, and includes, among a plurality of pixels corresponding to the first portion, pixels that do not correspond to an opaque second portion of the subject captured through the first portion. The image processing system also includes a first generation unit that generates a first virtual viewpoint image that includes the first portion using the first captured image and the viewpoint information. The first generation unit generates a second virtual viewpoint image that includes the second portion using a second captured image that includes the second portion and the viewpoint information. The image processing system also includes a second generation unit that generates a third virtual viewpoint image that includes the first portion and the second portion based on the first virtual viewpoint image and the second virtual viewpoint image. Note that the first generation unit and the second generation unit may be the same generation unit. Here, the first virtual viewpoint image is a virtual viewpoint image in which only a transparent or semi-transparent first portion of the subject is colored, and the second virtual viewpoint image is a virtual viewpoint image in which only an opaque second portion of the subject is colored. Therefore, by using the first virtual viewpoint image and the second virtual viewpoint image, a third virtual viewpoint image including the first portion and the second portion can be generated. Specifically, the third virtual viewpoint image may be generated by combining the first virtual viewpoint image and the second virtual viewpoint image. Note that the first virtual viewpoint image may include not only the first portion but also both the first portion and the second portion. Note that the subject refers to an object captured by multiple imaging devices in real space, and multiple objects may be collectively referred to as the subject, or a single object may be referred to as the subject. In the above, multiple objects are collectively referred to as the subject.
[0012] According to this aspect, color information of a region of a subject with high transmittance in the virtual viewpoint image can be determined using color information of a region of the captured image where no other subjects are shown behind the subject with high transmittance. Therefore, even if the position of the virtual viewpoint is set at a position different from the position of the imaging device, a natural virtual viewpoint image can be generated in which appropriate color information is set for the region of the subject with high transmittance.
[0013] In the image processing system, pixels in the first captured image that do not correspond to the second portion captured through the first portion correspond to specific pixels of the first portion in the first virtual viewpoint image, and the first generation means determines color information of the specific pixels in the first virtual viewpoint image using color information of pixels that do not correspond to the second portion captured through the first portion.
[0014] In addition, pixels in the first captured image that do not correspond to the second portion imaged through the first portion and specific pixels in the first portion in the first virtual viewpoint image correspond to specific components of the three-dimensional shape of the subject. If the three-dimensional shape of the subject is composed of a point cloud and the components are points, a certain point in the point cloud corresponds to pixels in the first captured image that do not correspond to the second portion imaged through the first portion and specific pixels in the first portion in the first virtual viewpoint image.
[0015] According to this aspect, the first generation means can determine the color information of the specific pixel in the first virtual viewpoint image using color information of a pixel that does not correspond to the second portion imaged through the first portion.
[0016] The second captured image may be a captured image including the second portion captured through the first portion. In other words, the first captured image used to generate the first virtual viewpoint image and the second captured image used to generate the second virtual viewpoint image may be the same captured image.
[0017] The first generation means determines color information of the second portion in the second virtual viewpoint image using color information obtained by removing color information corresponding to the first portion from color information of pixels corresponding to the second portion captured through the first portion in the first captured image. For example, when the second portion is visible through a semi-transparent first portion in the first captured image, the pixels corresponding to the second portion in the first captured image contain color information of the first portion and color information of the second portion. Therefore, by removing the color information of the first portion, it is possible to determine color information of the second portion from the first captured image.
[0018] The acquisition means acquires a plurality of captured images including the subject. Furthermore, the acquisition means acquires a plurality of pieces of transmittance information corresponding to the plurality of captured images using a trained model that receives the plurality of captured images and outputs a plurality of pieces of transmittance information indicating the transmittance of regions corresponding to the subject in the plurality of captured images. It is not necessary to acquire the transmittance information using a trained model, and the transmittance information can be acquired using an existing method. For example, the material of the subject in the captured image may be identified using image recognition technology, and a transmittance may be set for each material. The image processing system also includes an identification means that identifies the first captured image and the second captured image from the plurality of captured images based on the plurality of transmittance information.
[0019] According to this aspect, the first portion and the second portion in the captured image can be identified, and therefore the first captured image and the second captured image can be identified.
[0020] Furthermore, the first generation means generates shape information indicating the three-dimensional shape of the subject based on the multiple captured images, position information of the multiple imaging devices that captured the multiple captured images, and the multiple transmittance information. Here, the shape information includes transparent or semi-transparent first components and opaque second components. Note that, for a transparent or semi-transparent first portion of the subject in real space, the first components are components that constitute a transparent or semi-transparent portion of the three-dimensional shape representing the subject in virtual space. Note that, for an opaque second portion of the subject in real space, the second components are components that constitute an opaque portion of the three-dimensional shape representing the subject in virtual space. Therefore, the multiple first components in virtual space correspond to the first portion in real space, and the multiple second components in virtual space correspond to the second portion in real space.
[0021] According to this aspect, a 3D model of a subject including a first component and a second component in a virtual space can be generated from a subject including a first portion and a second portion in a real space.
[0022] The identification means identifies, among the plurality of captured images, a captured image that includes the first portion in an imaging range of the imaging device and in which the second component is not present on a line passing through the first component from a position in virtual space corresponding to a position in real space of the imaging device as a first captured image. The identification means also identifies, among the plurality of captured images, a captured image that includes the second portion in an imaging range of the imaging device as a captured image including the second portion.
[0023] Furthermore, the identification means identifies, in each of the plurality of captured images, a region corresponding to a subject whose transmittance is equal to or greater than a threshold as a transparent or semi-transparent region, and identifies a region corresponding to a subject whose transmittance is less than the threshold as an opaque region. When the transmittance of each pixel of a captured image is calculated using an existing method, the captured image can be divided into transparent or semi-transparent regions and opaque regions. Note that how regions are identified relative to the threshold is not limited to the above. For example, a region whose transmittance is greater than the threshold may be identified as a transparent or semi-transparent region, and a region whose transmittance is equal to or less than the threshold may be identified as an opaque region.
[0024] The first generation means generates the first components using the transparent or semi-transparent regions in the plurality of captured images, and generates the second components using the opaque regions in the plurality of captured images, and generates shape information indicating the three-dimensional shape of the subject using these components.
[0025] The second generating means generates the third virtual viewpoint image by removing a background color from the area of the first virtual viewpoint image corresponding to the first portion and combining the first virtual viewpoint image with the second virtual viewpoint image. The background color here refers to color information of the background seen through a subject with high transmittance in the first captured image.
[0026] This allows the background color of the real space included in the first captured image to be removed from the first virtual viewpoint image. As a result, the third virtual viewpoint image does not include color information of the background of the real space, so that when a separate background model is set in the virtual space, a natural virtual viewpoint image is generated.
[0027] The means included in the image processing system may be controlled by one computer or by multiple computers, and the means included in the image processing system may be recorded in one computer program or in multiple computer programs.
[0028] <Example> Hereinafter, examples for carrying out the present disclosure will be described with reference to the drawings. Note that the following examples do not limit the present disclosure, and not all of the combinations of features described in the examples are necessarily essential to the solutions of the present disclosure. Note that the same components will be described with the same reference numerals, and duplicate descriptions will be omitted.
[0029] A virtual viewpoint image is an image generated by a user freely manipulating the position and orientation of a virtual camera, and is also called a free viewpoint image, an arbitrary viewpoint image, etc. Unless otherwise specified, the term "image" will be explained as including the concepts of both moving images and still images.
[0030] The viewpoint information used to generate a virtual viewpoint image is information indicating the position and orientation (line of sight direction) of the virtual viewpoint. Specifically, the viewpoint information is a parameter set including a parameter indicating the three-dimensional position of the virtual viewpoint and a parameter indicating the orientation of the virtual viewpoint in the pan, tilt, and roll directions. Note that the content of the viewpoint information is not limited to the above. For example, the parameter set serving as viewpoint information may include a parameter indicating the size of the field of view (angle of view) of the virtual viewpoint. Furthermore, the viewpoint information may have multiple parameter sets. For example, the viewpoint information may have multiple parameter sets corresponding to multiple frames constituting a moving image of the virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of multiple consecutive time points.
[0031] The image processing system described below has multiple imaging devices that capture images of an imaging area from multiple directions. The imaging area may be, for example, a stadium where sports such as soccer or karate are held, or a stage where a concert or a play is held. The multiple imaging devices are installed at different positions surrounding the imaging area and capture images synchronously. Note that the multiple imaging devices do not need to be installed around the entire periphery of the imaging area; depending on installation space restrictions, they may be installed only around a portion of the periphery of the imaging area. Furthermore, the number of imaging devices is not limited to the example shown in the figure. For example, if the imaging area is a soccer stadium, approximately 30 imaging devices may be installed around the stadium. Furthermore, imaging devices with different functions, such as telephoto cameras and wide-angle cameras, may be installed.
[0032] In this embodiment, the multiple image capturing devices are cameras each having an independent housing and capable of capturing images from a single viewpoint. However, this is not limiting, and two or more image capturing devices may be configured in the same housing. For example, a single camera equipped with multiple lens groups and multiple sensors and capable of capturing images from multiple viewpoints may be installed as the multiple image capturing devices.
[0033] A virtual viewpoint image is generated, for example, by the following method. First, multiple images (multiple captured images) are obtained by capturing images from different directions using multiple imaging devices. Next, a foreground image in which a foreground region corresponding to a predetermined object, such as a person or a ball, is extracted, and a background image in which a background region other than the foreground region is extracted are obtained from the multiple captured images. Furthermore, a foreground model representing the three-dimensional shape of the predetermined object and texture data for coloring the foreground model are generated based on the foreground image, and texture data for coloring a background model representing the three-dimensional shape of a background, such as a stadium, is generated based on the background image. Then, the texture data is mapped to the foreground model and background model, and rendering is performed according to the virtual viewpoint indicated by the viewpoint information, thereby generating a virtual viewpoint image. However, the method for generating a virtual viewpoint image is not limited to this, and various methods can be used, such as a method of generating a virtual viewpoint image by projective transformation of captured images without using a three-dimensional model.
[0034] A foreground image is an image in which an object region (foreground region) is extracted from an image captured by an imaging device. An object extracted as a foreground region is a dynamic object (moving body) that moves (its absolute position and shape can change) when images are captured from the same direction in a time series. Examples of objects include players, referees, and other people on the field where a sport is being played, such as a ball in a ball game, or singers, musicians, performers, and presenters in a concert or entertainment event.
[0035] A background image is an image of at least a region (background region) different from the foreground object. Specifically, a background image is an image in which the foreground object has been removed from the captured image. Furthermore, the background refers to an imaged object that remains stationary or nearly stationary when images are captured from the same direction in chronological order. Examples of such imaged objects include a stage for a concert, a stadium where an event such as a sport is held, a structure such as a goal used in a ball game, or a field. However, the background is at least a region different from the foreground object, and the imaged object may include other objects in addition to the object and background.
[0036] A virtual camera is a virtual camera that is different from the multiple imaging devices actually installed around the imaging area, and is a concept for conveniently explaining a virtual viewpoint related to the generation of a virtual viewpoint image. That is, a virtual viewpoint image can be considered to be an image captured from a virtual viewpoint set in a virtual space associated with the imaging area. The position and orientation of the viewpoint in the virtual image capture can be expressed as the position and orientation of the virtual camera. In other words, a virtual viewpoint image can be said to be an image that simulates an image captured by a camera if it were assumed that the camera were located at the virtual viewpoint set in space.
[0037] Example 1 In this embodiment, for a semi-transparent subject, an area where another subject exists behind the semi-transparent subject is detected, and the corresponding area is selected so as not to be used for drawing.
[0038] 1 is a diagram showing an example of the configuration of an image processing system in Example 1. The image processing system includes an imaging device 101, an image processing device 108, and an output device 107. The image processing device 108 includes a transmittance map generation unit 102, a three-dimensional shape estimation unit 103, a virtual viewpoint acquisition unit 104, a texture selection unit 105, and a virtual viewpoint image generation unit 106.
[0039] The image processing system generates a virtual viewpoint image representing a scene from a specified virtual viewpoint based on multiple images captured by multiple imaging devices and a specified virtual viewpoint. The virtual viewpoint image in this embodiment is also called a free viewpoint video, but is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by a user. For example, the virtual viewpoint image may also include an image corresponding to a viewpoint selected by a user from multiple options. In this embodiment, the virtual viewpoint is mainly specified by a user operation, but the virtual viewpoint may also be automatically specified based on the results of image analysis, etc. In this embodiment, the virtual viewpoint image is mainly described as a moving image, but the virtual viewpoint image may also be a still image. Note that the image processing system may be configured such that each component is composed of a single electronic device or multiple electronic devices.
[0040] The image capture device 101 represents a plurality of physical cameras. These physical cameras are arranged at different positions and capture images of a subject from multiple viewpoints in a synchronized manner. The captured images and viewpoint information (external parameters, internal parameters, image size, focal length) of the image capture devices are then transmitted to the transmittance map generation unit 102 and the texture selection unit 105. There is no particular limit on the number of cameras. The external parameters of the image capture device include position information indicating the position of the image capture device and orientation information indicating the orientation of the image capture device. The image capture range of each image capture device can be identified by using the viewpoint information.
[0041] The transmittance map generation unit 102 generates a transmittance map by converting the image captured by the imaging device 101 into transmittance information representing the transmittance of each pixel of the captured image of the subject. The transmittance map is a multi-value mask with higher values corresponding to higher transmittance of the subject in the captured image. For example, for the scene shown in FIG. 3(a), the image captured by camera 302b is shown in FIG. 3(b). For this captured image, the transmittance map generated by the transmittance map generation unit 102 is shown in FIG. 3(c). The transmittance map can be generated by calculating the transmittance of the foreground using a background image or background color acquired in advance, as in Patent Document 1, for example. Alternatively, the transmittance map may be inferred using a machine learning technique. Of course, opacity may be calculated instead of transmittance.
[0042] The three-dimensional shape estimation unit 103 estimates a three-dimensional shape including transparency information using the transparency map generated by the transparency map generation unit 102. The three-dimensional shape estimation method is not particularly limited, and may be, for example, a volume intersection method or a stereo method. To include transparency information in the three-dimensional shape, the following process may be used, for example: The transparency map is binarized based on multiple different thresholds to generate multiple foreground maps representing foreground regions in the captured image. The threshold may be, for example, a median value. Alternatively, a histogram of transmittance may be created and its minimum point may be used as the threshold. The foreground region in the captured image is a two-dimensional region in which opaque voxels exist in the three-dimensional shape to be estimated, as viewed from the viewpoint of the imaging device. The background region is a two-dimensional region in which only transparent voxels exist. The foreground map is an image that represents opaque regions (foreground regions) and transparent regions (background regions) in binary. Depending on the threshold setting, semi-transparent regions of the subject in the captured image are determined as foreground or background regions. Multiple three-dimensional shapes are then estimated using the foreground maps with the respective thresholds. In the multiple obtained three-dimensional shapes, the semi-transparent regions of the subject are represented by opaque or transparent voxels depending on the threshold. Transparency information for the subject can be obtained by referencing the differences in the three-dimensional shapes. Hereinafter, this embodiment will describe processing when two thresholds are set. Using one threshold, all semi-transparent regions of the subject are designated as the foreground region, and a three-dimensional shape is estimated in which the entire subject is made up of opaque voxels, including both semi-transparent and opaque regions. Hereinafter, this three-dimensional shape will be referred to as the semi-transparent foreground three-dimensional shape. In other words, the semi-transparent foreground three-dimensional shape is a three-dimensional shape that includes both semi-transparent and opaque regions. Using another threshold, all semi-transparent regions of the subject are designated as the background region, and a three-dimensional shape showing only the opaque regions is estimated by identifying voxels from the opaque regions of the subject. Hereinafter, this three-dimensional shape will be referred to as the opaque foreground three-dimensional shape. In this embodiment, the semi-transparent foreground three-dimensional shape and the opaque foreground three-dimensional shape are collectively referred to as a three-dimensional shape containing transparency information. However, the number of thresholds may be one or more, and two or more types of three-dimensional shapes may be generated. In addition, in this embodiment, the three-dimensional shape is represented by the transparency or opacity of voxels, but for example, a value indicating whether a subject is present in every voxel may be stored, and the shape may be obtained by referencing the value.Since a translucent foreground three-dimensional shape includes both translucent and opaque regions, additional information indicating which region the component is included in may be added to each component that makes up the translucent foreground three-dimensional shape. This additional information may be expressed as a binary value, for example, 0 when the component is included in a translucent region and 1 when the component is included in an opaque region. Of course, these numbers may be set in reverse, or may be expressed as True or False.
[0043] The virtual viewpoint acquisition unit 104 acquires viewpoint information of a virtual viewpoint for drawing a virtual viewpoint image. The viewpoint information of the virtual viewpoint includes at least information such as the position of the virtual viewpoint, the line of sight direction from the virtual viewpoint, and the angle of view. The viewpoint information of the virtual viewpoint is associated with a frame number or a time code attached to a captured image. The viewpoint information of the virtual viewpoint is specified by an operator operating an input device such as a mouse or keyboard. Alternatively, viewpoint information of temporally consecutive virtual viewpoints generated in advance may be acquired from a storage device (not shown).
[0044] The texture selection unit 105 uses the captured image, the three-dimensional shape including transparency information, and viewpoint information of the virtual viewpoint to select a texture from the captured image to be used when generating an appearance from the virtual viewpoint in the virtual viewpoint image generation unit 106. In this embodiment, the captured image to be used as the texture for the two three-dimensional shapes from the three-dimensional shape estimation unit 103 is selected. Hereinafter, this process will be referred to as texture selection processing.
[0045] The virtual viewpoint image generation unit 106 acquires the three-dimensional shape acquired from the three-dimensional shape estimation unit 103, the transmittance map acquired from the transmittance map generation unit 102, the captured image acquired from the texture selection unit 105, and viewpoint information of the virtual viewpoint from the virtual viewpoint acquisition unit 104. Using the acquired information, a virtual viewpoint image including a semi-transparent region of the subject and a virtual viewpoint image including an opaque region are generated, and these are combined to generate a virtual viewpoint image including an opaque subject. The virtual viewpoint image may be drawn using, for example, the Z-sort method. Details will be described later using the flowchart of FIG. 7.
[0046] The output device 107 outputs the virtual viewpoint image generated by the virtual viewpoint image generating unit 106 and displays it on a display device such as a display, or may transmit it to a storage device such as a server.
[0047] The image processing device 108 is a PC or a tablet terminal, and may have a display means (not shown).
[0048] FIG. 2 is a flowchart showing the process of generating a virtual viewpoint image by the image processing system in this embodiment.
[0049] In S201, the image capturing device 101 captures images of a subject using a plurality of image capturing devices to obtain captured images. The obtained captured images are output to the transmittance map generating unit 102 and the texture selecting unit 105.
[0050] In S202, the transmittance map generation unit 102 generates a plurality of transmittance maps corresponding to a plurality of captured images using the trained model. The transmittance map is a map that represents the transmittance information of the subject in the captured image, and details will be explained later in the section on the transmittance map generation unit 102. The generated plurality of transmittance maps are output to the three-dimensional shape estimation unit 103 and the virtual viewpoint image generation unit 106.
[0051] In S203, the three-dimensional shape estimation unit 103 estimates a three-dimensional shape including transparency information. The three-dimensional shape including transparency information is used by the texture selection unit 105 and the virtual viewpoint image generation unit 106.
[0052] In S204, the virtual viewpoint acquisition unit 104 acquires viewpoint information of the virtual viewpoint. The viewpoint information of the virtual viewpoint is used by the texture selection unit 105 and the virtual viewpoint image generation unit .
[0053] In S205, the texture selection unit 105 selects a texture (captured image) to be used in the process of determining color information of each component of the three-dimensional shape. The selected texture is used by the virtual viewpoint image generation unit .
[0054] In S206, the virtual viewpoint image generating unit 106 generates a virtual viewpoint image. The generated virtual viewpoint image is output to the output device 107.
[0055] FIG. 3 is a diagram illustrating an overview of a scene in which this embodiment is implemented. FIG. 3(a) is a diagram illustrating a scene in which subjects 301a and 301b are captured by the imaging device 101. Note that part of the subject 301a includes a semi-transparent portion. Specifically, the dark gray portions are opaque portions, and the light gray portions are semi-transparent portions. The subject 301b is an opaque subject. Cameras 302a to 302c act as imaging devices to capture the subjects 301a and 301b. A virtual viewpoint 303 is located at the position shown in the figure, and the view from this virtual viewpoint is generated as a virtual viewpoint image. FIG. 3(b) shows an image captured by camera 302b. FIG. 3(c) is a transmittance map generated for the captured image of FIG. 3(b). In this embodiment, higher transmittance is displayed as black, and lower transmittance is displayed as white.
[0056] FIG. 4 is a flowchart showing the process of selecting a texture by the texture selection unit 105 in this embodiment.
[0057] In S401, depth information of the three-dimensional shape is calculated from each image capture device. The depth information is, for example, the distance from the image capture device to the surface of each three-dimensional shape. In this embodiment, there are two types of three-dimensional shapes: a semi-transparent foreground three-dimensional shape and an opaque foreground three-dimensional shape, and depth information is calculated for each of them. It is also possible to calculate and save depth information for each captured image in advance.
[0058] In S402, it is determined whether the processes of S403 to S406, which will be described later, have been completed for all target voxels. The target voxels may be all voxels, or only surface voxels. They may also be voxels visible from the virtual viewpoint. When the selection of captured images to be used as textures for all target voxels has been completed, the processing of the texture selection unit 105 ends. If not, the processes of S403 to S406 are performed for the remaining voxels.
[0059] In S403, it is determined whether the processes of S404 and S405 (described later) have been completed for all image capture devices for the target voxel for which the captured image to be used for texture is selected. If completed, the process proceeds to S406. If not completed, the processes of S404 and S405 are repeated for the remaining image capture devices.
[0060] In S404, back projection is performed from the target voxel to the target imaging device, and the distance from the target voxel to the target imaging device is obtained.
[0061] In S405, using the depth information from S401 and the distance from S404, it is determined whether or not the image captured by the target imaging device can be used as a texture when generating a virtual viewpoint image of the target voxel. The method of determination differs depending on the transparency information contained in the three-dimensional space. Details of the determination will be described later. The result of the determination may be a value indicating whether or not the target imaging device can be used as a texture for the target voxel, for example, flag information indicating true if it can be used and no if it cannot. The determination result may be stored as a voxel value, or may be saved in a separate table so as to be linked to the voxel.
[0062] In S406, a captured image to be used for the texture used when generating a virtual viewpoint image for a target voxel is selected according to the viewpoint information of the virtual viewpoint obtained by the virtual viewpoint acquisition unit 104. For example, the selection method involves selecting, as texture, one or more captured images from one or more image capture devices whose position or line of sight direction is closest to the virtual viewpoint from among the captured images determined to be usable in S405. At the same time, in addition to the captured image, a transmittance map of the corresponding captured image may be selected as the transmittance of the texture. When selecting, the distance or angle difference from the image capture device to the virtual viewpoint may be normalized and applied as a coefficient to the texture or captured image. Furthermore, in S406, the selection process may be performed after determining whether the target voxel is directly visible from the virtual viewpoint. The selection process for the captured image to be used for the texture may be performed only for voxels visible from the virtual viewpoint, and the texture selection process may be omitted for voxels outside the angle of view of the virtual viewpoint, voxels that are not on the surface, or voxels that are occluded by other voxels. The processes from S401 to S406 complete the selection of the captured image to be used for generating a virtual viewpoint of a three-dimensional shape as a texture, and the captured image to be used for the texture selected in S406 may be stored so as to be associated with each voxel of the three-dimensional shape. That is, a unique number or a unique captured image name representing the imaging device may be assigned to a captured image that can be used on a voxel-by-voxel basis.
[0063] The determination of captured images that can be used as textures in S405 will be described in detail with reference to Figures 5 and 6. The determination method differs depending on the type of three-dimensional shape. In this embodiment, of the depth information calculated in S402, the depth information up to the surface of the semi-transparent foreground three-dimensional shape of the three-dimensional shape estimation unit 103 is referred to as depth A, and the depth information up to the surface of the opaque foreground three-dimensional shape is referred to as depth B.
[0064] FIG. 5 illustrates the process of determining captured images that can be used as textures for translucent foreground shapes. In other words, this process selects captured images to be used to determine the color information of each component of a translucent foreground three-dimensional shape. First, we will explain the method for determining a translucent foreground three-dimensional shape using FIG. 5. FIG. 5 shows an overhead view of the surface voxels of the translucent foreground three-dimensional shape for the scene in FIG. 3(a). Two specific determinations are made for a translucent foreground three-dimensional shape. If both are met, the image captured by the target imaging device is determined to be usable as a texture. The first is a visibility determination to determine whether the target voxel is directly visible from the target imaging device. Visibility is required to determine that a texture can be used. If the depth A determined in S402 and the distance from the target voxel to the target imaging device determined in S404 are the same or less than a predetermined threshold, visibility is determined. The predetermined threshold can be, for example, a fixed value or a value determined based on the target subject. It can be determined based on the thickness and range of typical translucent clothing. Figure 5(a) shows an example where visibility is not determined because depth A and distance are greater than or equal to a predetermined threshold. Figure 5(b) shows an example where visibility is determined because depth A and distance match. The second determination is to determine whether there are other objects on the extension line (straight line) connecting the target imaging device to the target voxel. To determine that the texture can be used, there must be no other objects on the extension line. Figure 5(c) shows an example where there is another subject on the extension line, and the image captured by the target imaging device is determined to be unusable as a texture. Figure 5(d) shows an example where there are no other objects on the extension line. Figure 5(d) is also determined to be visible, so the image captured by the target imaging device is determined to be usable as a texture. The presence of other objects on the extension line can be determined, for example, by counting the number of voxels that form surfaces in the direction of the extension line. If the number is two or more than a predetermined number, it is determined that the extension line penetrates multiple subjects, and it is determined that there are other objects on the extension line. These two judgments allow the captured image to be selected so that the target voxel is directly visible, there are no other objects behind it in the semi-transparent area, and there is no mixing of texture colors, and it is determined that the image can be used as a texture.
[0065] FIG. 6 illustrates the process of determining whether a captured image can be used as a texture for an opaque foreground shape. FIG. 6 shows an overhead view of the scene in FIG. 3(a), with an opaque foreground 3D shape (a dark rectangle) and a semi-transparent foreground 3D shape (a light rectangle) superimposed. Two similar determinations are made for the opaque foreground 3D shape. If both are met, the image captured by the target imaging device is determined to be usable as a texture. The first is a visibility determination, similar to the determination for the 3D shape with the semi-transparent region in the foreground. The determination method can be the same. The second determination is whether depth A and depth B match or the difference is less than a predetermined threshold along the line connecting the target imaging device to the target voxel. A match or a difference less than the predetermined threshold is required for the image to be determined to be usable as a texture. For example, in FIG. 6(a), the difference between depth A and depth B is greater than or equal to the predetermined threshold, so the image captured by the imaging device is determined to be unusable as a texture. In Figure 6(b), depth A and depth B match and are visible, so the texture is determined to be usable. Also, as in Figure 6(c), the difference between depth A and depth B is less than a predetermined threshold, and depth B matches the distance from the target voxel to the target imaging device, so visibility is established and the texture is determined to be usable. These two determinations allow captured images to be selected in which the target voxel is directly visible, there are no other objects in front of the target voxel, or only thin, translucent objects are overlapping, and there is no intermixing of texture colors. These captured images are determined to be usable as textures. Furthermore, the format of the selected texture may be output by simultaneously assigning the determination results of S405 and S406 for the target voxel on a pixel-by-pixel basis for all captured images, including those determined to be unusable in S405 and S406. The determination results may be, for example, flag information with a value of true if the texture is usable and a value of no if the texture is overloaded. Alternatively, the results of the two determinations made in S405 may be assigned separately.
[0066] FIG. 7 is a flowchart showing the process of generating a virtual viewpoint image by the virtual viewpoint image generating unit 106 in this embodiment.
[0067] In S701, depth information indicating the depth (distance) from the virtual viewpoint to the three-dimensional shape including transparency information is calculated based on viewpoint information of the virtual viewpoint acquired from the virtual viewpoint acquisition unit 104. In this embodiment, depth information is calculated for each of the semi-transparent foreground three-dimensional shape and the opaque foreground three-dimensional shape.
[0068] In S702, a virtual viewpoint image of an opaque object is rendered based on the information of the selected texture using the captured image, transmittance map, opaque foreground three-dimensional shape, and depth information from the virtual viewpoint. The rendering method involves, for example, using depth information to identify voxels that constitute the surface of the three-dimensional shape when viewed from the virtual viewpoint. Then, for the identified surface voxels, the texture selected by the process described in FIG. 4 is used as the pixel color of the voxel to be rendered. Furthermore, if there are multiple selected textures for surface voxels, they may be blended to obtain the rendering color. The blending coefficient may be, for example, a normalized distance or angle difference from the imaging device to the virtual viewpoint. Furthermore, as in Patent Document 1, a transmittance map may be used to remove the background color from the selected texture. When rendering, in addition to color information, the transmittance of the selected texture may be used as the transmittance of the target voxel. For the scene shown in FIG. 3(a), the rendering result in S702 is shown in FIG. 8(a). This process is performed for all voxels determined to constitute the surface of the object.
[0069] In S703, a virtual viewpoint image of the translucent subject is drawn based on the information of the selected texture using the captured image, transmittance map, translucent foreground 3D shape, and depth information from the virtual viewpoint. To draw only the translucent area from the translucent foreground 3D shape, the translucent foreground 3D shape and the opaque foreground 3D shape are compared, and opaque voxels that exist only in the translucent foreground 3D shape are drawn. The drawing method can be the same as in S702. For the scene in Figure 3(a), the drawing result in S702 is shown in Figure 8(b).
[0070] In S704, the rendering result of the opaque object in S702 is overwritten with the rendering result of S703 in the area where only the translucent object exists, and the two are combined. The combining is performed based on the transmittance of the rendering result of S703. For the scene in FIG. 3(a), the resulting combination in S704 is as shown in FIG. 8(c). When rendering a translucent object in S703, the texture of the opaque object may be partially used depending on the threshold setting in S405 of the texture selection unit 105. For example, as shown in FIG. 6(c), if a thin translucent portion exists on the surface of an opaque object and the threshold setting in S405 is high, a texture containing the opaque object will be selected. However, since the rendering result of the opaque object is overwritten during the combining in S704, a natural virtual viewpoint is ultimately obtained. Note that before combining the rendering result of the opaque object with the rendering result of the translucent object, color information of the background image may be removed from the rendering result of the translucent object, i.e., from the area of the translucent object in the virtual viewpoint image containing the translucent object.
[0071] 9 is a hardware configuration diagram of the image processing device 108. The image processing device 108 has a calculation unit for performing image processing and three-dimensional shape generation, including a GPU (Graphics Processing Unit) 910 and a CPU (Central Processing Unit) 911. The image processing device 108 also has a storage unit including a ROM (Read Only Memory) 912, a RAM (Random Access Memory) 913, and an auxiliary storage device 914. The image processing device 108 also has a display unit 915, an operation unit 916, a communication I / F 917, and a bus 918.
[0072] The CPU 911 controls the entire image processing device 108 using computer programs and data stored in the ROM 912 and RAM 913, thereby realizing each function of the image processing device 108. The CPU 911 also operates as a display control unit that controls the display unit 915 and an operation control unit that controls the operation unit 916.
[0073] The GPU 910 can perform efficient calculations by processing a larger amount of data in parallel. Therefore, in this embodiment, the transmittance map generation unit 102, the three-dimensional shape estimation unit 103, the texture selection unit 105, and the virtual viewpoint image generation unit 106 use the GPU 910 in addition to the CPU 911. When executing a program, calculations may be performed by either the CPU 911 or the GPU 910 alone, or the CPU 911 and the GPU 910 may work together to perform calculations.
[0074] The image processing device 108 may have one or more pieces of dedicated hardware different from the CPU 911, and the dedicated hardware may execute at least a part of the processing by the CPU 911. Examples of the dedicated hardware include an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), and a DSP (digital signal processor).
[0075] The ROM 912 stores programs that do not require modification. The RAM 913 temporarily stores programs and data supplied from an auxiliary storage device 914, and data supplied from the outside via a communication I / F 917. The auxiliary storage device 914 is configured, for example, by a hard disk drive, and stores various data such as image data and audio data.
[0076] The display unit 915 is configured with, for example, a liquid crystal display, an LED, etc., and displays a GUI (Graphical User Interface) etc. for the user to operate the information processing device. The operation unit 916 is configured with, for example, a keyboard, a mouse, a joystick, a touch panel, etc., and accepts operations by the user and inputs various instructions to the CPU 911.
[0077] The communication I / F 917 is used for communication between the information processing device and an external device. For example, when the information processing device is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 917. When the information processing device has a function for wireless communication with an external device, the communication I / F 917 is equipped with an antenna. The bus 918 connects each part of the information processing device to transmit information.
[0078] In this embodiment, for a semi-transparent subject, a transmittance map is generated, a three-dimensional shape including transparency information is estimated, and a texture to be used for generating a virtual viewpoint image is selected from a captured image, and a virtual viewpoint image is generated. This embodiment makes it possible to use an appropriate texture, and a natural virtual viewpoint can be obtained.
[0079] In this embodiment, a three-dimensional shape including transparency information is generated by using multiple binary three-dimensional shapes with transparent regions as the foreground or background, but it may also be generated as a single multi-valued three-dimensional shape. In that case, when each process is performed by the texture selection unit 105 and the virtual viewpoint image generation unit 106, a threshold value is set for the voxel value to be the foreground or background, and the process is performed after binarization.
[0080] <Example 2> In the first embodiment, the texture selection unit 105 does not use a captured image containing an overlapping area between a semi-transparent object and another object when rendering the captured image, so that an appropriate captured image is used as the texture. However, depending on the position and orientation of the virtual viewpoint and the image capture device, and the positional relationship between the image capture device and the object, if such captured images are usable, a more natural virtual viewpoint image may be obtained. For example, if an image capture device close to the virtual viewpoint is determined to be unusable and an image capture device located far from the virtual viewpoint is used, the resolution of the texture may decrease, creating an unnatural feeling. Furthermore, it is possible that the appearance of an anisotropic object from the virtual viewpoint cannot be accurately reproduced. Therefore, in the second embodiment, for a captured image containing an overlapping area between a semi-transparent object and another object, the captured image of the overlapping area is corrected and made usable as a texture to generate a virtual viewpoint image.
[0081] 10 is a diagram showing the overall configuration of an image processing system. Hereinafter, differences between this embodiment and embodiment 1 will be mainly described, and components with the same names as those in embodiment 1 will not be described as they are the same. The difference between this embodiment and embodiment 1 is that a captured image correction unit 109 is added. This device may be configured by one electronic device or by multiple electronic devices.
[0082] The captured image correction unit 109 selects a camera for correcting the texture based on the information of the selected captured image, and corrects the captured image that cannot be used as texture using information on the position and angle of view of the imaging device 101. The corrected image becomes usable as texture.
[0083] FIG. 11 is a flowchart showing the process of the captured image correction unit 109 correcting a captured image.
[0084] In S1101, a captured image to be corrected is selected. The captured image to be used for the texture to be corrected is selected using viewpoint information of the virtual viewpoint and the determination results of S405 and S406 assigned to each pixel of the captured image by the texture selection unit 105. That is, the determination results of usability based on visibility and usability based on the relationship between the position of the virtual viewpoint and the position of the image capture device are used. For example, if visibility is determined to be present in S405 but another object is present on the extension line from the target image capture device to the target voxel, and the captured image is determined to be unusable as a texture, the captured image is selected as a correction target. Furthermore, based on the determination result of S406, the captured image not selected as a texture is selected as a correction target. Furthermore, the image capture device closest to the position or orientation of the virtual viewpoint may be selected based on viewpoint information of the virtual viewpoint.
[0085] In S1102, the imaging device that captured the image to be corrected is used as another virtual viewpoint, and a virtual viewpoint image of an opaque subject from the virtual viewpoint at the position of the imaging device is generated in the same manner as in S702 of the virtual viewpoint image generation unit 106. This generates an appearance when no semi-transparent subject is present from the imaging device.
[0086] In S1103, the captured image to be corrected and the virtual viewpoint image obtained in S1102 are used to perform correction so as to remove the color of other objects overlapping the texture of the semi-transparent object. The correction method, as in Patent Document 1, for example, calculates a transmittance map for the foreground image by using the captured image as a foreground image and the virtual viewpoint image as a background image. The transmittance map is used to remove the color of the virtual viewpoint image, which serves as the background image, from the captured image. The captured image corrected by the captured image correction unit 109 is changed to a texture that can be used, and is used by the virtual viewpoint image generation unit 106.
[0087] As a result, it is possible to correct the texture of areas where other objects overlap with semi-transparent areas of a captured image. This increases the number of textures that can be used and allows multiple captured images to be blended for coloring, resulting in a more natural virtual viewpoint.
[0088] The present disclosure can also be realized by providing a program that implements one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more functions.
[0089] The disclosure of this embodiment includes the following configurations, methods, and programs.
[0090] (Configuration 1) an acquisition means for acquiring viewpoint information indicating a position of a virtual viewpoint and a line of sight direction from the virtual viewpoint; a first generation means for generating a first virtual viewpoint image including a transparent or semi-transparent first portion of a subject, using the viewpoint information and a first captured image including a plurality of pixels corresponding to the first portion that do not correspond to an opaque second portion of the subject captured through the first portion; and for generating a second virtual viewpoint image including the second portion, using a second captured image including the second portion and the viewpoint information; a second generating means for generating a third virtual viewpoint image including the first portion and the second portion based on the first virtual viewpoint image and the second virtual viewpoint image; A system comprising:
[0091] (Configuration 2) a pixel in the first captured image that does not correspond to the second portion captured through the first portion corresponds to a specific pixel of the first portion in the first virtual viewpoint image; The system described in configuration 1 is characterized in that the first generation means determines color information of the specific pixel in the first virtual viewpoint image using color information of a pixel that does not correspond to the second portion imaged through the first portion.
[0092] (Configuration 3) The system described in configuration 2 is characterized in that pixels in the first captured image that do not correspond to the second part imaged through the first part and specific pixels of the first part in the first virtual viewpoint image correspond to specific components of the three-dimensional shape of the subject.
[0093] (Configuration 4) The system according to any one of configurations 1 to 3, wherein the second captured image is a captured image including the second portion captured through the first portion.
[0094] (Configuration 5) 5. The system according to any one of configurations 1 to 4, wherein the first captured image and the second captured image are the same captured image.
[0095] (Configuration 6) The system described in any one of configurations 1 to 5, characterized in that the first generation means determines color information of the second part in the second virtual viewpoint image using color information obtained by removing color information corresponding to the first part from color information of pixels corresponding to the second part imaged through the first part in the first captured image.
[0096] (Configuration 7) the acquisition means acquires a plurality of captured images including the subject, and acquires a plurality of pieces of transmittance information corresponding to the plurality of captured images using a trained model that receives the plurality of captured images as input and outputs a plurality of pieces of transmittance information indicating the transmittance of regions corresponding to the subject in the plurality of captured images; The system described in any one of configurations 1 to 6, further comprising an identification means for identifying the first captured image and the second captured image based on the plurality of transmittance information of the plurality of captured images.
[0097] (Configuration 8) the first generation means generates shape information indicating a three-dimensional shape of the subject including a transparent or semi-transparent first component and an opaque second component based on the plurality of captured images, position information of a plurality of imaging devices that captured the plurality of captured images, and the plurality of pieces of transmittance information; The system described in Configuration 7 is characterized in that the identification means identifies, among the plurality of captured images, an captured image that includes the first portion in the imaging range of the imaging device and in which the second component does not exist on a straight line passing through the first component from a position in virtual space corresponding to a position in real space of the imaging device, as the first captured image, and identifies an captured image that includes the second portion in the imaging range of the imaging device, as an captured image that includes the second portion.
[0098] (Configuration 9) the specifying means specifies, in each of the plurality of captured images, an area corresponding to a subject whose transmittance is equal to or greater than a threshold as a transparent or semi-transparent area, and specifies an area corresponding to a subject whose transmittance is less than the threshold as an opaque area; The system described in configuration 8 is characterized in that the first generation means generates the first components using the transparent or semi-transparent areas in the plurality of captured images, and generates the second components using the opaque areas in the plurality of captured images, thereby generating shape information indicating the three-dimensional shape of the subject.
[0099] (Configuration 10) The system described in any one of configurations 1 to 9, characterized in that the second generation means generates the third virtual viewpoint image by removing background color from an area corresponding to the first portion of the first virtual viewpoint image and combining the first virtual viewpoint image and the second virtual viewpoint image.
[0100] (method) an acquisition step of acquiring viewpoint information indicating a position of a virtual viewpoint and a line of sight direction from the virtual viewpoint; a first generation step of generating a first virtual viewpoint image including a transparent or semi-transparent first portion of a subject, using the viewpoint information and a first captured image including a plurality of pixels corresponding to the first portion, the first portion including pixels that do not correspond to an opaque second portion of the subject captured through the first portion; and generating a second virtual viewpoint image including the second portion using the viewpoint information and a second captured image including the second portion; a second generation step of generating a third virtual viewpoint image including the first portion and the second portion based on the first virtual viewpoint image and the second virtual viewpoint image; An image processing method comprising:
[0101] (program) A computer program for controlling each means of the system according to any one of configurations 1 to 10 by a computer. [Explanation of symbols]
[0102] 102 Transmittance map generator 103 3D shape estimation part 104 Virtual viewpoint acquisition unit 105 Texture Selection 106 Virtual viewpoint image generation unit
Claims
1. an acquisition means for acquiring viewpoint information indicating a position of a virtual viewpoint and a line of sight direction from the virtual viewpoint; a first generation means for generating a first virtual viewpoint image including a transparent or semi-transparent first portion of a subject, using the viewpoint information and a first captured image including a plurality of pixels corresponding to the first portion, the pixels including pixels that do not correspond to an opaque second portion of the subject captured through the first portion; and for generating a second virtual viewpoint image including the second portion, using the viewpoint information and a second captured image including the second portion; a second generating means for generating a third virtual viewpoint image including the first portion and the second portion based on the first virtual viewpoint image and the second virtual viewpoint image; An image processing system comprising:
2. a pixel in the first captured image that does not correspond to the second portion captured through the first portion corresponds to a specific pixel in the first portion in the first virtual viewpoint image; The image processing system according to claim 1, characterized in that the first generation means determines the color information of the specific pixel in the first virtual viewpoint image using color information of a pixel that does not correspond to the second portion imaged through the first portion.
3. The image processing system described in claim 2, characterized in that pixels in the first captured image that do not correspond to the second part captured through the first part and specific pixels of the first part in the first virtual viewpoint image correspond to specific components of the three-dimensional shape of the subject.
4. 2. The image processing system according to claim 1, wherein the second captured image is a captured image including the second portion captured through the first portion.
5. 2. The image processing system according to claim 1, wherein the first captured image and the second captured image are the same captured image.
6. The image processing system described in claim 1, characterized in that the first generation means determines the color information of the second part in the second virtual viewpoint image using color information obtained by removing color information corresponding to the first part from color information of pixels corresponding to the second part imaged through the first part in the first captured image.
7. the acquisition means acquires a plurality of captured images including the subject, and acquires a plurality of pieces of transmittance information corresponding to the plurality of captured images using a trained model that receives the plurality of captured images as input and outputs a plurality of pieces of transmittance information indicating the transmittance of regions corresponding to the subject in the plurality of captured images; 2. The image processing system according to claim 1, further comprising: a specifying unit that specifies the first captured image and the second captured image based on the plurality of pieces of transmittance information from among the plurality of captured images.
8. the first generation means generates shape information indicating a three-dimensional shape of the subject including a transparent or semi-transparent first component and an opaque second component based on the plurality of captured images, position information of a plurality of imaging devices that captured the plurality of captured images, and the plurality of pieces of transmittance information; The image processing system described in claim 7, characterized in that the identification means identifies, among the multiple captured images, an captured image that includes the first portion in the imaging range of the imaging device and in which the second component does not exist on a straight line passing through the first component from a virtual space position corresponding to a real space position of the imaging device as the first captured image, and identifies an captured image that includes the second portion in the imaging range of the imaging device as an captured image including the second portion.
9. the specifying means specifies, in each of the plurality of captured images, an area corresponding to a subject whose transmittance is equal to or greater than a threshold as a transparent or semi-transparent area, and specifies an area corresponding to a subject whose transmittance is less than the threshold as an opaque area; The image processing system according to claim 8, characterized in that the first generation means generates the first components using the transparent or semi-transparent areas in the plurality of captured images, and generates the second components using the opaque areas in the plurality of captured images, thereby generating shape information indicating the three-dimensional shape of the subject.
10. The image processing system described in claim 1, characterized in that the second generation means generates the third virtual viewpoint image by removing background color from the area corresponding to the first portion of the first virtual viewpoint image and combining the first virtual viewpoint image and the second virtual viewpoint image.
11. an acquisition step of acquiring viewpoint information indicating a position of a virtual viewpoint and a line of sight direction from the virtual viewpoint; a first generation step of generating a first virtual viewpoint image including a transparent or semi-transparent first portion of a subject, using the viewpoint information and a first captured image including a plurality of pixels corresponding to the first portion, the first portion including pixels that do not correspond to an opaque second portion of the subject captured through the first portion; and generating a second virtual viewpoint image including the second portion using the viewpoint information and a second captured image including the second portion; a second generation step of generating a third virtual viewpoint image including the first portion and the second portion based on the first virtual viewpoint image and the second virtual viewpoint image; An image processing method comprising:
12. A computer program for controlling each unit of the image processing system according to any one of claims 1 to 10 by a computer.
Citation Information
Patent Citations
Method and device for chromakey processing
JP1994225329A