Image processing device, image processing method, and program
The image processing device addresses the issue of non-stationary obstacles by synthesizing object and occluding regions, enabling accurate three-dimensional shape data generation despite moving interferences.
Patent Information
- Application Number
- JP2021199189
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2041-12-08
AI Technical Summary
Existing methods for generating three-dimensional shape data are compromised by non-stationary obstacles, such as moving spectators, which cannot be effectively addressed by prior technologies.
An image processing device that identifies and generates a mask image by synthesizing the regions of both the object and other objects, using a method that includes detecting occluding regions and generating integrated masks to prevent occlusions, thereby allowing for accurate three-dimensional shape data creation.
Enables the generation of high-accuracy three-dimensional shape data even when non-stationary obstacles are present, ensuring precise representation of the object of interest.
Smart Images

Figure 0007802511000001 
Figure 0007802511000002 
Figure 0007802511000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to generating data based on captured images. [Background technology]
[0002] There is a method for generating three-dimensional shape data of an object using a mask image representing a two-dimensional silhouette of the object generated from multiple captured images taken by multiple image capture devices, and camera parameters. When generating three-dimensional shape data using this method, if an obstacle exists between the object for which three-dimensional shape data is to be generated and the image capture devices, the accuracy of generating the three-dimensional shape of the object may be reduced.
[0003] Patent Document 1 discloses a method for preventing defects in the three-dimensional shape of an object caused by the object being occluded by a stationary structure when the obstacle is the structure. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2019-106145 Summary of the Invention [Problem to be solved by the invention]
[0005] A person who is not stationary, such as a spectator, may become an obstacle that blocks an object for which three-dimensional shape data is to be generated. The method of Patent Document 1 cannot prevent defects from occurring in the three-dimensional shape when an object for which three-dimensional shape data is to be generated is blocked by a non-stationary obstacle. [Means for solving the problem]
[0006] The image processing device of the present disclosure includes: an acquisition means for acquiring a captured image;There may be other objects that may occlude the object for which three-dimensional shape data is to be generated, but are not the object for which three-dimensional shape data is to be generated. The aforementioned Identifying a first region in the captured image No. 1 Identification means; a detection means for detecting a region of the object to be generated from the captured image and detecting a region of the other object from the first region specified by the specification means; a synthesis means for generating a mask image by synthesizing the region of the object to be generated and the region of the other object; and a synthesis means for generating a mask image by using the mask image. generating means for generating three-dimensional shape data of the object to be generated; The area of the other object is a rectangular area including the other object. It is characterized by: [Effects of the Invention]
[0007] According to the technology of the present disclosure, it is possible to generate three-dimensional shape data of an object with high accuracy. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 2 is a diagram showing an example of the arrangement of imaging devices. [Figure 2] FIG. 1 is a diagram showing an example of the hardware configuration of a virtual viewpoint image generating device. [Figure 3] FIG. 1 is a diagram showing an example of the functional configuration of a virtual viewpoint image generating device. [Figure 4] FIG. 2 is a diagram showing an example of the functional configuration of an image processing unit. [Figure 5] FIG. 4 is a diagram showing an example of a captured image. [Figure 6] 10 is a flowchart for explaining processing by an image processing unit. [Figure 7] 5A and 5B are diagrams for explaining an occlusion candidate region and an object detection region in a captured image. [Figure 8] FIG. 10 is a diagram for explaining a shielding region. [Figure 9] 10A and 10B are diagrams showing examples of a foreground shape mask and a combined mask; [Figure 10] 1A and 1B are diagrams for explaining a method for generating a three-dimensional model using a volume intersection method. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, the technology of the present disclosure will be described in detail based on embodiments with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the technology of the present disclosure is not limited to the configurations shown in the drawings.
[0010] <Embodiment 1> [About virtual viewpoint images] There is a method in which multiple imaging devices are installed at different locations to capture images from multiple viewpoints in a time-synchronized manner, and the multiple images obtained by the capture are used to generate an image that represents a view from a virtual viewpoint that is independent of the viewpoints of the actual imaging devices. The image that represents a view from the virtual viewpoint generated by this method is called a virtual viewpoint image. The virtual viewpoint image allows a user to view highlight scenes of a sport such as soccer from various angles, thereby providing the user with a higher sense of realism than a normal captured image. Note that the virtual viewpoint image may be a video or a still image. In the following embodiment, the virtual viewpoint image will be described as a video.
[0011] A virtual viewpoint image is generated by generating three-dimensional shape data (also called a three-dimensional model) that represents the three-dimensional shape of a foreground object, placing the three-dimensional model in the background, and coloring the three-dimensional model to represent the view from the virtual viewpoint. The generation of the three-dimensional model and the rendering of the background use data based on captured images obtained by multiple image capture devices whose postures are determined in advance, calibrated, and capable of capturing images in a time-synchronized manner by inputting a common synchronization signal and time code. The object for which the three-dimensional model is generated is a subject that can be viewed from any angle of the virtual viewpoint, such as a player on a stadium field.
[0012] Fig. 1 is a diagram showing an example of the installation of imaging devices 101a-h. The cameras serving as imaging devices are arranged so as to capture the entire imaging space (stadium) in which an object for which a three-dimensional model is to be generated exists, as shown in Fig. 1. The imaging devices 101a-h then output captured images, to which a unique camera ID and a time code common to the camera array made up of the imaging devices 101a-h are assigned, as input images to a virtual viewpoint image generation device 200, which will be described later.
[0013] The multiple image capturing devices 101a-h are preferably installed in positions where obstacles that are not the target of generating a 3D model are not captured in the images. However, due to placement restrictions or to obtain high-quality texture images, some image capturing devices are installed in low positions near the object for which a 3D model is to be generated.
[0014] [Hardware configuration] 2 is a diagram showing the hardware configuration of a virtual viewpoint image generation device 200, which is an image processing device that generates a virtual viewpoint image based on images captured by multiple imaging devices. The virtual viewpoint image generation device 200 includes a CPU 201, a ROM 202, a RAM 203, an input I / F 205, a communication I / F 204, and a bus 206.
[0015] The CPU 201 uses computer programs and data stored in the ROM 202 and RAM 203 to control the entire virtual viewpoint image generation device 200, thereby realizing each function of the virtual viewpoint image generation device 200 shown in FIG. 3. The CPU 201 also performs arithmetic processing on an input image input from an input I / F 205. The virtual viewpoint image generation device 200 may be configured to include one or more dedicated hardware components different from the CPU 201 and to execute at least a part of the processing by the CPU 201. Examples of the dedicated hardware components include a processor used for image processing and control, an ASIC (application-specific integrated circuit), an FPGA (field programmable gate array), and a DSP (digital signal processor).
[0016] The ROM 202 stores programs that do not require modification, etc. The RAM 203 temporarily stores programs supplied from the ROM 202, data used to realize the functions of each functional block, and data supplied from the outside via a communication I / F 204. The input I / F 205 is a receiving unit such as SDI or HDMI (registered trademark) that acquires input images.
[0017] The communication I / F 204 is used for communication with an external device. For example, if the virtual viewpoint image generation device 200 is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 204. If the virtual viewpoint image generation device 200 has a function for wireless communication with an external device, the communication I / F 204 is equipped with an antenna. The bus 206 connects each unit of the virtual viewpoint image generation device 200 to transmit information.
[0018] Additionally, at least one of a display unit and an operation unit (not shown) may be included, or at least one of the display unit and the operation unit may exist as a separate external device. The display unit is configured, for example, by a liquid crystal display or LED, and displays a GUI (Graphical User Interface) or the like for the user to operate the virtual viewpoint image generation device 200. The operation unit is configured, for example, by a keyboard, mouse, joystick, touch panel, etc., and inputs various instructions to the CPU 201 in response to operations by the user. The CPU 201 operates as a display control unit that controls the display unit and an operation control unit that controls the operation unit.
[0019] [Functional configuration of the virtual viewpoint image generation device] 3 is a diagram showing an example of the functional configuration of the virtual viewpoint image generation device 200. The virtual viewpoint image generation device 200 includes an acquisition unit 301, a foreground extraction unit 302, an image processing unit 304, a background generation unit 303, a three-dimensional model generation unit 305, a control unit 307, and a drawing unit 306.
[0020] The acquisition unit 301 acquires captured images obtained by the multiple image capture devices 101a to 101h capturing images in a time-synchronized manner.
[0021] The foreground extraction unit 302 extracts a foreground region indicating the region of an object included in the captured image from the captured image of each of the imaging devices 101a-h. Then, it generates a mask image (called a foreground shape mask) that indicates the foreground region and non-foreground region of the captured image using binary values. It also generates a texture image (foreground texture) of the foreground object. The foreground extraction unit 302 assigns the camera ID and time code of each imaging device to the foreground texture and foreground shape mask. Then, it outputs the foreground texture and foreground shape mask to the three-dimensional model generation unit 305.
[0022] A background subtraction method is one method for extracting the foreground from a captured image. In this method, for example, an imaging environment in which no objects exist is captured and stored in advance as a background image. Then, an area in which the difference in pixel values between the captured image and the background image is greater than a threshold is determined to be the foreground. Note that the method for extracting the foreground is not limited to the method using background subtraction information. Other methods for extracting the foreground may also be used, such as a method using parallax, a method using feature values, or a method using machine learning.
[0023] The background generation unit 303 generates a background image to which a time code is assigned. The background image may be generated using the texture of a captured image with the same time code, or may be generated by assigning a time code to the data of an image to be used as the background of the virtual viewpoint image.
[0024] The image processing unit 304 determines areas (called "occluded areas") in which an object for which a three-dimensional model is to be generated may be occluded in each captured image of the image capturing devices 101a to 101d. The image processing unit 304 then outputs occluded area information, which is information indicating the position and shape of the occluded area. The occluded area information is added with time code information of the captured image to enable association with the captured image (frame). Furthermore, the camera ID of the image capturing device is added to indicate which image capturing device the occluded area corresponds to. Details of the processing by the image processing unit 304 will be described later.
[0025] The 3D model generation unit 305 acquires foreground texture, foreground shape mask, and masked area information data corresponding to all image capture devices that can be used to generate a 3D model. The foreground texture, foreground shape mask, background image, and masked area information each have a time code associated with the captured image. This allows the foreground texture, foreground shape mask, background image, and masked area information to be synchronized and processed using the respective data.
[0026] If the occluded area information includes information about an occluded area, the 3D model generation unit 305 merges the shape of the occluded area specified in the occluded area information with the foreground shape mask assigned the camera ID of the corresponding image capture device. The resulting mask image is called an integrated mask. The 3D model generation unit 305 also functions as a generation unit that generates integrated masks. Once the generation of integrated masks corresponding to all image capture devices is complete, a 3D model is generated based on the generated integrated masks using a volume intersection method. The 3D model generation unit 305 outputs the 3D model and foreground texture and notifies the rendering unit 306 that the generation of the 3D model is complete.
[0027] The rendering unit 306 acquires information on a background image, a three-dimensional model, a foreground texture, and a virtual viewpoint. Then, the rendering unit 306 colors the acquired three-dimensional model based on the foreground texture and superimposes the three-dimensional model on the background image. Then, the rendering unit 306 outputs an image projected onto two-dimensional coordinates from the virtual viewpoint as a virtual viewpoint image.
[0028] The control unit 307 controls each unit of the virtual viewpoint image generating device 200. For example, it acquires coordinate information indicating a virtual viewpoint designated by a user and performs control to generate a virtual viewpoint image corresponding to the virtual viewpoint. It also generates and outputs occlusion area detection information, which will be described later.
[0029] [Functional configuration of the image processing unit] 4 is a block diagram showing an example of the functions of the image processing unit 304. The image processing unit 304 includes a region setting unit 401, an object detection unit 402, and an occluded region determination unit 403.
[0030] The region setting unit 401 refers to the occluded region detection conditions and sets a region in the captured image where an occluding object is to be detected.
[0031] 5A and 5B are diagrams showing an example of an image captured by one of the imaging devices 101a to 101d. Fig. 5A shows an image 503 captured for camera calibration purposes, which is an image captured by capturing an image of a stadium, which is the imaging space, with no players or spectators present. The captured image includes a spectator seating area 501 and a court 502 in the background.
[0032] FIG. 5(b) shows a captured image 513 obtained by capturing an image of the image space during a competition. In other words, the captured image 513 is an example of a captured image for generating data such as a foreground shape mask used to generate a virtual viewpoint image. The captured image includes players 504 and 505 and spectators 506 to 512. If only players 505 and 505 are considered as foreground objects and a three-dimensional model is to be generated, depending on the position of the imaging device, the foreground objects may be partially occluded by other objects such as spectators 506 and 507. Objects such as spectators whose shapes and positions change and that may exist between the imaging device and the foreground objects (players) for which a three-dimensional model is to be generated are called occluding objects. An occluding object is not limited to a person; it can also be equipment whose posture or position changes.
[0033] The region setting unit 401 outputs occluding object detection information including information on the position and shape of a region (called an object detection region) in the captured image where an occluding object is to be detected. The occluding object detection information may also include a method for detecting an occluding object and filtering parameters used to determine the validity of the extracted object.
[0034] The object detection unit 402 detects objects from an object detection area in the captured image, and detects occluding objects from the detected objects.
[0035] The occluded area determination unit 403 determines an occluded area, which is an area that may occlude the foreground object, based on the detected occluding object. Details of the processing by the area setting unit 401, object detection unit 402, and occluded area determination unit 403 will be described using flowcharts.
[0036] [Details of image processing] Fig. 6 is a flowchart for explaining an example of the processing of the image processing unit 304. The series of processing shown in the flowchart of Fig. 6 is performed by the CPU of the virtual viewpoint image generation device 200 expanding program code stored in the ROM into the RAM and executing it. Also, some or all of the functions of the steps in Fig. 6 may be realized by hardware such as an ASIC or electronic circuit. Note that the symbol "S" in the explanation of each process indicates a step in the flowchart.
[0037] The following steps are performed on each of the captured images obtained by the imaging devices 101a to 101d, but the following description will explain the processing when the captured image of one of the imaging devices 101a to 101d is used as the input image.
[0038] In the case of a moving image, the process of the following flowchart is repeated to acquire frames that make up the moving image. Note that, for the process of the next frame and onward, at least one of S601 and S602 may be skipped.
[0039] In S601, the region setting unit 401 determines the position and shape of an occluded candidate region in the input image based on the occluded region detection information. A region in the input image where an occluding object may exist is called an occluded candidate region.
[0040] When determining an occlusion candidate region from the input image 701 in Fig. 7(a), for example, an area of the auditorium where an audience member who is an occlusion object may be present is determined as the occlusion candidate region. In this case, as shown in Fig. 7(b), the hatched area corresponding to the auditorium is set as the occlusion candidate region 702.
[0041] After the calibration is completed, the control unit 307 supplies the captured image used for the calibration to the image processing unit 304 and instructs the image processing unit 304 to perform initialization including setting of an occlusion candidate region, thereby executing the occlusion candidate region setting process. The initialization instruction includes parameters used for detecting an occlusion object.
[0042] For example, if the range of movement of an obstacle that may be an occluding object captured between the imaging device and the foreground object (player) is known, the occlusion candidate area is set based on the range of movement of the obstacle. Alternatively, the occlusion candidate area may be set from an area other than the area where the obstacle is not captured, based on an area where the obstacle is not captured. The range of movement of the obstacle can be estimated from the arrangement of the imaging devices 101a-d. For example, if there is an area where passersby pass or an area where equipment is installed between the foreground object and the camera, the area in the input image where passersby or equipment may be captured is the range of movement of the obstacle. The range of movement of the obstacle may be determined using a method of estimating the range of movement according to the characteristics of the obstacle whose shape has been detected. Alternatively, the occlusion candidate area may be set using a method of estimating an area where the shape of the area does not change for a certain period of time as a background area.
[0043] The region setting unit 401 may set the occlusion candidate region using information such as the shape of the handrails of the spectator seats or the color indicating the spectator seats. For example, if the imaging target is a stadium for rugby or the like, the occlusion candidate region may be set based on the sidelines drawn on the field. In this case, for example, an area a certain distance outside the sidelines is also included in the range for which a three-dimensional model is to be generated, so the occlusion candidate region is set to an area outside the range for which a three-dimensional model is to be generated.
[0044] In order to be able to set a candidate occlusion area based on information such as the color of the imaging space, the area setting unit 401 can be configured to include the operating mode or parameters used to detect the candidate occlusion area in the occlusion area detection information or external control instructions.
[0045] In addition, when a part of a pillar or beam whose shape does not change is captured between an imaging device whose position and posture are fixed and a foreground object (player), the pillar or beam can always be treated as an occluding object. In this case, the occluded area detection information may include a mask image that represents the shape of the pillar or beam in the captured image as the foreground area, or coordinate information that represents the shape of the pillar or beam, and a flag that indicates that the area of the pillar or beam is not included in the occlusion candidate area. Furthermore, the mask image that represents the shape of the pillar or beam may be used to generate an integrated mask by integrating it with a foreground shape mask.
[0046] Since the angle of view and the orientation of the camera are fixed, once an occlusion candidate area is set initially, the same occlusion candidate area will always be set in principle. For this reason, S601 may be skipped in the processing of the next frame and thereafter.
[0047] In S602, the region setting unit 401 sets a region (object detection region) for detecting an occluding object from the captured image based on the occluding candidate region, and saves shape information indicating the position and shape of the object detection region.
[0048] Figures 7(c) and 7(d) are diagrams showing an example of an object detection region set based on the occlusion candidate region 702 in Figure 7(b). The vertically lined regions in Figures 7(c) and 7(d) are object detection regions 703 and 704.
[0049] If the occlusion candidate region 702 in Figure 7(b) is set as the object detection region as is, when a spectator stands up, the spectator, who is an occluding object, may protrude outside the range of the occlusion candidate region 702 and occlude a player. For this reason, the object detection regions 703 and 704 in Figures 7(c) and 7(d) are set based on the region obtained by adding a preliminary region to the occlusion candidate region 702 in Figure 7(b).
[0050] Furthermore, when the shape of an occlusion candidate area is complex or when many occlusion candidate areas are scattered, it may be difficult to calculate preliminary areas according to the shapes of all occlusion candidate areas and calculate areas where a player may be occluded. For this reason, as shown in Figure 7(d), for example, the image may be divided into 12 rectangular areas in advance, as indicated by the dotted rectangles. Then, of the 12 rectangular areas, a rectangular area that includes an occlusion candidate area or a preliminary area for an occlusion candidate area may be set as the object detection area.
[0051] When the movable area of the occluding object is set as the occluding candidate area, the occluding candidate area may be set as the object detection area as is. In this case, only objects detected within the movable area of the occluding object may be treated as occluding objects.
[0052] In S603, the object detection unit 402 detects an object from the object detection area set in S602.
[0053] Fig. 8(a) is a diagram showing the processing results of this step. Fig. 8(a) shows the detection results of objects detected from the object detection area 703 of Fig. 7(c) in the input image of Fig. 5(b). Bounding boxes 801-807 indicate areas that include the detected objects (audience members 506-512). In this way, detected objects may be represented by bounding boxes.
[0054] The object detection unit 402 detects objects using, for example, color or shape. Objects may also be detected using the results of object detection from the object detection area. Examples of object detection methods include the background subtraction method described above.
[0055] When object detection is performed, the bounding box obtained as a result of object detection performed on the object detection area in the past input image may be saved as a history. In this case, if the size or shape of the bounding box changes, the area where the change occurs may be used to detect the object.
[0056] Alternatively, the object detection unit 402 may accumulate a history for a predetermined period in advance and use the accumulated history to detect an object. The history may be the input image or the results of predetermined image processing such as reduction, object detection, or statistical processing applied to the input image. The presence or absence of an object can be estimated by observing the difference between the accumulated history images or the results of image processing on past input images and the input image or the results of image processing on the input image for a certain period. For example, a histogram representing the frequency of pixel values in the object detection region may be calculated for each rectangular region in FIG. 7(d). Then, if the difference between the most frequent value up to the immediately preceding frame and the most frequent value of the input image exceeds a certain value, the corresponding region may be detected as an object region.
[0057] In S604, the object detection unit 402 determines whether filtering parameters are available.
[0058] If the filtering parameters are available (YES in S604), the process proceeds to S605. In S605, the object detection unit 402 determines invalid objects from among the objects detected in the object detection area in S603 based on the filtering parameters. The invalid objects are then deleted. The process then proceeds to S606.
[0059] The filtering parameters include, for example, a threshold value for the size of an object to be detected from an occluding object detection target area, and parameters related to the color and shape of an object to be determined as valid. When the filtering parameters are included in the object detection information, it is advisable to associate the position and shape information of the occluded candidate area with an object detection method according to the subject and store them in the occluded area detection information.
[0060] In S606, the object detection unit 402 determines, as an occluding object, an object that has not been invalidated from among the objects detected in the object detection area. The data format of the occluding object may be a mask image showing the shape of the detected occluding object, or data showing the position and size information of a bounding box including the detected occluding object.
[0061] If the filtering parameters are not available (NO in S604), S605 is skipped and the process proceeds to S606. That is, the object detection unit 402 determines all objects detected from the object detection area in S603 as occluding objects.
[0062] In S607, the occluded area determination unit 403 determines an area of the object detection area that includes an occluding object as an occluded area, and generates occluded area information indicating the position and size of the occluded area.
[0063] Assume that all of the detected objects shown by bounding boxes 801 to 807 in FIG. 8(a) are determined to be occluding objects. In this case, the bounding box areas may be determined as occluded areas 811 to 817 as they are, as shown in FIG. 8(b). When occluded areas are shown as bounding boxes as in FIG. 8(b), the amount of data for occluded area information indicating occluded areas can be reduced. Therefore, if it is desired to reduce the communication load imposed on output of occluded area information due to constraints on the implementation of a communication interface, etc., it is sufficient to determine occluded areas as shown in FIG. 8(b).
[0064] 8(c), the input image may be divided into predetermined rectangular regions in advance, and a set of rectangular regions including a region where an occluding object is detected may be determined as the occluding region 820. The shape of the occluding region may also be determined based on the shape obtained by applying a dilation process to the shape of the occluding object.
[0065] The occluded area information may be generated as a mask image similar to the foreground shape mask. In this case, the occluded area is represented as a foreground area, and the area other than the occluded area is divided into non-foreground areas, thereby generating a mask image representing the occluded area in the captured image.
[0066] In S608, the masked area determination unit 403 stores the masked area information.
[0067] On the other hand, when the foreground extraction unit 302 receives the captured image 513 of Figure 5(b) as an input image, it extracts the foreground based on the difference with the background image (captured image 503) that does not include the object of Figure 5(a), and generates a foreground shape mask.
[0068] Fig. 9(a) shows a foreground shape mask generated by extracting the foreground from the captured image of Fig. 5(b). In Fig. 9(a), the white area indicates the foreground area, and the black area indicates the non-foreground area, which is the area other than the foreground.
[0069] The three-dimensional model generation unit 305 generates an integrated mask by integrating the occluded area indicated by the occluded area information saved in S608 of the flowchart in FIG. 6 with the foreground shape mask generated by the foreground extraction unit 302.
[0070] 9(b) and (c) are diagrams showing examples of integrated masks, in which foreground areas are shown in white and other non-foreground areas are shown in black.
[0071] When the foreground shape mask of FIG. 9(a) is generated from the captured image of FIG. 5(b) and the occlusion areas 811 to 817 shown in FIG. 8(b) are determined from the captured image of FIG. 5(b), the three-dimensional model generation unit 305 generates the integrated mask shown in FIG. 9(b).
[0072] Furthermore, when the occluded area 820 shown in Fig. 8(c) is determined from the captured image of Fig. 5(b), the integrated mask shown in Fig. 9(c) is generated by the three-dimensional model generation unit 305. Note that when one of the areas divided in advance as shown in Fig. 8(c) is determined as the occluded area, the load of the process of generating the integrated mask using the occluded area can be reduced.
[0073] Then, the three-dimensional model generating unit 305 generates a three-dimensional model of the foreground object by the volume intersection method using the integrated mask instead of the foreground shape masks corresponding to the multiple image capturing devices.
[0074] Fig. 10 is a diagram illustrating the basic principle of the volume intersection method. Fig. 10(a) is a diagram of an image of a target object C, which is a foreground object, captured by an imaging device. A mask image including a two-dimensional silhouette (foreground region) of the target object C is obtained by binarizing the captured image based on the difference in color or brightness between the captured image of the target object C and the background image.
[0075] FIG. 10(b) shows a cone extending in three-dimensional space from the projection center (Pa) of the imaging device through each point on the contour of the two-dimensional silhouette Da. This cone is called the visual volume Va of the imaging device. FIG. 10(c) shows how a three-dimensional model of a foreground object is obtained from multiple visual volumes. As shown in FIG. 10(c), multiple visual volumes are obtained for each imaging device from a two-dimensional silhouette Da based on images captured synchronously by multiple different imaging devices at different positions. In generating a three-dimensional model using the visual volume intersection method, a three-dimensional model of the target object is generated by determining the intersection (common area) of the visual volumes corresponding to the multiple imaging devices.
[0076] A three-dimensional model is represented by a collection of voxels. Specifically, the target space for generation is tiled with voxels, which are tiny rectangular parallelepipeds. When each voxel in the target space for generation is back-projected onto the plane of each image capture device 101, voxels that are back-projected within the foreground region of the mask images of all image capture devices 101 are left as the foreground, and the rest are deleted. In this way, by removing voxels that do not fit within the foreground region of the mask image, a three-dimensional model of the foreground object is generated using voxels.
[0077] In this embodiment, by using an integrated mask in which the occluded area is treated as the foreground area and merged with the foreground shape mask, the area in the mask image where the foreground object is occluded can be treated as the foreground area. This allows a large amount of foreground area to be preserved. Therefore, even if a player for which a 3D model is to be generated is occluded, unnecessary deletion of voxels that should constitute the player can be prevented.
[0078] The occluded area information may be used to extract a foreground from an input image having a time code that matches the occluded area information. For example, the foreground shape mask in FIG. 9( a) is generated by extracting the foreground from the entire input image. Alternatively, the foreground extraction unit 302 may generate a foreground shape mask by extracting the foreground from an area of the input image other than the occluded area. In this case, by using the occluded area as the foreground area, an image identical to the integrated mask is generated, thereby reducing the processing load for generating the integrated mask. Furthermore, the area from which the foreground is extracted by the foreground extraction unit 302 is reduced, thereby reducing the processing load for foreground extraction by the foreground extraction unit 302. Furthermore, foreground objects from which a three-dimensional model is to be generated can be preferentially extracted.
[0079] The masked area information may also be used by the background generation unit 303 to generate a background image whose time code matches the masked area information. For example, an image corresponding to an area specified as the masked area information in the captured image may be used to update the background texture.
[0080] When spectators or staff members positioned close to a foreground object (player) are captured by multiple imaging devices and obscure the foreground object (player), it is often impossible to obtain the positions and shapes of the spectators or staff members in advance. In such cases, in order to generate a highly accurate three-dimensional model, it is necessary to increase the number of imaging devices or install the imaging devices in positions where the player is not obscured by the obstacle. According to this embodiment, even if a non-stationary obstacle is present between the imaging device and the foreground object (player) for which a three-dimensional model is to be generated, it is possible to prevent the obstacle from causing defects in the three-dimensional model of the foreground object (player).
[0081] It is also possible to generate a mask image by treating the entire obstacle's movable area in the captured image as a occluded area (foreground area) and then generate a three-dimensional model. However, in this case, the area from which the foreground area and non-foreground area effective for generating a three-dimensional model of the foreground can be obtained becomes small. On the other hand, in this embodiment, a portion of the object detection area, which is the obstacle's movable area, is treated as a occluded area, so that the area from which the foreground area and non-foreground area effective for generating a three-dimensional model can be obtained can be prevented from becoming small.
[0082] <Other embodiments> In the above-described embodiment, the virtual viewpoint image generation device 200 has been described as generating a three-dimensional model and a virtual viewpoint image, but the functions included in the virtual viewpoint image generation device 200 may be realized by one or more devices different from the virtual viewpoint image generation device 200. For example, the processes of extracting the foreground, image processing for generating occluded area information, generating the three-dimensional model, and generating the virtual viewpoint image may each be performed by a different device.
[0083] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]
[0084] 200 Virtual viewpoint image generation device 401 Area setting section 402 Object Detection Unit 403 Shielding area determination unit
Claims
1. An acquisition means for acquiring a captured image; an identification means for identifying a first region in the captured image in which other objects that may occlude the object for which three-dimensional shape data is to be generated and that are not the object for which three-dimensional shape data is to be generated may exist; a detection means for detecting a region of the object to be generated from the captured image and detecting a region of the other object from the first region identified by the identification means; a synthesis means for synthesizing a region of the object to be generated with a region of the other object to generate a mask image; a generation means for generating three-dimensional shape data of the object to be generated using the mask image; and The region of the other object is a rectangular region that includes the other object.
1. An image processing device comprising:
2. The identification means Identifying the first area based on the movable area of the other object 2. The image processing device according to claim 1, wherein:
3. The specifying means specifies the first area based on an area of the seating area.
2. The image processing device according to claim 1, wherein:
4. The other object is detected from the object detected in the first region.
4. The image processing device according to claim 1, wherein the image processing device is a computer.
5. By performing object detection from the first region, an object is detected from the first region.
5. The image processing device according to claim 4.
6. An object is detected from the first area based on information on the history of the captured images.
5. The image processing device according to claim 4.
7. An acquisition step of acquiring a captured image; a specifying step of specifying a first area in the captured image in which another object that may occlude the object for which three-dimensional shape data is to be generated, but is not a target for which three-dimensional shape data is to be generated, may exist; a detection step of detecting a region of the object to be generated from the captured image and detecting a region of the other object from the first region identified in the identification step; a combining step of combining a region of the object to be generated with a region of the other object to generate a mask image; a generation step of generating three-dimensional shape data of the object to be generated using the mask image. The region of the other object is a rectangular region that includes the other object. An image processing method comprising:
8. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 6.
Citation Information
Patent Citations
Generation device, generation method and program of three-dimensional model
JP2019106145A
Background model generation device, background model generation method, and background model generation program
JP2020112928A
Image processing device and program
JP2020135525A
Image processing device, image processing method, and program
JP2021056960A
Three-dimensional model generation method and device
JP2021099671A