Information processing apparatus, information processing method, and program

The information processing device automates the generation of 3D models by using mask images to exclude irrelevant silhouettes, addressing the inefficiency of manual deletion in conventional methods and enhancing processing efficiency.

JP2025175763APending Publication Date: 2025-12-03CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024082010
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Conventional methods for generating 3D models of dynamic objects, such as people, require manual deletion of unnecessary models post-estimation, which is time-consuming and inefficient.

Method used

An information processing device that acquires first and second mask images from captured images to generate a 3D model of a target object by identifying and excluding irrelevant silhouettes using a volume intersection method.

Benefits of technology

Facilitates easy and efficient generation of a 3D model of the target object by automating the exclusion of irrelevant areas, reducing data load and processing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025175763000001_ABST
    Figure 2025175763000001_ABST
Patent Text Reader

Abstract

To easily obtain a 3D model of a target object among objects in a photographing space.SOLUTION: An information processing apparatus acquires a first mask image showing silhouettes of all objects for a plurality of photographed images, and a second mask image based on the first mask image and showing silhouettes of objects other than an attention object among all objects, and on the basis of the acquired first mask image and second mask image, generates a 3D model representing a three-dimensional shape of the attention object.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to techniques for generating 3D models of objects. [Background technology]

[0002] In recent years, virtual viewpoint imaging technology has been gaining attention. This technology converts the entire space into 3D data by capturing the same scene from multiple viewpoints using multiple cameras, enabling viewing of the video from any viewpoint within the space. To convert the captured space into 3D data, shape estimation is performed to generate data (commonly referred to as a "3D model") representing the three-dimensional shapes of people and objects within the captured space. For example, when generating virtual viewpoint images for a theatrical performance, distracting people or objects, such as staff members other than performers, may appear within the captured space, increasing the data writing load during shape estimation and increasing the data size of the 3D model. To address this issue, if the distracting object is a stationary object such as a structure, a possible solution is to pre-generate a mask image for each camera that indicates the area representing the structure, and not generate a 3D model for the area of ​​the distracting object. Patent Document 1 discloses a technology that masks areas of structures or other objects that may obstruct the target object for shape estimation within each camera's image, and then estimates the shape of the target object by taking the masked areas into account. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2019-106145 Summary of the Invention [Problem to be solved by the invention]

[0004] As mentioned above, even with conventional technology, for static objects such as structures, shape estimation was performed taking their presence into account, which made it possible to reduce the load and data volume required when writing data when generating a 3D model of the target object. However, for dynamic objects such as obstructive people, the only option was to manually select and delete unnecessary 3D models after shape estimation, which was time-consuming and burdensome. [Means for solving the problem]

[0005] The information processing device according to the present disclosure is characterized by having an acquisition means for acquiring, from a plurality of captured images, a first mask image representing the silhouettes of all objects appearing in the captured images, and a second mask image based on the first mask image, the second mask image representing the silhouettes of all the objects other than a target object, and a generation means for generating a 3D model representing the three-dimensional shape of the target object based on the acquired first mask image and second mask image. [Effects of the Invention]

[0006] According to the present disclosure, a 3D model of a target object among objects in a shooting space can be easily obtained. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of an image processing system. [Figure 2] FIG. 10 is a diagram showing an example of camera installation. [Figure 3] FIG. 1 is a diagram showing a hardware configuration of an information processing apparatus. [Figure 4] 1 is a block diagram showing the functional configuration of a virtual viewpoint image generating device according to a first embodiment. [Figure 5] (a) to (c) are diagrams showing the basic principles of the volume intersection method. [Figure 6] 5 is a flowchart showing the flow of operations in the virtual viewpoint image generating device according to the first embodiment. [Figure 7A] 10A to 10C are diagrams illustrating a process of generating a non-focused OBJ mask image. [Figure 7B] 10A to 10C are diagrams illustrating a process of generating a non-focused OBJ mask image. [Figure 7C] 10A to 10C are diagrams illustrating a process of generating a non-focused OBJ mask image. [Figure 7D] 10A to 10C are diagrams illustrating a process of generating a non-focused OBJ mask image. [Figure 8] (a) shows the results of shape estimation using a foreground mask image, and (b) shows the state in which shape parts that depend on irrelevant areas shown in the non-focus OBJ mask image have been deleted from the shape estimation results in (a). [Figure 9] FIG. 10 is a block diagram showing the functional configuration of a virtual viewpoint image generating device according to a second embodiment. [Figure 10] 10 is a flowchart showing the flow of operations in the virtual viewpoint image generating device according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] The present invention will be described in detail below based on preferred embodiments thereof with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the illustrated configurations.

[0009] [Embodiment 1] FIG. 1 is a block diagram showing an example of the configuration of an image processing system according to this embodiment. The image processing system 100 includes a camera array 110 including multiple image capture devices (cameras), a control device 120, a foreground separation device 130, and a virtual viewpoint image generation device 140. Each of the multiple cameras constituting the camera array 110 synchronously captures an image of a common capture area, generating a captured image in which each pixel has, for example, an 8-bit pixel value for each RGB channel. The capture area may be, for example, a stadium where a sport such as soccer or karate is played, or a stage where a concert or play is played. FIG. 2 is a diagram showing how images are captured by multiple cameras installed to surround the capture area, with two people present in the capture area. Note that the multiple cameras do not need to be installed around the entire perimeter of the capture area; depending on installation space limitations, they may be installed only in a partial direction of the capture area. The number of cameras is not limited to the example shown in FIG. 2. For example, if the capture area is a vast space such as a soccer field, several tens to several hundred cameras may be installed around the stadium. Cameras with different functions, such as telephoto and wide-angle cameras, may also be installed. 2, the cameras are connected to the control device 120 and the foreground separation device 130 in a star topology, but a ring topology or bus topology using a daisy chain connection may also be used. Data on images captured from multiple viewpoints by each camera in the camera array 110 is sent to the control device 120 and the foreground separation device 130.

[0010] The control device 120 generates and manages camera parameters for each camera constituting the camera array 110 and supplies them to the virtual viewpoint image generation device 140. The camera parameters consist of external parameters representing the position and orientation (line of sight direction) of each camera and internal parameters representing the focal length and angle of view (captured area) of the lens of each camera, and are obtained through calibration. Calibration is a process for determining the correspondence between points in a three-dimensional world coordinate system acquired using multiple images of a specific pattern, such as a checkerboard, and corresponding two-dimensional points. The control device 120 also generates a camera path used to generate the virtual viewpoint image and supplies it to the virtual viewpoint image generation device 140. The camera path is information indicating the movement path of the virtual camera (virtual viewpoint) set continuously in time series, and is composed of information such as the position, orientation, and focal length of the virtual camera for each time (frame) specified by a time code. The operator sets this camera path by operating a joystick or the like while viewing a virtual three-dimensional space, in which the captured area is rendered as CG, on a UI screen. The camera path may be created simultaneously with the capture or may be created in advance. Furthermore, the control device 120 generates information for identifying, from among objects appearing in the captured image that may be used to generate a 3D model, an object of interest that will actually be used to generate a 3D model, distinguishing it from other objects. Hereinafter, this information will be referred to as "interested object information." Hereinafter, an object that may be used to generate a 3D model refers to a person or object that will be treated as a foreground region in the foreground-background separation process described below. A specific example of interest object information is an image captured at close range of a person or object that will be used as the interest object. Furthermore, information that represents the external characteristics of the object of interest, such as the uniform color and number of a sports player, or the costume color and accessories of an actor, can be used as interest object information.

[0011] The foreground separation device 130 performs foreground / background separation processing for each of the input multiple captured images for each frame. Specifically, the device distinguishes between foreground areas corresponding to people, etc. in the captured images and other background areas, and extracts the foreground areas from the captured images. Any method can be used to distinguish the foreground areas, such as background subtraction or inference using a convolutional neural network (CNN). In the case of background subtraction, the device calculates the difference between each captured image and a pre-prepared background image (an image captured without people, etc. in the foreground), and identifies the area corresponding to the difference as the foreground area. The foreground separation device 130 extracts only the identified foreground area from the captured images and generates a foreground image in which the foreground area is represented using three RGB channels, and a foreground mask image (also called a "silhouette image") in which the foreground area is represented using one black and white channel. 2, if two people appear in the captured image, the area in the captured image where the two people appear is the foreground area, and therefore the foreground mask image in this case represents the silhouettes of the two people. The generated foreground image and foreground mask image are output to virtual viewpoint image generation device 140.

[0012] The virtual viewpoint image generating device 140 generates a virtual viewpoint image based on the target OBJ information, a foreground image, a foreground mask image, camera parameters, and a camera path. Specifically, first, using the target OBJ information and the foreground mask image, shape data (hereinafter referred to as a "3D model") representing the three-dimensional shape of the target object indicated by the target OBJ information among multiple objects present in the shooting area is generated. Then, the positional relationship between each foreground image and the 3D model is determined from the camera parameters defined in the camera path, and the foreground image corresponding to the 3D model is mapped to generate a virtual viewpoint image when the target object is viewed from an arbitrary angle.

[0013] 2 is an example and is not limiting. For example, one computer may have the functions of multiple devices (e.g., foreground separation device 130 and virtual viewpoint image generation device 140). Alternatively, each camera module may have the function of foreground separation device 130, and each camera may be configured to supply data of a captured image and a foreground image and a foreground mask image corresponding to the captured image.

[0014] The control device 120, the foreground separation device 130, and the virtual viewpoint image generation device 140 are realized by a general computer (information processing device) equipped with a CPU for performing arithmetic processing, memory for storing arithmetic processing results, programs, etc. FIG. 3 is a diagram showing the hardware configuration of a general information processing device. The information processing device includes a CPU 311, a ROM 312, a RAM 313, an auxiliary storage device 314, a display unit 315, an operation unit 316, a communication I / F 317, and a bus 318. The CPU 311 is a processing unit that controls the entire information processing device and realizes various functions of the information processing device using computer programs and data stored in the ROM 312 and the RAM 313. Note that the information processing device may have one or more dedicated hardware components different from the CPU 311, and at least a portion of the processing by the CPU 311 may be executed by the dedicated hardware components. Examples of the dedicated hardware include an ASIC (application-specific integrated circuit), an FPGA (field-programmable gate array), and a DSP (digital signal processor). The ROM 312 stores programs that do not require modification. The RAM 313 temporarily stores programs and data supplied from the auxiliary storage device 314, and data supplied from the outside via the communication I / F 317. The auxiliary storage device 314 is configured, for example, with a hard disk drive or the like, and stores various data such as image data and audio data.

[0015] The display unit 315 is composed of, for example, a liquid crystal display, an LED, etc., and displays a GUI (Graphical User Interface) for the user to operate the information processing device. The operation unit 316 is composed of, for example, a keyboard, a mouse, a joystick, a touch panel, etc., and receives operations by the user and inputs various instructions to the CPU 311. The CPU 311 operates as a display control unit that controls the display unit 315 and an operation control unit that controls the operation unit 316. The communication I / F 317 is used for communication with devices external to the information processing device. For example, when the information processing device is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 317. When the information processing device has a function for wireless communication with an external device, the communication I / F 317 is equipped with an antenna. The bus 318 connects each unit of the information processing device to transmit information.

[0016] In this embodiment, the display unit 315 and the operation unit 316 are assumed to exist inside the information processing device, but at least one of the display unit 315 and the operation unit 316 may exist as a separate device outside the information processing device.

[0017] <Details of the virtual viewpoint image generation device> Next, we will explain the functions of the virtual viewpoint image generation device 140. Fig. 4 is a block diagram showing the functional configuration of the virtual viewpoint image generation device 140 according to this embodiment. The virtual viewpoint image generation device 140 of this embodiment has a data acquisition unit 401, a mask processing unit 402, a shape estimation unit 403, a visibility determination unit 404, and a rendering unit 405.

[0018] The data acquisition unit 401 receives, from the control device 120, the camera parameters of each camera constituting the camera array 110, a camera path for generating a virtual viewpoint image, and target OBJ information for identifying an object of interest in a captured image. The data acquisition unit 401 also receives, from the foreground separation device 130, data on captured images (multiple viewpoint images) obtained by synchronously capturing images by each camera of the camera array 110, and data on foreground images obtained by extracting foreground portions present in each captured image.

[0019] Based on the input foreground mask image and information about the object of interest, the mask processing unit 402 generates a mask image in which the silhouette area indicated by the foreground mask image, corresponding to objects other than the object of interest, is expressed in one black and white channel. Hereinafter, this mask image will be referred to as the "non-object of interest mask image."

[0020] The shape estimation unit 403 performs shape estimation based on the camera parameters of all cameras, the foreground mask image, and the non-focus OBJ mask image to generate a 3D model of the focused object. Shape estimation uses a volume intersection method or a method based on the volume intersection method. Figures 5(a) to 5(c) illustrate the basic principles of the volume intersection method. From an image of an object captured, a mask image is obtained that represents a two-dimensional silhouette of the object on the captured image (Figure 5(a)). A cone extending in three-dimensional space from the camera's projection center through each point on the contour of the mask image is then considered (Figure 5(b)). This cone is called the "view volume" of the object captured by the corresponding camera. Furthermore, the three-dimensional shape of the object is determined by calculating the common area of ​​multiple view volumes, i.e., the intersection of the view volumes (Figure 5(c)). During this shape estimation, the non-focus OBJ mask images are used to identify which parts of the silhouette area shown in the foreground mask image of each camera are irrelevant areas (i.e., silhouette areas of non-focus objects) that are not required for generating a 3D model of the focused object. Then, from the 3D shape estimated using the foreground mask image, shape portions that are determined to be foreground based solely on irrelevant regions indicated by the non-target OBJ mask images for each camera are deleted. This can be rephrased as a process in which, among the components (voxels or point clouds) of the 3D model obtained based on the foreground mask image, components that are recognized as part of the target object by any camera are left as they are, and the other components are deleted. Regarding the deletion of components of the 3D model, if a certain number (threshold) or more of cameras do not recognize the component as the target object, that component may be deleted. Alternatively, if a certain number (threshold) or more of cameras recognize the component as the shape of a non-target object, that component may be deleted.

[0021] The visibility determination unit 404 determines which portions of the three-dimensional shape of the target object are visible from each camera using the camera parameters and the non-target OBJ mask image for the 3D model of the target object obtained by the shape estimation unit 403. In this visibility determination, the visibility determination unit 404 treats portions corresponding to the silhouette region (irrelevant region) shown in the non-target OBJ mask image of the target camera as invisible to that camera. Treating the object shape corresponding to the silhouette region shown in the non-target OBJ mask image as invisible to that camera in this way is intended to enable accurate texture generation. For example, if an obstructing non-target object exists between the target object and the camera and occludes the target object, the occluded portion will not be captured in the image captured by that camera. However, the 3D model generated by the shape estimation unit 403 eliminates the influence of the non-target object and accurately represents the three-dimensional shape of the target object. Therefore, when visibility determination is performed on the generated 3D model based on the camera parameters of each camera, portions that would be occluded and invisible in the actual captured space will be determined to be visible. If texture generation is performed using the results of such visibility determination, image areas containing irrelevant objects (non-target objects) will be used for texture generation, resulting in the generation of an incorrect texture. Therefore, the above problem is solved by treating the areas corresponding to the irrelevant areas of each camera as invisible areas.

[0022] The rendering unit 405 generates a virtual viewpoint image for each frame based on the camera parameters, the camera path, the foreground image, the 3D model, and the visibility determination result. The generated virtual viewpoint image is output to the display unit 315 and displayed.

[0023] The functions of the virtual viewpoint image generating device 140 may be distributed among multiple information processing devices. For example, the function of the rendering unit 405 may be performed by another information processing device. Also, the 3D model resulting from the shape estimation performed by the shape estimation unit 403 and the texture information determined by the rendering unit 405 may be displayed on the display unit 315.

[0024] <Operation flow of the virtual viewpoint image generation device> Next, the flow of operations in the virtual viewpoint image generating device 140 according to this embodiment will be described with reference to the flowchart of FIG.

[0025] In S601, the data acquisition unit 401 acquires, from the control device 120, the camera parameters of each camera constituting the camera array 110 and target OBJ information indicating a target object for which the user wishes to generate a 3D model. This acquisition completes preparations for starting the process of generating a virtual viewpoint image. After synchronized shooting is performed by each camera, the captured images from each camera are sent to the foreground separation device 130, where foreground / background separation processing is performed, and a foreground image and a foreground mask image corresponding to the captured image from each camera are generated for each frame.

[0026] In the next step S602, the data acquisition unit 401 acquires a foreground image and a foreground mask image generated from the images captured by each camera from the foreground separation device 130. The data acquisition unit 401 also acquires a camera path for generating a virtual viewpoint image, which is referenced by the rendering unit 405, from the control device 120. Each step from S603 onwards is executed for each frame based on the time code attached to the foreground image.

[0027] In S603, the mask processing unit 402 generates the aforementioned non-attention OBJ mask image based on each foreground image, each foreground mask image, and attention OBJ information of the target frame. Specifically, first, an area corresponding to the attention object is detected in the foreground image using information such as color and shape that represents the appearance characteristics of the image or attention object as attention OBJ information. The detection method is not particularly limited, but two types of methods are described as examples. One is a method using segmentation. In the segmentation method, segmentation is performed on the foreground image to divide it into regions, and which of the divided regions corresponds to the attention object is identified based on the attention OBJ information. The other is a method using machine learning. In the machine learning method, the foreground image and attention OBJ information are input to a trained model obtained by prior training, and inference is performed to detect an area in the foreground image that corresponds to the attention object. Note that the method for detecting the attention object is not limited to these. Then, a mask image is generated that represents the area of ​​the attention object detected from the foreground image in one black and white channel from the silhouette area indicated by the foreground mask image. Hereinafter, this mask image will be referred to as the "obj mask image of interest." In the above case, if one of the two people is the object of interest for which a 3D model is to be generated and the other is an object not to be generated, the silhouette represented by the obj mask image of interest will be that of the person who is the object of interest. Next, the mask processing unit 402 compares the generated obj mask image of interest with the corresponding foreground mask image and performs processing to remove the silhouette area represented by the obj mask image of interest from the silhouette area represented by the foreground mask image. In this way, by removing the silhouette area represented by the obj mask image of interest from the silhouette area represented by the foreground mask image, a non-obj mask image representing the silhouette area of ​​the object for which a 3D model is not to be generated is obtained. Figures 7A to 7D are diagrams explaining the generation process of a non-obj mask image. In Figures 7A to 7D, images 700a to 700d show captured images obtained by four cameras with different viewpoints, and images 701a to 701d show foreground images corresponding to the captured images 700a to 700d.Furthermore, images 702a to 702d show foreground mask images corresponding to the photographed images 700a to 700d. Images 703a to 703d show target OBJ mask images showing the silhouette of the person 201, which is the target object, corresponding to the photographed images 700a to 700d. Furthermore, images 704a to 704d show non-target OBJ mask images showing the silhouette of the person 202, which is the non-target object, corresponding to the photographed images 700a to 700d. The non-target OBJ mask images obtained in this manner are output to the shape estimation unit 403.

[0028] In S604, the shape estimation unit 403 generates a 3D model of the target object based on the camera parameters of each camera, the foreground mask image of the target frame, and the non-target OBJ mask image. Specifically, shape estimation is first performed using the camera parameters of each camera and the foreground mask image. The shape estimation method used here is the volume intersection method (or a method based on the volume intersection method) described above in FIG. 5. This results in 3D models representing the three-dimensional shapes of all objects included in the input foreground image. FIG. 8(a) shows how a 3D model 801 of a person 201, which is the target object, and a 3D model 802 of a person 202, which is the non-target object, are obtained by shape estimation using the foreground mask images of each camera, including the foreground mask images 702a to 702d shown in FIGS. 7A to 7D. Next, from the 3D models of all objects thus obtained, components (e.g., voxels) that depend only on the silhouettes of the non-target objects are deleted. In this case, which of the silhouette regions shown in the foreground mask images of each camera will become irrelevant regions representing the silhouettes of non-target objects is identified using the non-target OBJ mask images. Fig. 8(b) shows the state after a deletion process has been performed on the two 3D models 801 and 802 shown in Fig. 8(a) to delete components (e.g., voxels) that depend only on the irrelevant regions shown in the non-target OBJ mask images 704a to 704d of Figs. 7A to 7D. In Fig. 8(b), only the 3D model 801 of the person 201, which is the target object, remains as it is (the 3D model 802' shown by the dashed line indicates that the 3D model 802 of the person 202, which is the non-target object, has been deleted by the deletion process).

[0029] In S605, the visibility determination unit 404 determines whether the 3D model of the target object obtained in S604 is visible from each camera (visibility determination) based on the camera parameters of each camera and the non-target OBJ mask image of the target frame. The visibility determination result is provided to the rendering unit 405 as visibility information. In this visibility determination, among the parts of the 3D model visible when viewed from the viewpoint of the target camera, parts corresponding to irrelevant areas indicated by the non-target OBJ mask image are determined to be invisible from the target camera (invisible area). The reason for treating such parts as invisible areas is as described above. That is, in the subsequent rendering process, it is to prevent the color values ​​of pixel areas in the foreground image that should be occluded by the non-target object from being erroneously used as the texture of the 3D model. The visibility determination result for each camera obtained in this way is output to the rendering unit 405 as visibility information.

[0030] In S606, the rendering unit 405 determines the texture when the 3D model is viewed from the virtual viewpoint specified by the camera path based on the visibility information, the camera parameters of each camera, and each foreground image, and generates a virtual viewpoint image in the target frame.

[0031] In S607, it is determined whether processing has been completed for all frames of the input foreground image and foreground mask image, and if there are any unprocessed frames, the process returns to S603 and continues. On the other hand, if processing has been completed for all frames, this flow ends. The generated virtual viewpoint images for all frames are output to the display unit 315. The above is the flow of operations in the virtual viewpoint image generation device 140 according to this embodiment.

[0032] <Variation 1> While the above description takes as an example a case where only one of two people appearing in a captured image is the object of interest, it is also possible to change the object of interest at a certain time (frame). In this case, the user inputs, via the control device 120, a time code of interest OBJ information corresponding to the object of interest to be newly generated as a target for 3D model generation. When it is determined in S607 that there is an unprocessed frame, the virtual viewpoint image generation device 140 checks whether new object of interest OBJ information has been input. If new object of interest OBJ information is detected, the process from S603 onward may be executed with the frame identified by the time code indicated by the object of interest information as the new object of interest. For example, in the case of the specific example described above, suppose the object of interest is switched from person 201 to person 202 midway through the process. If captured images 700a to 700d shown in FIGS. 7A to 7D are obtained after the switch, images 703a to 703d will show non-object of interest mask images representing the silhouette of person 201, the non-object of interest corresponding to captured images 700a to 700d. The images 704a to 704d show target OBJ mask images that represent the silhouette of the person 202, which is the target object corresponding to the photographed images 700a to 700d.

[0033] <Variation 2> There are cases where it is desired to generate 3D models for specific individuals among a large number of people appearing in a captured image. In this case, attention object information for each of the individuals may be generated by the control device 120 and input to the virtual viewpoint image generation device 140. Alternatively, one attention object information may include information for identifying each of the multiple objects. Then, based on this attention object information, the regions of the multiple attention objects are detected from the foreground image, and an attention object mask image representing the silhouettes of the multiple attention objects is generated based on the detection results. In other words, when there are multiple attention objects, the region obtained by integrating the silhouettes of the multiple attention objects becomes the silhouette region indicated by the attention object mask image. Furthermore, by comparing the generated attention object mask image with a foreground mask image corresponding to a captured image showing a large number of people, a non-attention object mask image is generated by excluding the silhouette region indicated by the attention object mask image from the silhouette region indicated by the foreground mask image. Then, the processing from step S604 onward may be performed based on the generated attention object mask image and non-attention object mask image.

[0034] As described above, according to this embodiment, even if there is an obstructing object in the shooting space that is treated as the foreground and that appears together with the target object in the captured image, it is possible to generate only the target 3D model.

[0035] [Embodiment 2] Next, when a captured image contains multiple objects for which a 3D model can be generated, a mode in which foreground / background separation processing is performed for each object to generate a 3D model of one or multiple objects of interest will be described as embodiment 2. Note that a description of the content common to embodiment 1 will be omitted, and the following will describe the differences.

[0036] <Details of the virtual viewpoint image generation device> FIG. 9 is a block diagram showing the functional configuration of a virtual viewpoint image generation device 140 according to this embodiment. The virtual viewpoint image generation device 140 of this embodiment includes a data acquisition unit 901, a foreground separation unit 902, a mask processing unit 903, a shape estimation unit 904, a visibility determination unit 905, and a rendering unit 906. The most significant difference from the first embodiment is the addition of a foreground separation unit 902. This foreground separation unit 902 has the same function as the foreground separation device 130 of the first embodiment, and performs foreground / background separation processing for each frame of a plurality of input captured images to generate a foreground image and a foreground mask image. Furthermore, the foreground separation unit 902 also generates a foreground mask image for each object (hereinafter referred to as an "OBJ-specific foreground mask image") based on information for distinguishing each object that can be the target of 3D model generation from other objects (hereinafter referred to as "OBJ information"). Here, the OBJ information is an image taken at close range of only people or objects that can be used to generate a 3D model, and in the case of a person, it is information that indicates the characteristics of the object, such as the color of the uniform or clothing, the number on the back, and accessories. The OBJ information is linked to ID information (such as a number) that is assigned to each object. As with the above-mentioned focus OBJ information, OBJ information may be generated for each object, or one piece of OBJ information may include information on the appearance characteristics and ID information of each of multiple objects.

[0037] The functions of each unit other than the foreground separation unit 902 (data acquisition unit 901, mask processing unit 903, shape estimation unit 904, visibility determination unit 905, and rendering unit 906) are basically the same as those in embodiment 1. Minor differences will be mentioned in the description of the operation flow below.

[0038] <Operation flow of the virtual viewpoint image generation device> Next, the flow of operations in the virtual viewpoint image generating device 140 according to this embodiment will be described with reference to the flowchart of FIG.

[0039] In S1001, the data acquisition unit 901 acquires the camera parameters of each camera constituting the camera array 110, as well as the above-mentioned OBJ information and target OBJ selection information from the control device 120. Here, the "target OBJ selection information" refers to the ID information of an object that the user has actually selected as a target for generating a 3D model from among all objects related to the acquired OBJ information. For example, the user selects an object for which they want to generate a 3D model from among multiple objects appearing in a captured image via a GUI (not shown) of the control device 120, and inputs the target OBJ selection information to the virtual viewpoint image generation device 140. By acquiring these various data, preparations are completed for starting the process of generating a virtual viewpoint image.

[0040] In S1002, the data acquisition unit 901 acquires captured images obtained by synchronous shooting from each camera. The data acquisition unit 901 also acquires, from the control device 120, a camera path for generating a virtual viewpoint image, which is referenced by the rendering unit 906. Each step from S1003 onwards is executed for each frame based on the time code attached to the captured image.

[0041] In S1003, the foreground separation unit 902 performs foreground / background separation processing on each captured image of the target frame to generate a foreground image and a foreground mask image. Furthermore, the aforementioned mask image for each OBJ is generated based on the generated foreground image and the OBJ information. When the object information indicates multiple objects and these multiple objects exist in a single foreground image, a mask image for each OBJ is generated for each object. The mask image for each OBJ can be generated using a method similar to the mask image for the target OBJ described in S603 of the first embodiment. Specifically, first, an area corresponding to the target object indicated by the OBJ information is detected from a single foreground image using information such as color and shape that indicate the image and characteristics of the target object. The detection method may be segmentation or inference using a trained model. Then, based on the detection results, a mask image for each OBJ is generated that shows the silhouette of the target object included in the foreground image. This processing is performed for each object in the foreground image, and as many mask images for each OBJ as there are objects indicated by the object information are generated. The generated foreground image is output to the rendering unit 906, the foreground mask image is output to the mask processing unit 903 and the shape estimation unit 904, and the OBJ-specific foreground mask image is output to the mask processing unit 903.

[0042] In S1004, the mask processing unit 903 generates a non-focus OBJ mask image based on each foreground mask image, the per-OBJ mask image, and the focus OBJ selection information of the target frame. Specifically, the mask processing unit 903 compares the per-OBJ mask image of the focus object related to the focus OBJ selection information with the foreground mask image, and identifies an area obtained by excluding the silhouette area of ​​the per-OBJ mask image from the silhouette area of ​​the foreground mask image. Then, a mask image is generated in which the identified area is treated as an irrelevant area (silhouette area) that is unnecessary for generating a 3D model of the focus object, and the generated mask image is used as the non-focus OBJ mask image. At this time, if there are multiple focus objects related to the focus OBJ selection information, a non-focus OBJ mask image is generated in which the area obtained by excluding the silhouette areas of the per-OBJ mask images of all the focus objects from the silhouette area of ​​the foreground mask image is used as the irrelevant area. The generated non-focus OBJ mask is output to the shape estimation unit 904.

[0043] The processes in S1005 to S1008 correspond to S604 to S607 in the flow of Fig. 6 in the first embodiment, and there is no particular difference between them, so a description thereof will be omitted. The above is the flow of operations in the virtual viewpoint image generation device 140 according to this embodiment.

[0044] The concept of <Modification 1> of the above-described first embodiment also applies to this embodiment. That is, the user may input, from the control device 120, target object selection information corresponding to a target object of interest for which a 3D model is to be generated, by specifying a time code. In this case, when it is determined in S1008 that there is an unprocessed frame, the virtual viewpoint image generation device 140 checks whether new target object selection information has been input. Then, when new target object selection information is detected, the process from S1003 onward may be executed, using the frame identified by the time code indicated by the target object selection information as the new target object.

[0045] As described above, according to this embodiment, when an obstructive object treated as the foreground exists in the shooting space and appears together with the target object in the shot image, it is possible to generate only the target 3D model.

[0046] (Other Examples) The present disclosure can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0047] The present disclosure also includes the following configurations and methods.

[0048] [Configuration 1] An information processing device characterized by comprising: an acquisition means for acquiring, from a plurality of captured images, a first mask image representing the silhouettes of all objects appearing in the captured images; and a second mask image based on the first mask image, the second mask image representing the silhouettes of all the objects other than a target object; and a generation means for generating a 3D model representing the three-dimensional shape of the target object based on the acquired first mask image and second mask image.

[0049] [Configuration 2] The information processing device according to configuration 1, characterized in that the second mask image is generated based on the first mask image and a third mask image representing a silhouette of the target object generated based on information about the target object.

[0050] [Configuration 3] the information about the target object is information for identifying the target object from among objects appearing in the plurality of captured images, and distinguishing the target object from other objects; In the acquisition means, the third mask image is generated based on information about the object of interest; The second mask image is generated and acquired by subtracting the silhouette area indicated by the third mask image from the silhouette area indicated by the first mask image. 3. The information processing device according to configuration 2.

[0051] [Configuration 4] the information for identifying the target object in distinction from other objects is information representing an external feature of the target object, the third mask image is generated using a foreground image obtained by extracting the foreground from the captured image and information representing the external characteristics of the target object. 4. The information processing device according to configuration 3.

[0052] [Configuration 5] 5. The information processing device according to configuration 4, wherein the information representing the external characteristics of the target object is an image of the target object photographed at a close distance, or information about the color or shape of the target object.

[0053] [Configuration 6] 6. The information processing device according to configuration 4 or 5, wherein the third mask image is generated by detecting an area corresponding to the object of interest from the foreground image based on information representing external features of the object of interest from an area obtained by segmentation of the foreground image.

[0054] [Configuration 7] 6. The information processing device according to configuration 4 or 5, wherein the third mask image is generated by detecting a region corresponding to the object of interest from the foreground image by inference using a trained model that receives as input information representing the foreground image and external features of the object of interest.

[0055] [Configuration 8] the image processing device further comprises a foreground separation means for performing a foreground / background separation process on the plurality of captured images to generate the first mask image, and for generating a fourth mask image representing a silhouette of each object based on the generated first mask image and information for identifying each of all objects constituting the foreground by distinguishing them from other objects, the information about the object of interest is ID information for identifying an object selected by a user from among all objects constituting the foreground; In the acquisition means, a third mask image is generated based on the ID information and the fourth mask image; The second mask image is generated and acquired by subtracting the silhouette area indicated by the third mask image from the silhouette area indicated by the first mask image. 3. The information processing device according to configuration 2.

[0056] [Configuration 9] 9. The information processing device according to configuration 8, wherein the information for identifying each of all objects constituting the foreground and distinguishing them from other objects is information representing the external characteristics of each of all of the objects.

[0057] [Configuration 10] 10. The information processing device according to configuration 8 or 9, wherein the fourth mask image is generated by detecting an area corresponding to the object of interest from the foreground image based on information representing external features of the object of interest from an area obtained by segmentation of the foreground image obtained by foreground / background separation processing on the captured image.

[0058] [Configuration 11] 10. The information processing device according to configuration 8 or 9, wherein the fourth mask image is generated by detecting an area corresponding to the object of interest from the foreground image by inference using a trained model that inputs a foreground image obtained by foreground / background separation processing on the captured image and information representing external features of the object of interest.

[0059] [Configuration 12] The information processing device according to any one of configurations 1 to 11, wherein the generating means generates a 3D model of the target object by deleting a shape portion that depends on a silhouette area indicated by the second mask image from a three-dimensional shape obtained by shape estimation using the first mask image.

[0060] [Configuration 13] The information processing device according to configuration 12, wherein the generating means generates a 3D model of the target object by deleting a shape portion that depends only on a silhouette area indicated by the second mask image from a three-dimensional shape obtained by shape estimation using the first mask image.

[0061] [Configuration 14] 14. The information processing device according to any one of configurations 1 to 13, further comprising a determination means for determining visibility of a 3D model generated by the generation means from a plurality of photographing devices corresponding to the plurality of photographed images using the second mask image.

[0062] [Configuration 15] 15. The information processing device according to configuration 14, further comprising a rendering means for determining a texture when the 3D model generated by the generating means is viewed from a virtual viewpoint based on visibility information that is the result of the determination by the determining means, and generating a virtual viewpoint image.

[0063] [Method 1] an acquisition step of acquiring a first mask image representing silhouettes of all objects constituting a foreground from a plurality of captured images, and a second mask image representing silhouettes of all the objects other than the target object; a generation step of generating a 3D model representing a three-dimensional shape of the target object based on the acquired first mask image and the acquired second mask image; An information processing method comprising:

[0064] [Configuration 16] A program for causing a computer to execute the information processing method described in Method 1.

Claims

1. an acquisition means for acquiring, from a plurality of photographed images, a first mask image representing silhouettes of all objects appearing in the photographed images, and a second mask image based on the first mask image, the second mask image representing silhouettes of all the objects except a target object; a generation means for generating a 3D model representing a three-dimensional shape of the target object based on the acquired first mask image and the acquired second mask image; An information processing device comprising:

2. 2 . The information processing device according to claim 1 , wherein the second mask image is generated based on the first mask image and a third mask image representing a silhouette of the target object that is generated based on information about the target object.

3. the information about the target object is information for identifying the target object from among objects appearing in the plurality of captured images, and distinguishing the target object from other objects; In the acquisition means, the third mask image is generated based on information about the target object; the second mask image is generated and acquired by subtracting the silhouette region indicated by the third mask image from the silhouette region indicated by the first mask image.

3. The information processing apparatus according to claim 2, wherein:

4. the information for identifying the target object in distinction from other objects is information representing an external feature of the target object, the third mask image is generated using a foreground image obtained by extracting a foreground from the captured image and information representing external features of the target object.

4. The information processing apparatus according to claim 3,

5. 5. The information processing apparatus according to claim 4, wherein the information representing the external characteristics of the target object is an image of the target object photographed at a close distance, or information about the color or shape of the target object.

6. 5. The information processing device according to claim 4, wherein the third mask image is generated by detecting an area corresponding to the object of interest from the foreground image based on information representing external features of the object of interest from an area obtained by segmentation of the foreground image.

7. 5. The information processing device according to claim 4, wherein the third mask image is generated by detecting a region corresponding to the object of interest from the foreground image by inference using a trained model that receives as input information representing the foreground image and external features of the object of interest.

8. the image processing device further comprises a foreground separation means for performing a foreground / background separation process on the plurality of photographed images to generate the first mask image, and for generating a fourth mask image representing a silhouette of each object based on the generated first mask image and information for identifying each of all objects constituting the foreground in a manner distinguishable from other objects, the information about the target object is ID information that identifies an object selected by a user from among all objects that make up the foreground; In the acquisition means, a third mask image is generated based on the ID information and the fourth mask image; the second mask image is generated and acquired by subtracting the silhouette region indicated by the third mask image from the silhouette region indicated by the first mask image.

3. The information processing apparatus according to claim 2, wherein:

9. 9. The information processing apparatus according to claim 8, wherein the information for identifying each of all the objects constituting the foreground in a manner that distinguishes them from other objects is information that represents the external characteristics of each of the objects.

10. 9. The information processing device according to claim 8, wherein the fourth mask image is generated by detecting an area corresponding to the target object from the foreground image based on information representing external features of the target object from an area obtained by segmentation of the foreground image obtained by foreground / background separation processing on the captured image.

11. 9. The information processing device according to claim 8, wherein the fourth mask image is generated by detecting a region corresponding to the object of interest from the foreground image by inference using a trained model that inputs a foreground image obtained by foreground / background separation processing on the captured image and information representing external features of the object of interest.

12. 2. The information processing device according to claim 1, wherein the generating means generates a 3D model of the target object by deleting a shape portion that depends on a silhouette area indicated by the second mask image from a three-dimensional shape obtained by shape estimation using the first mask image.

13. 13. The information processing device according to claim 12, wherein the generating means generates a 3D model of the target object by deleting a shape portion that depends only on a silhouette area indicated by the second mask image from the three-dimensional shape obtained by shape estimation using the first mask image.

14. 2. The information processing device according to claim 1, further comprising a determination unit that determines visibility of the 3D model generated by the generation unit from a plurality of image capturing devices corresponding to the plurality of captured images using the second mask image.

15. 15. The information processing apparatus according to claim 14, further comprising a rendering means for determining a texture when the 3D model generated by the generating means is viewed from a virtual viewpoint based on visibility information resulting from the determination by the determining means, and generating a virtual viewpoint image.

16. an acquisition step of acquiring a first mask image representing silhouettes of all objects constituting a foreground from a plurality of captured images, and a second mask image representing silhouettes of all the objects other than a target object; a generation step of generating a 3D model representing a three-dimensional shape of the target object based on the acquired first mask image and the acquired second mask image; An information processing method comprising:

17. A program for causing a computer to function as the information processing device according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Generation device, generation method and program of three-dimensional model

    JP2019106145A