Image processing system
By strategically selecting cameras with appropriate focal lengths and positions to generate foreground and background images for virtual viewpoint images, the image processing system reduces processing load while maintaining image quality, effectively addressing the challenge of increasing camera resolution and number.
Patent Information
- Application Number
- JP2024037956
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-02-12
AI Technical Summary
Increasing camera resolution and the number of cameras to improve virtual viewpoint image quality leads to a significant increase in processing load for generating material data such as foreground and three-dimensional shape data.
The image processing system selectively uses cameras with appropriate focal lengths and positions to generate foreground and background images, and uses these images to generate three-dimensional shape data, thereby reducing the processing load while maintaining image quality.
This approach allows for the appropriate generation of material data while minimizing the increase in processing load, thereby enhancing the efficiency of the image processing system.
Smart Images

Figure 0007693879000002 
Figure 0007693879000003 
Figure 0007693879000004
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing system Mu and the like.
Background Art
[0002] There has been proposed an apparatus that generates a virtual viewpoint image viewed from a virtual viewpoint specified by a user from images captured by a plurality of cameras constituting a photographing system. The image processing system disclosed in Patent Document 1 generates a foreground image, a background image, three-dimensional shape data, etc. from images captured by a plurality of cameras. The foreground image, the background image, and the three-dimensional shape data are material data for generating a virtual viewpoint image. The image processing system acquires the material data based on the virtual viewpoint specified by the user and reproduces the virtual viewpoint image.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In order to improve the image quality of the virtual viewpoint image, it is necessary to improve the image quality of the foreground image and the accuracy of the three-dimensional shape data. For this purpose, it is required to increase the camera resolution and the number of cameras. However, increasing the camera resolution and the number of cameras causes a problem that the processing load for generating material data such as the foreground image and the three-dimensional shape data increases. The present invention has been made in view of the above-described problems, and an object thereof is to appropriately generate material data while reducing an increase in processing load.
Means for Solving the Problems
[0005] The image processing system of the present invention is a foreground image used for generating a foreground region of a virtual viewpoint image corresponding to a virtual viewpoint based on a first image acquired by a first imaging means, and shows a foreground included in the first image. The system generates a foreground image and outputs the first generation of the foreground image. device Based on a second image acquired by a second imaging means having a focal length shorter than the focal length of the first imaging means, the virtual viewpoint image of generates a background image used for generating a background region, which is a background image showing the background included in the second image, and outputs the second generation of the background image. device And A third generation device that receives the foreground image output from the first generation device and the background image output from the second generation device, and generates three-dimensional shape data of the foreground based on the foreground image output by the first generation device; has, and the second generation device does not output a foreground image showing the foreground included in the second image. which is used for generating the three-dimensional shape data of the foreground Foreground image , to the third generation device Do not output configured as follows It is characterized by this.
Effect of the Invention
[0006] According to the present invention, it is possible to appropriately generate material data while reducing an increase in processing load.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0008] Hereinafter, embodiments according to the present invention will be described with reference to the drawings. (First Embodiment) The material generation device 10 according to the first embodiment will be described. The material generation device 10 generates a foreground image, a background image, 3D shape data (hereinafter referred to as a 3D model), etc., which are material data of a virtual viewpoint image, using camera images acquired from a plurality of cameras. Generally, in order to generate a 3D model, a more accurate 3D model can be generated by using camera images of many cameras. On the other hand, if all the camera images of many cameras are stored, the data to be stored will increase and the disk access bandwidth will become large. The material generation device 10 of the present embodiment reduces the data to be stored by selecting cameras that shoot the foreground image and the background image to be stored as material data from a plurality of cameras while generating a 3D model with high accuracy using camera images from all cameras.
[0009] FIG. 1 is a diagram showing an example of the outline of a photographing system composed of a plurality of cameras. In the imaging system, a plurality of cameras 1A to 1T image the target space to generate material data for virtual viewpoint images. Here, the plurality of cameras 1A to 1T are arranged in an ellipse so as to surround the target space. However, the number of cameras and the camera arrangement are not limited. The images captured by each of the cameras 1A to 1T are input to the material generation device 10. Note that information regarding the camera, such as the posture of the camera, the position of the camera, the angle of view of the camera, the distance to the imaging target, the focal length, the width of the image, the height of the image, the aperture value, the shutter speed, and the ISO sensitivity, is set in advance in the material generation device 10.
[0010] FIG. 2 is a diagram showing an example of the configuration of the material generation device 10. The configuration shown in FIG. 2 is realized by the CPU of the material generation device 10 executing a program recorded in the non-volatile memory. The material generation device 10 includes an image acquisition unit 100, a storage target camera selection unit 101, a foreground / background separation unit 102, a 3D model generation unit 103, a data storage unit 104, and a data output unit 105. The image acquisition unit 100 acquires camera images from each camera. The image acquisition unit 100 outputs the acquired camera images to the foreground / background separation unit 102.
[0011] The storage target camera selection unit 101 selects, based on information regarding the camera set in advance, a camera that captures a foreground image to be stored as material data and a camera that captures a background image to be stored as material data. In the present embodiment, a camera that captures a foreground image to be stored as material data is referred to as a foreground storage target camera, and a camera that captures a background image to be stored as material data is referred to as a background storage target camera. The storage target camera selection unit 101 outputs the identification information of the foreground storage target camera and the identification information of the background storage target camera to the foreground / background separation unit 102. The identification information is information for uniquely identifying the camera, and for example, a camera ID in the imaging system can be used.
[0012] Here, a method for the storage target camera selection unit 101 to select a camera will be described. When arranged on the circumference of an ellipse for photographing a target space as in the photographing system of this embodiment, the memory target camera selection unit 101 does not redundantly select cameras that are close in position or orientation (posture) among the plurality of cameras. Specifically, when selecting a foreground memory target camera among the cameras arranged as shown in FIG. 1, it is preferable to select every other (or every one or more) cameras such as camera 1A, 1C, 1E,... based on the position of the cameras, and not to select adjacent cameras. Also, when selecting a foreground memory target camera, cameras with different focal lengths or different resolutions may be selected.
[0013] On the other hand, when selecting a background memory target camera, since the consistency of color and luminance of the background is important, a camera with a short focal length that can photograph as large an area of the target space as possible is selected. Specifically, when selecting a background memory target camera among the cameras shown in FIG. 1, it is preferable to select the cameras arranged at the four corners of cameras 1C, 1I, 1M, and 1S. These four cameras have a short focal length and a long distance to the imaging target, and are suitable for photographing background images. Therefore, as the background memory target camera, a camera with a shorter focal length than the foreground memory target camera or a camera with a longer distance to the center of the imaging target than the foreground memory target camera is selected.
[0014] The memory target camera selection unit 101 preferably selects a camera based on information about the camera, but is not limited to such a method. For example, it may be selected based on at least any one of information on lens performance such as lens resolution, less aberration, color reproducibility, and sensor information indicating whether it is a high-sensitivity camera. For example, when there is no large difference in the positions of the cameras, the memory target camera selection unit 101 can compare the lens performance and select a camera with a lens having higher performance such as resolution and aberration. Also, a rectangular parallelepiped virtually placed in the target space can be projected onto the camera image by three-dimensional calculation, and selected from the overlapping degree of the imaging areas of each camera.
[0015] The foreground / background separation unit 102 separates the foreground and the background based on the camera image output from the image acquisition unit 100. Specifically, the foreground / background separation unit 102 generates a background image using a plurality of frames in all camera images. For example, the foreground / background separation unit 102 compares images of a plurality of frames to detect a moving area and a non-moving area, and updates the background image only with the non-moving area to generate the latest background image. Also, for all cameras, the foreground / background separation unit 102 compares the camera image output for each camera with the background image generated from the camera image, and determines a pixel with a difference in pixel value equal to or greater than a preset threshold as a foreground pixel. The foreground / background separation unit 102 generates a binary foreground silhouette image indicating whether or not it is a foreground area based on the foreground pixels, and outputs it to the 3D model generation unit 103. The foreground silhouette image is the same size as the image output from the camera.
[0016] Also, the foreground / background separation unit 102 acquires the identification information of the foreground storage target camera selected by the storage target camera selection unit 101, and outputs the foreground image generated from the camera image of the foreground storage target camera to the data storage unit 104. The data storage unit 104 stores the foreground image output from the foreground / background separation unit 102. Here, the foreground image is a group of images obtained by cutting out a foreground area obtained by combining consecutive foreground pixels into a rectangle based on the generated foreground silhouette image. Also, the foreground / background separation unit 102 acquires the identification information of the background storage target camera selected by the storage target camera selection unit 101, and outputs the background image generated from the camera image of the background storage target camera to the data storage unit 104. The data storage unit 104 stores the background image output from the foreground / background separation unit 102.
[0017] The 3D model generation unit 103 generates a 3D model based on the foreground silhouette image output from the foreground / background separation unit 102. In this embodiment, a 3D model is generated using the volumetric intersection method. The volumetric intersection method tiles a rectangular parallelepiped of a certain size in the generation target space of the 3D model, projects each rectangular parallelepiped three-dimensionally onto the foreground silhouette image of each camera, and determines whether it overlaps with the foreground pixels. By repeatedly performing the process of determining the overlap with the foreground pixels and determining whether it is a rectangular parallelepiped that forms the 3D model, a 3D model is generated. For example, when the 3D model generation unit 103 determines that a certain rectangular parallelepiped is a rectangular parallelepiped that forms the 3D model for all cameras, it retains it as the 3D model. The 3D model generation unit 103 determines all the rectangular parallelepipeds, retains the rectangular parallelepipeds that construct the 3D model in the generation target space, and outputs the point cloud data indicating the 3D model constructed with its coordinate information to the data storage unit 104. Also, the 3D model generation unit 103 outputs the foreground silhouette image to the data storage unit 104. The data storage unit 104 stores the point cloud data and the foreground silhouette image. However, the 3D model generation unit 103 may output elements such as regions or voxels instead of the point cloud data.
[0018] The data output unit 105 acquires and outputs the necessary data from the data storage unit 104 in response to a request from an external device. Here, the request from the external device is an output request for data such as the foreground image, 3D model, etc. stored in the data storage unit 104 and the time information associated with the data. The time information includes the time when the camera image that is the source of the data was captured.
[0019] Next, the hardware configuration of the material generation device 10 will be described with reference to FIG. 13. Note that the hardware configurations of the image processing device and the image generation device described later are the same as the configuration of the material generation device 10 described below. The material generation device 10 includes a CPU 1311, a ROM 1312, a RAM 1313, an auxiliary storage device 1314, a communication I / F 1315, and a bus 1316.
[0020] The CPU 1311 controls the entire material generation device 10 using computer programs and data stored in the ROM 1312 and the RAM 1313, thereby realizing each function of the material generation device 10 shown in FIG. 2. Note that the material generation device 10 may have one or more dedicated hardware different from the CPU 1311, and at least a part of the processing by the CPU 1311 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor).
[0021] The ROM 1312 stores programs and the like that do not require modification. The RAM 1313 temporarily stores programs and data supplied from the auxiliary storage device 1314, and data and the like supplied from the outside via the communication I / F 1315. The auxiliary storage device 1314 is composed of, for example, a hard disk drive, and stores various data such as image data and audio data. The communication I / F 1315 is used for communication with devices external to the material generation device 10. For example, when the material generation device 10 is connected to an external device by wire, a communication cable is connected to the communication I / F 1315. When the material generation device 10 has a function of wireless communication with an external device, the communication I / F 1315 includes an antenna. The bus 1316 connects each part of the material generation device 10 and transmits information.
[0022] Next, an example of the processing by the material generation device 10 will be described with reference to the flowchart of FIG. 3. The flowchart of FIG. 3 is realized by the CPU of the material generation device 10 executing a program recorded in a non-volatile memory. In S101, the foreground / background separation unit 102 separates the foreground and the background based on the camera image, generates a foreground image, a background image, and a foreground silhouette image, and outputs the foreground silhouette image to the 3D model generation unit 103.
[0023] In S102, the foreground / background separation unit 102 outputs a foreground image generated from the camera image of the foreground memory target camera selected by the memory target camera selection unit 101 to the data storage unit 104. The data storage unit 104 stores the foreground image. In S103, the foreground / background separation unit 102 outputs a background image generated from the camera image of the background memory target camera selected by the memory target camera selection unit 101 to the data storage unit 104. The data storage unit 104 stores the background image. In S104, the 3D model generation unit 103 generates a 3D model based on the foreground silhouette image, and outputs the 3D model and the foreground silhouette image to the data storage unit 104. The data storage unit 104 stores the 3D model and the foreground silhouette image.
[0024] In this way, the data storage unit 104 does not store all the foreground images and background images, but stores the foreground images and background images obtained from the images captured by the selected cameras among the plurality of cameras. Therefore, the data to be stored can be reduced, and the disk access bandwidth can be made small. Also, the memory target camera selection unit 101 can appropriately generate material data while reducing an increase in processing load by selecting a camera suitable for obtaining a foreground image and a background image based on information about the camera and the like. Also, when the 3D model generation unit 103 generates a 3D model, since the 3D model is generated from the camera images of more cameras than the foreground memory target camera, here, the camera images of all the cameras, a 3D model can be generated with higher accuracy. Note that the 3D model may be generated based only on the foreground silhouette image generated from the camera image of the foreground memory target camera.
[0025] Note that the 3D model generation unit 103 may determine whether the output point cloud data is visible from each camera, generate visible determination information in advance, and the data storage unit 104 stores the visible determination information. By storing the visible determination information and using it when generating a virtual viewpoint image, the virtual viewpoint image can be generated at high speed. Here, a method for determining whether each point of the point cloud data generated by the 3D model generation unit 103 is visible from each camera used for foreground processing will be described. The 3D model generation unit 103 projects the points constituting the 3D model three-dimensionally onto each camera image or foreground silhouette image, and determines that it is visible when it overlaps with the pixel indicating the foreground pixel. When generating a virtual viewpoint image, the 3D model from the virtual viewpoint specified by the user is drawn. Specifically, the foreground image is colored. To draw the 3D model from the virtual viewpoint, based on the visibility determination information indicating whether the 3D model is visible from the camera, and the position and orientation of the virtual viewpoint and the camera, a camera that captures the foreground image to be colored is selected. Therefore, in the present embodiment, when the 3D model generation unit 103 performs visibility determination, by setting only the foreground storage target camera as the target of visibility determination, the processing load related to visibility determination can be reduced, and the data when storing the visibility determination information can be reduced.
[0026] Also, in the present embodiment, the case where the storage target camera selection unit 101 determines and selects a camera has been described, but it is not limited to this case. For example, when configuring the above-described imaging system and the material generation device 10, the foreground storage target camera, the background storage target camera, and the camera for generating the 3D model may be set in advance, and the storage target camera selection unit 101 may select a camera based on the preset information. Also, in the present embodiment, the case where the foreground silhouette image is used to generate the 3D model has been described, but it is not limited to this case. For example, a model may be separately generated using a depth camera that performs distance measurement, and coloring may be performed using the foreground image obtained by being captured by the above-described foreground storage target camera.
[0027] (Second Embodiment) Next, the material generation device 20 according to the second embodiment will be described. Generally, when generating a 3D model, the processing load increases as the number of cameras increases. In the above-described view volume intersection method, since a determination process is performed for each camera to determine whether a foreground is reflected or not, the load of the determination process increases in proportion to the number of cameras. On the other hand, if the number of cameras is too small, when generating a 3D model, the range not reflected in the cameras increases, and for an object that actually exists, the 3D model is generated in a different form, resulting in a deterioration in the quality of the 3D model.
[0028] In addition, in order to increase the image resolution of the object reflected as the foreground, some cameras of the imaging system may be cameras equipped with lenses having a long focal length to generate an object with high resolution. In this case, the camera has a narrow shooting range. When generating a 3D model, it includes a process of projecting a certain point onto the camera image of the camera by three-dimensional calculation. A camera with a narrow shooting range is more likely to be determined not to be reflected as a result of the calculation, resulting in an increase in the processing load. The material generation device 20 of the present embodiment reduces the processing load when generating a 3D model by selecting a camera to be photographed for generating the 3D model from a plurality of cameras.
[0029] FIG. 4 is a diagram showing an example of the configuration of the material generation device 20. The same components as those in the first embodiment are denoted by the same reference numerals and the description thereof is omitted. The 3D model generation unit 203 generates a 3D model using the camera image of the camera selected by the model generation camera selection unit 207. Hereinafter, the camera selected by the model generation camera selection unit 207 is referred to as a 3D model generation target camera. In the view volume intersection method, when determining whether it is a rectangular parallelepiped forming a 3D model, representative points of the rectangular parallelepiped are projected onto the foreground silhouette image of each camera by three-dimensional calculation. In the present embodiment, only the 3D model generation target camera is the target of the projection process. Therefore, the process of projecting onto the foreground silhouette image and the determination process of determining whether it is a rectangular parallelepiped can be reduced, and the processing load when generating a 3D model can be reduced.
[0030] The model generation camera selection unit 207 selects a camera to be used for the process of generating a 3D model based on information regarding cameras set in advance. Here, a method by which the model generation camera selection unit 207 selects a camera will be described. When cameras are arranged in an elliptical shape for photographing a target space as in the photographing system similar to the first embodiment, the model generation camera selection unit 207 does not redundantly select cameras that are close in position and orientation among the plurality of cameras. That is, the model generation camera selection unit 207 selects a camera based on the position and orientation of the camera. Further, the model generation camera selection unit 207 may evaluate the photographing range based on information regarding the focal length and the distance to the photographing target in consideration of reducing the processing load for generating the 3D model, and select a camera with a wider photographing range. For example, a camera with a wider photographing range than the foreground memory target camera can be selected. Furthermore, the model generation camera selection unit 207 may select a camera based on the same information as in the first embodiment. The method by which the model generation camera selection unit 207 selects a camera is not limited to the method described above, and the method described above may be used alone or in combination for selection. For example, the model generation camera selection unit 207 may select a camera based only on information regarding the photographing range.
[0031] Next, an example of the process by the material generation apparatus 20 will be described with reference to the flowchart of FIG. 3. Since the process different from the first embodiment is S104, S104 will be described. In S104, the 3D model generation unit 103 generates a 3D model from the camera image photographed by the 3D model generation target camera selected by the model generation camera selection unit 207. That is, in the present embodiment, the 3D model generation unit 103 stores, in S104, the silhouette image generated from the camera image photographed by the 3D model generation target camera and the 3D model generated from the silhouette image.
[0032] In this way, when the 3D model generation unit 103 generates a 3D model by generating a 3D model from an image captured by the camera selected by the model generation camera selection unit 207, the processing load can be reduced. That is, the 3D model generation unit 103 can reduce the processing load when generating a 3D model by omitting processing on an image captured by a camera not selected by the model generation camera selection unit 207. Further, the model generation camera selection unit 207 can suppress a decrease in the quality of the 3D model by selecting a camera suitable for generating a 3D model based on information about the camera and the like. Note that the foreground memory target camera, the background memory target camera described in the first embodiment, and the 3D model generation target camera described in this embodiment may be selected independently or may be selected repeatedly.
[0033] (Third Embodiment) Next, the image processing system 1 according to the third embodiment will be described. The image processing system 1 includes an image processing device 30 connected to each of a plurality of cameras, and a material generation device 40 that generates material data of a virtual viewpoint image using the camera images output from the image processing device 30. The image processing device 30 of this embodiment separates the foreground and the background based on the camera image, and reduces the transmission amount by transmitting only the region extracted as the foreground to the material generation device 40. Further, the material generation device 40 of this embodiment selects a camera that captures the foreground image to be stored as the material data and a camera that captures the background image to be stored as the material data, and reduces the transmission amount by causing the image processing device 30 to transmit the foreground image or the background image for each image processing device 30.
[0034] FIG. 5 is a diagram showing an example of the configuration of the image processing system 1 having the image processing device 30 and the material generation device 40. Note that the same configuration as that in the first embodiment is denoted by the same reference numeral and the description thereof is omitted. The configuration of each device is realized by the CPU executing a program recorded in the nonvolatile memory. The image processing device 30 is composed of a plurality of units, each of which is connected to a camera of the imaging system. Therefore, in the image processing system 1, there are as many image processing devices 30 as the number of cameras. The image processing device 30 processes the camera image and transmits various data to the material generation device 40. The image processing device 30 includes an image acquisition unit 301, a foreground / background separation unit 302, and a data transmission unit 303. The image acquisition unit 301 acquires a camera image from one camera. The image acquisition unit 301 outputs the acquired camera image to the foreground / background separation unit 302.
[0035] The foreground / background separation unit 302 generates a background image, a foreground image cut out in a rectangle, and a foreground silhouette image from the camera image. This process is the same as that of the foreground / background separation unit 102 in the first embodiment. The foreground / background separation unit 302 outputs the foreground silhouette image to the data transmission unit 303. Here, the foreground / background separation unit 302 receives, from the transmission camera selection unit 401 of the material generation device 40, the identification information of the camera that captures the foreground image to be stored as material data and the identification information of the camera that captures the background image to be stored as material data. In this embodiment, the camera that captures the foreground image to be stored as material data is referred to as the foreground transmission target camera, and the camera that captures the background image to be stored as material data is referred to as the background transmission target camera. The identification information is information for uniquely identifying the camera, and for example, the camera ID in the imaging system can be used. When the identification information of the received foreground transmission target camera matches the identification information of the connected camera, the foreground / background separation unit 302 outputs the generated foreground image to the data transmission unit 303. Also, when the identification information of the received background transmission target camera matches the identification information of the connected camera, the foreground / background separation unit 302 outputs the generated background image to the data transmission unit 303. Note that depending on the selection method by the transmission camera selection unit 401, it may output both the foreground image and the background image.
[0036] The data transmission unit 303 transmits the data output from the foreground / background separation unit 302 to the material generation device 40. From the data transmission unit 303 to the material generation device 40, they may be connected, for example, by a LAN (Local Area Network) cable or the like. Also, the data transmission unit 303 transmits using the network as TCP packet data defined by TCP / IP which combines Transmission Control Protocol (TCP) and Internet Protocol (IP). However, the physical cable and protocol used for transmission are not limited, and for example, Serial Degital Interface or the like may be used. Also, the data transmission unit 303 may transmit by wireless communication such as Wi-Fi defined by the IEEE802.11 standard.
[0037] The material generation device 40 includes a transmission camera selection unit 401, a received data processing unit 402, a 3D model generation unit 103, a data storage unit 104, and a data output unit 105. The transmission camera selection unit 401 selects a foreground transmission target camera and a background transmission target camera from the cameras of the imaging system. The transmission camera selection unit 401 transmits the identification information of the foreground transmission target camera and the identification information of the background transmission target camera to each image processing device 30. Here, the method by which the transmission camera selection unit 401 selects a camera can be selected based on the same information as in the first embodiment. The received data processing unit 402 receives a foreground image, a background image, and a foreground silhouette image from each image processing device 30. The received data processing unit 402 outputs the foreground image and the background image to the data storage unit 104. The data storage unit 104 stores the foreground image and the background image output from the received data processing unit 402. Also, the received data processing unit 402 outputs the foreground silhouette image to the 3D model generation unit 103. The 3D model generation unit 103 generates a 3D model based on the foreground silhouette image output from the received data processing unit 402.
[0038] In this way, the data transmission unit 303 of the image processing apparatus 30 transmits the foreground image and the background image generated from the image captured by the camera selected by the transmission camera selection unit 401. Therefore, the amount of data transmission can be reduced for the entire image processing system. Note that the transmission camera selection unit 401 selects a camera suitable for capturing the foreground image and the background image based on information about the camera and the like, so that the data transmission unit 303 can transmit more effective data.
[0039] (Fourth Embodiment) Next, the image processing system 2 according to the fourth embodiment will be described. In this embodiment, the camera information of the foreground memory target camera, the background memory target camera, and the 3D model generation target camera is stored, and when generating a virtual viewpoint image, the camera information of the camera used for each process is given as metadata. A user viewing the virtual viewpoint image can estimate the quality of the virtual viewpoint image by checking the camera information used for each process. For example, an image in which 3D model generation is performed using more cameras can be estimated to generate a 3D model with higher accuracy and better image quality. Therefore, it becomes easier for the user to search for desired data by using the camera information.
[0040] FIG. 6 is a diagram showing an example of the configuration of the image processing system 2 having the material generation apparatus 50 and the image generation apparatus 60. Note that the same components as those in the first and second embodiments are denoted by the same reference numerals and the description thereof is omitted. The configuration of each apparatus is realized by the CPU executing a program recorded in a non-volatile memory. The material generation apparatus 50 includes an image acquisition unit 100, a foreground / background separation unit 102, a data storage unit 104, a data output unit 105, a 3D model generation unit 203, a memory target camera selection unit 501, and a model generation camera selection unit 502.
[0041] The memory target camera selection unit 501 outputs the camera information of the foreground memory target camera and the camera information of the background memory target camera to the data storage unit 104 in addition to the function of the memory target camera selection unit 101 of the first embodiment. The data storage unit 104 stores the camera information output from the memory target camera selection unit 501. In this case, the data storage unit 104 stores the foreground image output from the foreground memory target camera and the camera information of the foreground memory target camera in association with each other, and stores the background image output from the background memory target camera and the camera information of the background memory target camera in association with each other. The model generation camera selection unit 502 outputs the camera information of the 3D model generation target camera to the data storage unit 104 in addition to the function of the model generation camera selection unit 207 of the second embodiment. The data storage unit 104 stores the camera information output from the model generation camera selection unit 502. In this case, the data storage unit 104 stores the 3D model output from the 3D model generation target camera and the camera information of the 3D model generation target camera in association with each other. Here, the camera information includes information on the lens performance, information on the lens performance, and sensor information in addition to the information on the camera described above.
[0042] The image generation device 60 includes a data acquisition unit 601, a virtual viewpoint image file generation unit 602, a use camera metadata adding unit 603, and a file output unit 604. The data acquisition unit 601 acquires the material data of the virtual viewpoint image from the material generation device 50 and outputs it to the virtual viewpoint image file generation unit 602. Also, the data acquisition unit 601 acquires the camera information of the foreground memory target camera, the background memory target camera, and the 3D model generation target camera, which are respectively associated with the material data, from the material generation device 50, and outputs it to the used camera metadata adding unit 603. When the data acquisition unit 601 acquires the material data, it designates the time information and the data type associated with the material data to the data output unit 105. Here, as the time information, for example, date, hour, minute, second, frame, etc. are designated. Also, as the data type, foreground image, background image, and 3D model are designated. Therefore, the data acquisition unit 601 can acquire the foreground image, background image, and 3D model of a certain frame, and the virtual viewpoint image file generation unit 602 can generate the virtual viewpoint image in that frame.
[0043] The virtual viewpoint image file generation unit 602 generates a virtual viewpoint image based on the input virtual viewpoint information and converts it into a video file. The virtual viewpoint image file generation unit 602 outputs the generated video file to the used camera metadata adding unit 603. Note that the virtual viewpoint information at least includes information regarding the position and orientation of the virtual viewpoint. The virtual viewpoint information is input by the user or the operator using a UI (not shown). However, the virtual viewpoint information may be automatically set by the image generation device 60. Here, an example of a method for generating a virtual viewpoint image will be described. The virtual viewpoint image file generation unit 602 projects the points constituting the model onto the virtual viewpoint image for the 3D model visible from the input virtual viewpoint. The color of the projected points is colored using the foreground image of the camera with an angle close to the virtual viewpoint. Thereby, a foreground image seen from the virtual viewpoint can be generated. Also, the virtual viewpoint image file generation unit 602 generates a background image seen from the virtual viewpoint by projecting the 3D model onto the virtual viewpoint image using the background model for the background image. The virtual viewpoint image file generation unit 602 synthesizes the background image and the foreground image to generate a virtual viewpoint image.
[0044] Next, an example of a method for converting into a video file will be described. First, the virtual viewpoint image file generation unit 602 compresses the virtual viewpoint image into a video according to the H.265 (ISO / IEC 23008-2 HEVC) standard. Further, the virtual viewpoint image file generation unit 602 converts it into a video file in the form of an mp4 file in the MPEG-4 Part14 format defined in ISO / ICE 14496-14:2003, ISO / ICE JTC1.
[0045] The used camera metadata adding unit 603 adds, as file metadata, each process used camera information including the camera information of the camera used in each process. The adding method in each standard will be described later. The file output unit 604 outputs the generated video file.
[0046] FIG. 7 is a diagram showing an example of each process used camera information 700. Each process used camera information 700 has, at the head, the number of used camera information and a pointer 701 to each used camera information, and then the used camera information 702 for each process is arranged. Here, the used camera information 702 includes, for example, the camera information of all the cameras used when generating the foreground image as material data. The used camera information 702 is composed of a process ID 703, the number of cameras used in the process 704, and camera information 705. The camera information 705 is information for one camera and includes at least one of the following information. That is, the attitude parameter 706, the position parameter 707, the field angle parameter 708, the distance to the shooting target 709, the focal length 710, the width of the image 711, the height of the image 712, the aperture value 713, the shutter speed 714, the ISO sensitivity (speed rate) 715, etc. are included. The attitude parameter 706 is represented by a quaternion expressed by, for example, the following formula (1).
[0047]
Equation
[0048] Here, the left side of the semicolon is the real part, and x, y, and z represent the imaginary parts. Assume that the position parameter 707 is the coordinates (x, y, z) of a three-dimensional coordinate system with the origin on the world coordinates being (0, 0, 0). Assume that the angle-of-view parameter 708 is the horizontal angle-of-view of the virtual viewpoint. The expression for the horizontal angle-of-view is not limited. For example, it may be expressed as an angle in the range of 0 degrees to 180 degrees, or it may be expressed in terms of the focal length when the standard for a 35mm film camera is 50mm. In addition, other camera information may include information such as color temperature, focal length of the lens, lens model number, camera model number, etc. However, the camera information is not limited to the above-mentioned information.
[0049] Next, an example of the method for attaching metadata in each standard will be described. Here, the method for attachment in the case of the Camera Image Machinery Industry Association Standard DC-008-2012 Image File Format Standard Exif2.3 for Digital Still Cameras will be described. FIG. 8(a) is a diagram showing an example of attaching each processing-used camera information 700 to the format defined by the Exif standard. FIG. 8(a) defines a Processing Using Camera Image File Directory (hereinafter referred to as PUC IFD) 802 and stores each processing-used camera information 700. The PUC IFD Pointer 801 is a pointer indicating the PUC IFD 802.
[0050] FIG. 9 is a diagram showing an example of the tag information of the PUC IFD. The version of the PUC tag starts with a value of 1 and represents the version of the data format that follows. The attitude parameter is represented by a quaternion, and each value of the above-mentioned real part and imaginary part is represented by a 4-byte signed integer. The position parameter is such that each value of the coordinates x, y, z from the origin of the arena is represented by a 4-byte signed floating-point number. The angle-drawing parameter represents the horizontal angle and is represented by a 4-byte signed floating-point number. The distance to the object to be photographed indicates the distance in millimeters and is represented by a 4-byte unsigned integer. The focal length is represented in millimeters and is represented by a fraction of a 4-byte unsigned integer. The first is the numerator and the second is the denominator. The width of the image indicates the width of the captured image by the camera and is represented by a 4-byte signed integer. The height of the image indicates the height of the captured image by the camera and is represented by a 4-byte unsigned integer. The aperture value indicates the aperture value of the shooting parameter of the camera and is represented by a fraction of a 4-byte unsigned integer. The first is the numerator and the second is the denominator. The shutter speed indicates the shutter speed of the shooting parameter of the camera and is represented by a fraction of a 4-byte signed integer. The first is the numerator and the second is the denominator. The ISO sensitivity (speed rate) indicates the ISO sensitivity of the shooting parameter of the camera and is represented by a 2-byte unsigned integer. However, the order and data length of the above-mentioned information are not limited. Also, it may be composed of some of the above-mentioned information.
[0051] Figure 8(b) is a diagram showing an example of storing each processing-used camera information 700 using the undefined APP3(811), which is an undefined APPn marker that can be arbitrarily used by a vendor or an industry group and is not defined in the Exif specification. In this way, an area for storing each processing-used camera information 700 can be added and defined to the existing Exif specification, which is a still image format, to generate a virtual viewpoint image with virtual viewpoint parameters.
[0052] Next, the method of attachment in the case of the ISO / IEC 14496-12 (MPEG-4 Part12) ISO base media file format (hereinafter referred to as ISO BMFF) specification will be described. ISO BMFF treats a file in units of boxes that store information indicating size and type and data. Figure 10(a) is a diagram showing an example of the structure of a box. Note that, as shown in Figure 10(b), it is also possible to have a structure in which a box is included as data within a box.
[0053] Figure 11 is a diagram showing an example of the data structure of an ISO BMFF file. An ISO BMFF file is composed of boxes such as ftyp1101 (File Type Compatibility Box), moov1102 (Movie Box), and mdat1103 (Media Data Box). The box ftyp1101 contains information about the file format, such as that the file is an ISO BMFF file, the version of the box, the name of the manufacturer that created the file, etc. The box moov1102 contains metadata such as a time axis and addresses for managing media data. The box mdat1103 contains the media data that is actually played as a video.
[0054] Figure 12 is a diagram showing an example of attaching each processing used camera information 700 to the box moov1102. As shown in Figure 12(a), each processing used camera information 700 can be attached to the box meta1201 that shows the meta information of the entire file. Also, in the case of a video file edited by stitching together different images for each track, as shown in Figure 12(b), it is also possible to attach each processing used camera information 700 to the box meta1202 of each track.
[0055] For example, aligned(8) class MetaBox (handler_type) extends FullBox('meta', version = 0, 0) { HandlerBox(handler_type) theHandler; PrimaryItemBox primary_resource; / / optional DataInformationBox file_locations; / / optional ItemLocationBox item_locations; / / optional ItemProtectionBox protections; / / optional ItemInfoBox item_infos; / / optional IPMPControlBox IPMP_control; / / optional ItemReferenceBox item_refs; / / optional ItemDataBox item_data; / / optional Virtual_ViewPoint_Camera_Info / / optional Box other_boxes[]; / / optional } In Virtual_ViewPoint_Camera_Info represents the camera information 700 used in each process. This box is Box Type: 'vvci' Container: Meta box ('meta') Mandatory: No Quantity: Zero or one The syntax is aligned(8) class ItemLocationBox extends FullBox('vvci',version,0) { unsigned int(32) offset_size; unsigned int(32) length_size; unsigned int(32) base_offset_size; if (version == 1) { unsigned int(32) index_size; } else { unsigned int(32) reserved; } for (i = 0; i < 4; i++) { int(32) Rotetion Quaternion[i]; / / Attitude parameter } for (i = 0; i < 3; i++) { float(32) Translation_Vector[i]; / / Position parameter } float(32) Horizontal_Angle; / / Horizontal angle parameter unsigned int(32) Distance_to_Target; / / Distance to the shooting target unsigned int(32) Focal_Length_Numerator / / Focal length, numerator unsigned int(32) Focal_Length_Denominator / / Focal length, denominator unsigned int(32) Image_Width; / / Image width unsigned int(32) Image_Length; / / Image height unsigned int(32) Aperture_Value_Numerator / / Aperture value, numerator unsigned int(32) Aperture_Value_Denominator / / Aperture value, denominator int(32) Shutter_Speed_Numerator / / Shutter speed, numerator int(32) Shutter_Speed_Denominator / / Shutter speed, denominator short(16) ISO_Speed_Rating / / ISO sensitivity } It becomes.
[0056] In this way, the data storage unit 104 of the material generation device 50 stores the camera information of the camera used in each process when generating the material data of the virtual viewpoint image. Therefore, the image generation device 60 can attach the camera information of the camera used in each process as file metadata. In this way, by attaching the camera information of the camera used in each process to the file, the user can confirm the quality of the virtual viewpoint image based on the attached camera information. For example, when the user plays a video to which the camera information of the camera used in each process is attached, the user can confirm the video quality by checking the camera information attached to the video with a property or the like using an image playback device (not shown). Also, when the user searches for a file, it can be used for search conditions such as the number of cameras used for 3D model generation being 30 or more, and higher quality videos can be searched for. Note that in this embodiment, the case where the still image format is the Exif standard or the ISO BMFF standard has been described, but it is not limited to this case, and other standards or unique formats may also be used. Also, the values and expressions of the respective parameters are not limited to the cases described above.
[0057] (Other Embodiments) The present invention is also realized by executing the following processing. That is, a program that realizes the functions of the above-described embodiment is supplied to a system or device via a network or various recording media, and a computer (such as a CPU or MPU) of the system or device reads and executes the program code for control. In this case, the program and the recording medium on which the program is recorded constitute the present invention.
[0058] As described above, the present invention has been described in detail based on the above-described embodiments, but the present invention is not limited to the above-described embodiments, and various forms within the scope not departing from the gist of the present invention are also included. Furthermore, each of the above-described embodiments merely shows one embodiment of the present invention, and the embodiments can be combined as appropriate.
Explanation of Reference Numerals
[0059] 10: Material generation device 100: Image acquisition unit 102: Foreground / background separation unit
Claims
1. a first generation device that generates a foreground image used for generating a foreground area of a virtual viewpoint image corresponding to a virtual viewpoint, the foreground image indicating a foreground included in the first image, based on a first image acquired by a first imaging means, and outputs the foreground image; a second generation device that generates a background image used for generating a background region of the virtual viewpoint image, the background image indicating a background included in the second image, based on a second image acquired by a second image capturing means having a focal length shorter than a focal length of the first image capturing means, and outputs the background image; a third generating device that receives the foreground image output from the first generating device and the background image output from the second generating device, and generates three-dimensional shape data of the foreground based on the foreground image output by the first generating device; The image processing system is characterized in that the second generating device is configured not to output to the third generating device a foreground image showing the foreground contained in the second image and used to generate three-dimensional shape data of the foreground.
2. 2. The image processing system according to claim 1, wherein a distance from said second image capturing means to a subject to be photographed is longer than a distance from said first image capturing means to said subject to be photographed.
3. The third generating device determines that an element included in the three-dimensional shape data of the foreground is visible from the first photographing means when a projection point onto the first image of the element overlaps with a pixel of the foreground included in the first image; 3. The image processing system according to claim 1, wherein the virtual viewpoint image is generated based on a result of the determination.
4. An image processing system as described in any one of claims 1 to 3, further characterized in that it has a fourth generation device that generates a foreground region of the virtual viewpoint image based on a three-dimensional model of the foreground and the foreground image, generates a background region of the virtual viewpoint image based on a three-dimensional model of the background and the background image, and synthesizes the foreground region of the virtual viewpoint image and the background region of the virtual viewpoint image.
Citation Information
Patent Citations
Image processing system, image processor, control method, and program
JP2017211828A
Generation device, generation method and program of virtual view point image
JP2018107793A
Image generation device, image generation method and program
JP2019003320A
Information processing apparatus, image processing system, control method, and program
JP2019022151A