Information processing device, method and program performed by the information processing device
The system addresses image quality issues in virtual viewpoint generation by excluding non-static background regions and retransmitting images to enhance three-dimensional model accuracy and reduce data volume, resulting in high-quality virtual images.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-06
- Publication Date
- 2026-05-19
AI Technical Summary
Existing methods for generating virtual viewpoint images from multiple viewpoints suffer from deteriorated image quality due to difficulties in extracting foreground silhouettes when overlapping with non-static backgrounds, leading to inaccuracies and increased data volume.
A system that generates a three-dimensional model by excluding regions where foreground extraction is difficult and retransmits images from these areas, using a viewing volume cross-eyed method and perspective projection to enhance image quality.
This approach enables the generation of high-quality virtual viewpoint images with improved accuracy and reduced data volume, particularly in replay and highlight scenes, by incorporating retransmitted images to complete the three-dimensional model.
Smart Images

Figure 2026082049000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, a method performed by the information processing apparatus, and a program.
Background Art
[0002] There is known a technique of installing a plurality of imaging devices at different positions, synchronously photographing an area (or subject) where at least a part overlaps at a plurality of viewpoints, and generating a virtual viewpoint image from a three-dimensional model generated using a plurality of viewpoint images obtained by the photographing.
[0003] Non-Patent Document 1 describes a volume intersection method for generating a three-dimensional model based on the silhouette of the foreground extracted from images photographed from a plurality of viewpoints (hereinafter, photographed images).
[0004] Here, an area where the difference in color and luminance from the background is larger than a set value is extracted as the silhouette of the foreground. However, if there are non-static objects such as spectators or electronic display boards in the background, the difference in color and luminance between the foreground and the background changes, resulting in misdetection and undetected of the silhouette.
[0005] Therefore, in order to suppress a decrease in the accuracy of the three-dimensional model and an unnecessary data amount, in an area where the background is not static, etc., the three-dimensional model is generated and colored by the silhouette of another imaging device and photographed images without extracting the silhouette of the foreground, and a virtual viewpoint image is generated.
Prior Art Documents
Non-Patent Documents
[0006]
Non-Patent Document 1
[0007] However, with the aforementioned technology, if the foreground overlaps with an area where it is difficult to extract the foreground silhouette, a three-dimensional model is generated and colorized using an image captured by a separate imaging device only for the foreground in that overlapping area. As a result, the image quality of the three-dimensional model deteriorates, and consequently, the image quality of the virtual viewpoint image deteriorates.
[0008] This disclosure was made in view of the above-mentioned issues and provides a technology for appropriately generating virtual viewpoint images. [Means for solving the problem]
[0009] To solve this problem, for example, the information processing apparatus of the present disclosure has the following configuration: a receiving means for receiving an image showing a foreground region included in a captured image; a generating means for generating a three-dimensional model of the foreground based on the image showing the foreground; and a identifying means for identifying a region in the captured image that corresponds to the three-dimensional model of the foreground but is not included in the received image, wherein the receiving means receives a part of the captured image based on the identified region. [Effects of the Invention]
[0010] According to this disclosure, virtual viewpoint images can be appropriately generated. [Brief explanation of the drawing]
[0011] [Figure 1]A block diagram showing the overall configuration of the virtual viewpoint image generation system in the embodiment. [Figure 2] Functional block diagram of the three-dimensional model generation device in the first embodiment. [Figure 3] Functional block diagram of the imaging device in the first embodiment. [Figure 4] A diagram showing a flowchart of the imaging process of the imaging device in the first embodiment. [Figure 5] A diagram showing an example of the arrangement of multiple imaging devices in a virtual viewpoint image generation system. [Figure 6] A diagram showing an example of an image captured by the imaging device in the first embodiment. [Figure 7] A diagram showing an example of the data format of the transmitted data in the embodiment. [Figure 8A] This figure shows a flowchart of the data reception process of the three-dimensional model generation device according to the first embodiment. [Figure 8B] This figure shows a flowchart of the data retransmission request processing of the three-dimensional model generation device of the first embodiment. [Figure 9] A diagram showing the generated three-dimensional model projected onto an image from an imaging device. [Figure 10] Functional block diagram of the three-dimensional model generation device in the second embodiment. [Figure 11] This figure shows a flowchart of the data reception process of the three-dimensional model generation device according to the second embodiment. [Figure 12] A diagram showing an example of the importance of the area to be photographed in the second embodiment. [Figure 13] A block diagram showing the hardware configuration of the information processing device of the embodiment. [Modes for carrying out the invention]
[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the present disclosure, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.
[0013] (Virtual viewpoint image generation system) The virtual viewpoint image generation system of the embodiment installs a plurality of imaging devices at different positions, synchronously captures an area (also referred to as a shooting area) where at least a part including the subject overlaps at a plurality of viewpoints, and uses a plurality of captured images (also referred to as a plurality of viewpoint images) obtained by the shooting to generate a three-dimensional model and generate a virtual viewpoint image.
[0014] The virtual viewpoint image generation system can generate virtual viewpoint images for viewing highlight scenes of soccer and basketball from various angles, for example, by a technique of generating virtual viewpoint images from a plurality of viewpoint images. Thereby, the virtual viewpoint image generation system can generate an image that gives a higher sense of presence to the user than a normal image.
[0015] The virtual viewpoint image generation system aggregates the images captured by a plurality of imaging devices in a three-dimensional model generation device such as a server, and the three-dimensional model generation device performs processes such as generation and rendering of a three-dimensional model, and transmits it to the user terminal, thereby realizing generation and browsing of virtual viewpoint images based on images of a plurality of viewpoints.
[0016] (Model generation) The virtual viewpoint image generation system generates a three-dimensional model, for example, by a volume intersection method based on the silhouette of the foreground from images captured from a plurality of viewpoints. Here, the foreground may be an object to be a three-dimensional model. For example, if the shooting object is sports, the athlete performing the sports is the object of the three-dimensional model and thus becomes the foreground. The silhouette refers to the area occupied by the foreground in the captured image.
[0017] The virtual viewpoint image generation system generates a three-dimensional model of the foreground by using the viewing volume cross-section method to find the common region of the frustum formed by passing through each point on the contour of the foreground silhouette image from the projection center of multiple imaging devices. Furthermore, the virtual viewpoint image generation system represents the three-dimensional model as a collection of voxels, which are tiny cubes.
[0018] (coloring) The virtual viewpoint image generation system colors the foreground three-dimensional model by assigning the corresponding pixel color of the captured image to the voxels that make up the generated three-dimensional model, based on the position and orientation of the imaging device. The virtual viewpoint image generation system may generate a virtual viewpoint image by perspective projection transformation to an arbitrary viewpoint in three-dimensional space.
[0019] (Silhouette extraction) A virtual viewpoint image generation system may generally extract areas of foreground silhouettes where the difference in color and brightness between the foreground and background is greater than a set value. However, if there are non-stationary objects in the background, such as spectators or electronic billboards, the difference in color and brightness between the foreground and background changes, leading to false detections and failures to detect objects.
[0020] In order to suppress a decrease in model accuracy and the amount of unnecessary data, the virtual viewpoint image generation system of this embodiment does not extract the foreground silhouette in areas where the background is not static (also called non-extracted areas), and instead superimposes a silhouette based on an image cut out from the captured image (also called a cut-out image) onto those areas to generate a three-dimensional model.
[0021] <First Embodiment> A virtual viewpoint image generation system of the first embodiment will be described with reference to the drawings.
[0022] In this embodiment, we describe a process that enables the generation of a high-quality three-dimensional model by retransmitting images of foreground regions that were not transmitted because they were included in areas where it is difficult to extract the foreground silhouette based on the generated three-dimensional model.
[0023] This embodiment improves accuracy by retransmitting unreceived captured images based on a three-dimensional model of the foreground generated in real time, thus enabling the generation of high-quality virtual viewpoint images in replay scenes and highlight scenes even when the foreground overlaps with spectators, electronic billboards, etc.
[0024] (System Configuration) Figure 1 is a block diagram showing the overall configuration of a virtual viewpoint image generation system 1, which includes a three-dimensional model generation device and an imaging device, according to an embodiment.
[0025] The virtual viewpoint image generation system 1 includes an imaging device group 15, a control device 11, a three-dimensional model generation device 12, a storage device 13, and a rendering device 14. At least one of the imaging device group 15 and the control device 11 may be an example of an external device.
[0026] The imaging device group 15 has multiple imaging devices 10a to 10q. The multiple imaging devices 10a to 10q capture a shooting area including one or more foregrounds (also called subjects) from various angles and positions. In the following description, when it is not necessary to distinguish between the multiple imaging devices 10a to 10q, they may be referred to as imaging device 10. Imaging device 10 may be, for example, a digital camera that captures a shooting area including a foreground and generates captured image data such as video, still images, and footage. In the following description, the term "image" may include video, still images, footage, and their data.
[0027] The imaging device 10 extracts the foreground region from the captured image, excluding the region indicated by the non-foreground region extraction mask (also called the non-extracted region), which is a region in which the foreground is not extracted. Based on the extracted foreground region, the imaging device 10 generates a foreground mask image and a foreground texture image and outputs them to the three-dimensional model generation device 12.
[0028] The non-foreground region extraction mask may be a binary mask image in which the areas where the foreground region is not extracted (because the background is not static and extraction of the foreground region is difficult) are represented in white, and the other areas in black.
[0029] The foreground mask image may be a binary image in which the silhouette region representing at least a portion of the foreground from the captured image is shown in white and the background region in black. Alternatively, the foreground mask may be a binary image in which the silhouette region representing at least a portion of the foreground from the captured image is shown in black and the background region in white.
[0030] The foreground texture image may be an image that shows at least a portion of the foreground generated from the captured image. The foreground texture image is an example of a foreground image. In order to reduce the amount of data, the foreground texture image may be an image in which everything except the region showing at least a portion of the foreground (or the silhouette region showing at least a portion of the foreground) from the captured image is shown in a single color (e.g., black), or an image that has been cropped to show only the region showing at least a portion of the foreground.
[0031] The imaging device 10 of the imaging device group 15 stores the captured images. When the imaging device 10 receives a request from the control device 11 to retransmit the captured images, it extracts the requested time and region from the stored images and outputs it to the three-dimensional model generation device 12.
[0032] The imaging device group 15 calculates the amount of data to be output to the three-dimensional model generation device 12 and outputs it according to the request of the control device 11.
[0033] The control device 11 may be an information processing device also called a computer. The control device 11 calculates camera parameters indicating the position and orientation of each of the multiple imaging devices 10 and outputs them to the three-dimensional model generation device 12.
[0034] The control device 11 extracts areas where the background is not static from the images captured by the multiple imaging devices 10, generates a non-foreground region extraction mask in advance, and transmits it to the multiple imaging devices 10 and the three-dimensional model generation device 12. The non-foreground region extraction mask may also be created manually.
[0035] The control device 11 receives a retransmission request from the three-dimensional model generation device 12 and outputs a request to the imaging device group 15 for retransmission of the captured image.
[0036] The control device 11 controls the start and end of transmission and notifies the three-dimensional model generation device 12.
[0037] The control device 11 acquires the output data amount of the imaging device group 15 from the imaging device 10 of the imaging device group 15. Alternatively, the control device 11 may continuously acquire the output data amount from the imaging device 10 of the imaging device group 15.
[0038] The three-dimensional model generation device 12 may be an information processing device also known as a computer. The three-dimensional model generation device 12 generates a three-dimensional model for generating a virtual viewpoint image from images captured by multiple imaging devices 10. In this embodiment, the three-dimensional model generation device 12 may generate not only a three-dimensional model but also a virtual viewpoint image from the three-dimensional model.
[0039] In the specific generation of a three-dimensional model, the three-dimensional model generation device 12 receives camera parameters and a non-foreground region extraction mask from the control device 11, and also receives a foreground mask image and a foreground texture image from the imaging device group 15.
[0040] The 3D model generation device 12 generates a 3D model by the viewing volume cross-eyed method, excluding the region indicated by the non-foreground region extraction mask from the received foreground mask image and camera parameters. The 3D model generation device 12 outputs the generated 3D model, along with the foreground texture image, foreground mask image, and camera parameters, to the storage device 13.
[0041] Furthermore, the three-dimensional model generation device 12 of this embodiment also generates a three-dimensional model of the foreground of the region excluded by the region indicated by the non-foreground region extraction mask. Specifically, the three-dimensional model generation device 12 generates a three-dimensional model of the region based on the foreground texture image (hereinafter also referred to as the extracted image) of the image extracted from the excluded region from a previously captured image. The previously captured image referred to here may be the captured image that was the source from which the received foreground mask image and foreground texture image were generated. Also, the foreground mask image and foreground texture image and the previously captured image may be from the same time of capture.
[0042] The storage device 13 may be an information processing device also called a computer. The storage device 13 stores and stores camera parameters, the three-dimensional model generated by the three-dimensional model generation device 12, and foreground mask images and foreground texture images calculated by each of the multiple imaging devices 10.
[0043] The rendering device 14 may be an information processing device, also known as a computer. The rendering device 14 obtains a foreground texture image, a three-dimensional model, and camera parameters from the storage device 13.
[0044] The rendering device 14 generates a colored three-dimensional model by assigning pixel values of a foreground texture image corresponding to each voxel constituting the three-dimensional model based on camera parameters, and generates a virtual viewpoint image by performing a perspective projection transformation to an arbitrary viewpoint in three-dimensional space.
[0045] (Regarding camera parameters) Camera parameters include both external and internal parameters.
[0046] External parameters include rotation matrices and translation matrices. External parameters are parameters that indicate the position and orientation of the imaging device. Internal parameters include the focal length and optical center of the imaging device. Internal parameters are parameters that indicate the field of view of the imaging device and the size of the imaging sensor.
[0047] Camera parameters are determined by using the correspondence between points in a three-dimensional world coordinate system, obtained from multiple images of a specific pattern captured on a checkerboard or similar object, and points in two dimensions that correspond to those points in the three-dimensional world coordinate system.
[0048] (Functional configuration of a three-dimensional model generation device) Figure 2 is a functional block diagram of the three-dimensional model generation apparatus 12 in the first embodiment.
[0049] The receiving unit 101 receives the foreground mask image and the foreground texture image from the imaging device group 15. The receiving unit 101 determines whether or not it has received the foreground mask image and the foreground texture image. The receiving unit 101 may also receive the captured image.
[0050] The re-reception determination unit 102 determines whether the received data is re-received data. Re-received data is not necessarily data received for the first time; for example, it may be data received in the past. Re-received data does not need to be a perfect match with previously received data; for example, it is sufficient if the shooting time and shooting area of the image data included in both data are the same. Furthermore, re-received data may be a foreground texture image that includes the silhouette region of the foreground of the unreceived area, as described later. The foreground texture image here may be a part of the foreground texture image generated from a single captured image that includes at least the silhouette region of the foreground of the unreceived area. For example, the re-reception determination unit 102 may determine whether the received data is re-received data by looking at the retransmission flag included in the newly received data. Alternatively, the re-reception determination unit 102 may determine whether the newly received data is an image from the first time it was received or an image received in the past, based on information such as the shooting time in the received data.
[0051] The re-reception determination unit 102 outputs the received image data to the model generation unit 103 if it is not the most recent image data, including the image data captured in real time.
[0052] The re-reception determination unit 102 outputs the received data to the data update unit 107 in order to update the foreground texture image if the re-reception data includes image data that was received in the past.
[0053] The model generation unit 103 generates a three-dimensional model using the viewing volume cross-eyed method by excluding the region indicated by the non-foreground region extraction mask from the captured image, based on foreground mask images and foreground texture images that indicate silhouette regions received from multiple imaging devices 10. The model generation unit 103 may also generate a three-dimensional model by superimposing the re-received image onto the excluded region (including the unreceived region) based on re-received data of foreground images (e.g., foreground texture images) included in the region excluded by the non-foreground region extraction mask (including the unreceived region). The model generation unit 103 outputs the generated three-dimensional model, the received foreground mask image, and the foreground texture image to the storage device 13.
[0054] The related data holding unit 104 receives and holds camera parameters from multiple imaging devices 10 and non-foreground region extraction masks from the control device 11.
[0055] The unreceived area calculation unit 105 projects the three-dimensional foreground model generated by the model generation unit 103 onto either a foreground mask image or a foreground texture image generated from images captured by multiple imaging devices 10. The unreceived area calculation unit 105 calculates and identifies the unreceived area, which includes the foreground region included in the captured image that is not included in the foreground mask image or foreground texture image. The region not included in the foreground mask image or foreground texture image may be the foreground silhouette region within the region indicated by the non-foreground region extraction mask. The unreceived area calculation unit 105 may calculate the region where the foreground silhouette region and the non-foreground region extraction mask overlap in the image onto which the three-dimensional model is projected as the unreceived area.
[0056] The retransmission request unit 106 sends retransmission request data to the control device 11 requesting the retransmission of the re-received image, which includes a foreground texture image that includes the silhouette region of the foreground of the unreceived area. The retransmission request unit 106 may send retransmission request data to the control device 11, which includes information on the unreceived area calculated by the unreceived area calculation unit 105, identification information of the imaging device 10, and the time of shooting.
[0057] The data update unit 107 updates the foreground texture image stored in the storage device 13 using the retransmitted data (also called re-received data) based on the retransmission request data. Specifically, based on the imaging device 10 and shooting time indicated by the re-received data, the data update unit 107 superimposes the re-received image, including the retransmitted foreground texture image, onto the unreceived area of the corresponding foreground texture image read from the storage device 13, and writes the superimposed foreground texture image to the storage device 13 to update it. The data update unit 107 may also read and update a foreground texture image from the storage device 13 with the same shooting time as indicated by the re-received data.
[0058] The request determination unit 120 determines whether or not there are any requests stored in the queue.
[0059] If there is requested data in the queue, the transmission acquisition unit 121 acquires the current transmission data of the imaging device group 15 from the control device 11.
[0060] The transmission amount determination unit 122 determines whether the transmission amount of the transmitted data is below a threshold. The transmission amount here may be the amount of data transmitted between the imaging device 10 and the three-dimensional model generation device 12, or the total amount of data transmitted between multiple imaging devices 10 and three-dimensional model generation devices 12.
[0061] The data management unit 123 determines whether the number of data stored in the queue is above a threshold. If the data management unit 123 determines that the number of data stored in the queue is above a threshold, it deletes data so that the number of data stored in the queue falls below the threshold.
[0062] <Functional Configuration of Imaging Device> Figure 3 is a functional block diagram of the imaging device 10 in the first embodiment.
[0063] The imaging unit 111 synchronizes the time within the imaging device group 15 and captures images of the shooting area.
[0064] The foreground extraction unit 112 extracts the foreground silhouette by excluding the non-foreground region extraction mask area from the image captured by the imaging unit 111, and generates a foreground mask image and a foreground texture image.
[0065] The foreground extraction unit 112 sends a foreground mask image showing the extracted foreground silhouette and a foreground texture image cut out from the captured image based on the foreground silhouette to the output unit 116.
[0066] The storage unit 113 stores the images captured by the imaging unit 111.
[0067] The retransmission request receiving unit 114 receives a retransmission request for previously transmitted image data from the control device 11.
[0068] The cropping unit 115 acquires the captured image from the storage unit 113 at the time of capture included in the retransmission request information. The cropping unit 115 crops the acquired image using the image region included in the retransmission request information, and sends the foreground texture image extracted from the cropped image (hereinafter also referred to as the cropped image) to the output unit 116. The cropped image is an example of a re-received image.
[0069] The output unit 116 outputs a foreground mask image and a foreground texture image indicating the region of the foreground silhouette extracted by the foreground extraction unit 112, and a cropped image cut out from the captured image by the cropping unit 115, to the three-dimensional model generation device 12. The output unit 116 may also output the captured image to the three-dimensional model generation device 12.
[0070] <Processing flow in the imaging device> Figure 4 shows a flowchart of the imaging process of the imaging device according to the first embodiment. The imaging process shown in Figure 4 is started, for example, when a user inputs an instruction to start imaging.
[0071] In S1001, the imaging unit 111 of each imaging device 10 captures the shooting area and generates an image while the time is synchronized.
[0072] Figure 5 shows an example of the arrangement of multiple imaging devices in a virtual viewpoint image generation system. In this embodiment, the multiple imaging devices 10 of the imaging device group 15 are arranged around the shooting area 30, which includes a soccer field where the foreground 31 (players) and spectators 32 are located, as shown in Figure 5. Each imaging unit 111 of the multiple imaging devices 10 photographs the shooting area 30 from various angles in a synchronized manner. The imaging unit 111 takes photographs in units called frames, which are equal divisions of one second.
[0073] In step S1002, the foreground extraction unit 112 extracts the foreground region by excluding the region indicated by the non-foreground region extraction mask from the captured image, and generates a foreground mask image and a foreground texture image. Figure 6 shows an example of a captured image taken by the imaging device 10.
[0074] Figure 6(a) shows an example of an image 60 captured by the imaging device 10h. The captured image includes multiple foreground elements 31, which are athletes. Of the multiple foreground elements 31, a portion of the uppermost foreground element 31a overlaps with the background, which consists of spectators 32.
[0075] Figure 6(b) shows an image 61 captured by the imaging device 10h including a non-foreground region extraction mask 40. The non-foreground region extraction mask 40 can be described as a mask of the non-extracted region, which is the region where the foreground is not extracted. The region where the spectators are located (spectator seating area) becomes the non-foreground region extraction mask 40, which is the region where the foreground is not extracted. Of the multiple foregrounds 31, the uppermost foreground 31a overlaps with the non-foreground region extraction mask 40.
[0076] Figure 6(c) shows a foreground mask image 62 obtained by the foreground extraction unit 112 of the imaging device 10h, which extracts the foreground region (also called the silhouette region). The region shown by the non-foreground region extraction mask 40 is an asymmetric region for foreground extraction, so a portion of the foreground 31a within the non-foreground region extraction mask 40 is not included in the foreground mask image 62.
[0077] Figure 6(d) shows the foreground texture image 63 output by the foreground extraction unit 112 of the imaging device 10h. The foreground texture image 63 is an image obtained by extracting the portion of the foreground mask image 62 from the captured image 61. In other words, the foreground texture image 63 can be said to be an image of the silhouette region.
[0078] In S1003, the retransmission request receiving unit 114 determines whether or not it has received a retransmission request from the control device 11. If the retransmission request receiving unit 114 determines that it has not received a retransmission request, it proceeds to S1005. On the other hand, if the retransmission request receiving unit 114 determines that it has received a retransmission request, it proceeds to S1004.
[0079] In S1004, the cropping unit 115 reads the captured image taken at the requested time from the storage unit 113, and generates a cropped image, which is a foreground texture image from which the silhouette region is extracted by cutting out the requested unreceived region indicated by the retransmission request from the captured image. In other words, the cropped image may be a past captured image, and may be a foreground texture image taken at the same time as the image from which the unreceived region was extracted.
[0080] In S1005, the output unit 116 outputs data from at least one of the foreground mask image and foreground texture image extracted in S1002, and the extracted region image cut out in S1004, to the three-dimensional model generation device 12.
[0081] Figure 7 shows an example of the data format of the transmission data that the output unit 116 transmits to the three-dimensional model generation device 12 in S1005.
[0082] The transmitted data includes and is delimited by image data. The transmitted data includes information such as the imaging device identification number, time, type, and retransmission flag. The imaging device identification number may be a number that identifies the imaging device that generated the corresponding image data. The time may be the time when the corresponding image data was captured. The type may indicate the type of the corresponding image data, in this case, one of the following: foreground mask image, foreground texture image, etc. The retransmission flag may indicate whether the corresponding image data is retransmitted data or real-time data.
[0083] Transmission data 2000a and transmission data 2000b represent the data of the foreground mask image and foreground texture image processed from images captured in real time in S1002. Transmission data 2000c represents the data of the foreground texture image extracted in S1004 from previously captured images stored based on a retransmission request.
[0084] Since transmission data 2000a and transmission data 2000b are data relating to images captured in real time in S1002, the retransmission flag is set to 0.
[0085] Since the transmitted data 2000c is data related to the foreground texture image (extracted image) extracted from past captured images accumulated in S1004, the retransmission flag is set to 1.
[0086] In the flowchart of the process in Figure 4, image extraction in response to a retransmission request and the output of the extracted image are performed in sync with the capture, but the processing method is not limited to this. For example, the processing method may involve performing image extraction and output asynchronously with real-time capture.
[0087] <Processing flow in a 3D model generation device> Figure 8 is a flowchart showing the processing of the three-dimensional model generation device 12 of this embodiment.
[0088] Figure 8A is a flowchart showing the data reception process of the three-dimensional model generation device 12 of the first embodiment.
[0089] In S1101, the related data holding unit 104 receives and holds pre-generated camera parameters and non-foreground region extraction masks for multiple imaging devices 10 from the control device 11.
[0090] In S1102, the receiving unit 101 determines whether or not it has received a transmission termination notification from the control device 11. If the receiving unit 101 determines that the transmission has ended, it terminates the receiving process of the three-dimensional model generation device 12. If the receiving unit 101 determines that the transmission has not ended, it proceeds to S1103.
[0091] In S1103, the receiving unit 101 determines whether or not there is data received from the imaging device group 15. The received data may be transmission data including, for example, a foreground mask image, a foreground texture image, and a cropped image (a past foreground texture image). If the receiving unit 101 determines that there is data received, it proceeds to S1104. On the other hand, if the receiving unit 101 determines that there is no data received, it returns to S1102.
[0092] In S1104, the re-reception determination unit 102 determines whether the received data is re-received data based on the retransmission flag included in the received transmission data. If the re-reception determination unit 102 does not determine that the received transmission data is re-received data, the process proceeds to S1105. On the other hand, if the re-reception determination unit 102 determines that the received data is re-received data, the process proceeds to S1111.
[0093] In S1105, the model generation unit 103 generates a three-dimensional model of the foreground using the viewing volume cross-eyed method based on the received foreground mask image and non-foreground region extraction mask. The uppermost foreground 31a shown in Figure 6 overlaps with the non-foreground region extraction mask in the imaging device 10h, so a portion of it is excluded in the foreground mask image. However, the model generation unit 103 can generate a three-dimensional model of the entire foreground 31a using the foreground mask image from another imaging device.
[0094] In S1106, all imaging devices 10 sequentially perform the processing from S1107 onwards.
[0095] In S1107, the unreceived area calculation unit 105 projects the three-dimensional foreground model generated by the model generation unit 103 in S1105 based on camera parameters, etc., onto either a foreground mask image or a foreground texture image generated from images captured by multiple imaging devices 10.
[0096] In S1108, the unreceived area calculation unit 105 determines whether or not there is an unreceived area. The unreceived area calculation unit 105 determines the presence or absence of an unreceived area by checking whether or not there is an overlapping area between the foreground area and the non-foreground area extraction mask onto which the three-dimensional model is projected. If the unreceived area calculation unit 105 determines that there is an unreceived area, the process proceeds to S1108a. On the other hand, if the unreceived area calculation unit 105 determines that there is no unreceived area, the process proceeds to S1110.
[0097] In S1108a, the unreceived area calculation unit 105 calculates the area where the foreground area and the non-foreground area extraction mask overlap as the unreceived area.
[0098] Figure 9 shows the generated three-dimensional model projected onto the image of the imaging device 10h. Here, the image may be the captured image, the foreground mask image, or the foreground texture image. The unreceived area calculation unit 105 calculates the unreceived area 50 based on the overlapping area between the non-foreground area extraction mask 40 and the area onto which the three-dimensional model of the foreground 31 is projected (in this case, the upper area of the foreground 31a). In this embodiment, the unreceived area calculation unit 105 calculates the unreceived area 50 as the bounding rectangle of the overlapping area.
[0099] In S1109, the retransmission request unit 106 stores retransmission request data for the foreground texture image of the unreceived region calculated by the target imaging device in a queue. The retransmission request data includes data such as the unreceived region calculated by the unreceived region calculation unit 105, the identification information of the imaging device, and the time of capture.
[0100] In S1110, once all imaging devices 10 have completed processing, the process returns to S1102.
[0101] In S1111, the data update unit 107 superimposes the re-received foreground texture image (which is also the cropped image) onto the foreground texture image read from the storage device 13, based on the source imaging device and shooting time of the re-transmitted data (data of the re-received image). The data update unit 107 then writes the superimposed foreground texture image to the storage device 13 and updates it. Here, the shooting time of the superimposed foreground texture image may be the same as the shooting time of the re-received cropped image. After processing in S1111, the process returns to S1102.
[0102] Figure 8B is a flowchart showing the data retransmission request processing of the three-dimensional model generation device 12 of the first embodiment.
[0103] In S1121, the receiving unit 101 determines whether or not it is receiving data from the imaging device group 15. If the receiving unit 101 determines that it is receiving data, it proceeds to S1122. On the other hand, if the receiving unit 101 determines that it is not receiving data, it proceeds to S1128.
[0104] In S1122, the request determination unit 120 determines whether or not there is retransmission request data stored in the queue. The retransmission request data stored in the queue may be the request for retransmission and the data necessary for retransmission. If the request determination unit 120 determines that there is no retransmission request data in the queue, it returns to S1121. On the other hand, if the request determination unit 120 determines that there is retransmission request data in the queue, it proceeds to S1123.
[0105] In S1123, the transmission acquisition unit 121 acquires the current transmission amount of the imaging device group 15 from the control device 11.
[0106] In S1124, the transmission volume determination unit 122 determines whether the transmission volume of the transmitted data is below a threshold. Whether the transmission volume is below a threshold is just one example of a condition. If the transmission volume determination unit 122 determines that the transmission volume is greater than the threshold, it returns to S1123 without requesting retransmission in order to prioritize real-time processing. On the other hand, if the transmission volume determination unit 122 determines that the transmission volume is below the threshold, it proceeds to S1125.
[0107] In S1125, the retransmission request unit 106, since the transmission amount is below the threshold and there is sufficient transmission bandwidth, retrieves retransmission request data from the queue and requests the imaging device group 15, via the control device 11, to immediately retransmit the foreground texture image of the unreceived region at the requested time. In this embodiment, the retransmission request unit 106 requests immediate retransmission of the unreceived region, but it is not limited to this and may also request retransmission at a specified time (timing) a certain period of time after the current time.
[0108] In S1126, the data management unit 123 determines whether the number of data stored in the queue is greater than or equal to a threshold. If the data management unit 123 determines that the number of data stored in the queue is greater than or equal to a threshold, it proceeds to S1127. On the other hand, if the data management unit 123 determines that the number of data stored in the queue is less than a threshold, it returns to the process in S1121.
[0109] In S1127, the data management unit 123 deletes data so that the number of data items stored in the queue falls below a threshold, and then returns to S1121.
[0110] In this embodiment, the data management unit 123 reduces the number of data points below a threshold by deleting data stored in the queue, but the management of the number of data points is not limited to this. The data management unit 123 may also prevent the number of data points stored in the queue from becoming excessively large by stopping transmission or other means.
[0111] If the receiving unit 101 determines in S1121 that it is not currently receiving data, the request determination unit 120 determines in S1128 whether or not there is retransmission request data in the queue. If the request determination unit 120 determines that there is retransmission request data in the queue, it proceeds to S1123. If the request determination unit 120 determines that there is no retransmission request data in the queue, it terminates the retransmission request process.
[0112] In this embodiment, the data retransmitted from the imaging device is a foreground texture image, but the data to be retransmitted is not limited to this. For example, the data to be retransmitted may be image data including a foreground mask image, obtained by extracting the foreground from the captured image for which retransmission was requested.
[0113] In retransmission of the foreground mask image, foreground extraction is limited to the circumscribing rectangle of the foreground that overlaps with the non-foreground region extraction mask. Therefore, the foreground mask image does not become so large as to affect transmission, and the foreground model can be made more accurate with the new foreground mask image.
[0114] Furthermore, although this embodiment describes the enhancement of virtual viewpoint image quality through real-time processing, it is not limited to this, and the method may also be applied to non-real-time processing of virtual viewpoint image generation to enhance the quality of virtual viewpoint images.
[0115] As described above, this embodiment projects the generated three-dimensional model of the foreground onto either the foreground mask image or the foreground texture image of the imaging device 10, and retransmits the cropped image (foreground texture image) which includes the silhouette of the unreceived area where foreground extraction is difficult, as re-received data. As a result, this embodiment can generate a three-dimensional model that includes the unreceived area using the cropped image, thereby suppressing differences in resolution and color in the unreceived area where foreground extraction is difficult and enabling the generation of a three-dimensional model. Consequently, this embodiment can generate a high-quality three-dimensional model in replay scenes and highlight scenes with low latency, and can improve image quality while suppressing unnaturalness in the virtual viewpoint image using the three-dimensional model.
[0116] In this embodiment, the area where the non-foreground region extraction mask 40 and the area obtained by projecting the three-dimensional model onto either the foreground mask image or the foreground texture image overlap is calculated as the unreceived region, making it possible to calculate the unreceived region easily and accurately.
[0117] This embodiment uses a foreground texture image of the image extracted from the unreceived region as the re-received data. This allows this embodiment to reduce the amount of data in the re-received data.
[0118] In this embodiment, retransmission is requested when the transmission amount falls below a threshold. As a result, this embodiment can appropriately retransmit according to the transmission amount, thereby reducing the load on transmission while achieving the above effects.
[0119] (Second Embodiment) A second embodiment will be described with reference to the drawings. In this embodiment, the process of setting the priority for retransmission requests based on the position of the foreground will be described.
[0120] The effect of this embodiment is that, in order to prioritize the high accuracy of important foreground elements, even under high transmission load conditions, important foreground elements are prioritized and generated as high-quality virtual viewpoint images in replay scenes and highlight scenes.
[0121] <Processing block> Figure 10 is a functional block diagram of the three-dimensional model generation apparatus 12 in the second embodiment. Functions similar to those in the first embodiment are given the same numbers and their descriptions are omitted.
[0122] The importance setting unit 108 sets the importance level for the foreground three-dimensional model generated by the model generation unit 103. The importance setting unit 108 may set the importance level according to the position of the foreground three-dimensional model in the captured image. Specifically, the importance setting unit 108 may set the importance level for each area of the imaging region.
[0123] The unreceived area calculation unit 105 projects the three-dimensional model onto either a foreground mask image or a foreground texture image generated from images captured by multiple imaging devices 10, calculates the area that overlaps with the non-foreground area extraction mask, associates the importance level set for the projected three-dimensional model with it, and outputs it to the retransmission request unit 106.
[0124] The retransmission request unit 106 requests the imaging device group 15 to retransmit the re-received image according to its importance.
[0125] <Processing Flow> Figure 11 is a flowchart showing the data reception process of the three-dimensional model generation device 12 of the second embodiment.
[0126] In S2101, the importance setting unit 108 sets, as an initial setting, region importance information in which importance is assigned to the three-dimensional model for each area of the shooting region.
[0127] Figure 12 shows an example of setting region importance information, which indicates the importance of the area to be photographed. In this example, the area to be photographed is a soccer field. The importance setting unit 108 sets the importance of the three-dimensional model of the high-importance region 33, which is likely to be a highlight scene near the goal, to high importance. The importance setting unit 108 sets the importance of the three-dimensional model of the low-importance region 34, which is not near the goal, to low importance.
[0128] Sections S2102 to S2103 are the same as sections S1102 to S1103 in the first embodiment, so their explanation is omitted.
[0129] In S2104, the importance setting unit 108 generates a three-dimensional model using the viewing volume cross-section method and assigns importance to the three-dimensional model based on the region importance information and the position of the three-dimensional model.
[0130] In the example shown in Figure 12, the importance setting unit 108 sets the three-dimensional model of the foreground 41, which is located in the low importance region 34, to low importance. The importance setting unit 108 also sets the three-dimensional model of the foreground 42, which is located in the high importance region 33, to high importance.
[0131] Sections S2105-S2107 and S2109 are the same as S1106-S1108 and S1110 in the first embodiment, so their explanation is omitted.
[0132] In S2107a, the unreceived area calculation unit 105 calculates the area where the foreground area and the non-foreground area extraction mask overlap as the unreceived area. The unreceived area calculation unit 105 assigns the importance of the three-dimensional model set in S2104 to the foreground area extraction mask and the retransmission request for the area where the three-dimensional model overlaps.
[0133] In S2108, the retransmission request unit 106 stores the retransmission request data for the foreground texture of the unreceived area calculated by the target imaging device in a queue in order of importance.
[0134] In the second embodiment, since requests are stored according to their importance, the foreground texture images of the three-dimensional models with high importance can be updated preferentially. Even in cases with low latency, such as replay scenes, important foreground elements such as the area in front of the goal can be prioritized to generate high-quality three-dimensional models, enabling the generation of high-quality virtual viewpoint images.
[0135] In the second embodiment, importance was assigned according to the position of the foreground, but the method of setting importance is not limited to this. For example, the importance setting unit 108 may assign importance to the foreground three-dimensional model according to the distance between a specific object, such as a ball, and the foreground three-dimensional model.
[0136] As described above, this embodiment makes it possible to accelerate the improvement of accuracy of the important foreground three-dimensional model by setting the priority of retransmission requests based on the foreground position.
[0137] <Information Processing Device> The hardware configuration of the information processing device 200 will be explained using Figure 13. Figure 13 is a block diagram showing the hardware configuration of the information processing device 200. The information processing device 200 is an example consisting of an imaging device 10, a control device 11, a three-dimensional model generation device 12, a storage device 13, and a rendering device 14. Note that the information processing device 200 may consist of only a part of the imaging device 10, the control device 11, the three-dimensional model generation device 12, the storage device 13, and the rendering device 14, and some parts of the configuration may be omitted.
[0138] The information processing device 200 includes a CPU 211, ROM 212, RAM 213, auxiliary storage device 214, display unit 215, operation unit 216, communication interface 217, and bus 218.
[0139] The CPU 211 controls the entire information processing device 200 by reading computer programs and data stored in either the ROM 212 or the auxiliary storage device 214, loading them into the RAM 213, and executing them. For example, the CPU 211 implements the functions of the imaging device 10 and the three-dimensional model generation device 12 shown in Figures 2, 3, and 10 by executing programs. The CPU 211 may also operate as a display control unit that controls the display unit 215 and an operation control unit that controls the operation unit 216. The information processing device 200 may have one or more dedicated hardware components separate from the CPU 211, and at least a portion of the functions and processing performed by the CPU 211 may be executed by the dedicated hardware. Examples of dedicated hardware include ASICs (Application-Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), and DSPs (Digital Signal Processors).
[0140] The information processing device 200 may have other processors such as an MPU (Micro Processing Unit), GPU (Graphics Processing Unit), NPU (Neural Processing Unit), and QPU (Quantum Processing Unit) in place of or in addition to the CPU 211.
[0141] ROM212 may be a non-volatile storage device that stores programs and other data that do not require modification.
[0142] RAM213 may be a memory capable of reading and writing data at a higher speed than ROM212, etc. RAM213 temporarily stores data such as programs supplied from auxiliary storage device 214, parameters necessary for program execution, image data that is the target of program processing, and data supplied from the outside via communication I / F217.
[0143] The auxiliary storage device 214 may be a large-capacity non-volatile storage device such as a hard disk drive or an SSD (Solid State Drive). The auxiliary storage device 214 stores various types of data, such as programs, image data, and audio data.
[0144] The display unit 215 is composed of, for example, a liquid crystal display and LEDs, and displays a GUI (Graphical User Interface) for the user to operate the information processing device 200.
[0145] The control unit 216 includes, for example, a keyboard, mouse, joystick, or touch panel. The control unit 216 receives operations from the user and inputs various instructions to the CPU 211.
[0146] The communication interface 217 is used for communication between the information processing device 200 and external devices. For example, if the information processing device 200 is connected to an external device by a wired connection, the communication cable is connected to the communication interface 217. If the information processing device 200 has a function for wireless communication with an external device, the communication interface 217 is equipped with an antenna.
[0147] Bus 218 connects the various parts of the information processing device 200 to transmit information.
[0148] In this embodiment, the display unit 215 and the operation unit 216 are assumed to be located inside the information processing device 200, but at least one of the display unit 215 and the operation unit 216 may be located outside the information processing device 200 as a separate device.
[0149] (Other variations) In the above-described embodiment of the virtual viewpoint image generation system 1, the multiple imaging devices 10, control devices 11, three-dimensional model generation device 12, storage device 13, and rendering device 14 are configured as separate devices, but the configuration of the virtual viewpoint image generation system 1 is not limited to this. Some or all of the multiple imaging devices 10, control devices 11, three-dimensional model generation device 12, storage device 13, and rendering device 14 may be configured within the same device.
[0150] In the embodiments described above, an example was given in which the re-received data includes a cropped image, which is a foreground texture image of an image obtained by cropping out the unreceived region of the captured image. However, the re-received data is not limited to this. For example, the re-received data may include data that includes a foreground mask image showing the silhouette region of the unreceived region of the captured image. In this case, the three-dimensional model generation device may generate a foreground texture image from the foreground mask image of the re-received data to generate a three-dimensional model. Furthermore, the re-received data may be a captured image obtained by cropping out the unreceived region from the captured image. In this case, the three-dimensional model generation device may generate a foreground mask image and a foreground texture image from the cropped captured image to generate a three-dimensional model.
[0151] In the above-described embodiment, an example was given in which the unreceived area calculation unit 105 calculates the unreceived area by projecting a three-dimensional model onto either the foreground mask image or the foreground texture image, but the method for calculating the unreceived area is not limited to this. For example, the unreceived area calculation unit 105 may project a three-dimensional model onto the captured image and calculate the area where the foreground area onto which the three-dimensional model is projected and the non-foreground area extraction mask overlap as the unreceived area.
[0152] In the above-described embodiment, the three-dimensional model generation device received data generated by the imaging device and related data via a control device, but the method of receiving data is not limited to this. For example, the three-dimensional model generation device may receive data directly from the imaging device.
[0153] (Other embodiments) This disclosure can also be implemented by supplying a program that implements one or more of the functions of the embodiments described above to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. Furthermore, this disclosure can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions.
[0154] The disclosures herein include the following information processing devices, methods performed by the information processing devices, and programs. (Item 1) A receiving means for receiving an image showing the foreground region included in the captured image, A generation means for generating a three-dimensional model of the foreground based on an image showing the foreground, The captured image includes a means for identifying regions that are not included in the received image, among the regions corresponding to the three-dimensional model of the foreground. The receiving means receives a portion of the captured image based on the identified region. An information processing device characterized by the following: (Item 2) The identifying means identifies a region corresponding to the three-dimensional model of the foreground that is not included in the received image by projecting the three-dimensional model of the foreground. The information processing device described in item 1, characterized by the features described herein. (Item 3) A portion of the captured image that is received is generated by cutting out the portion corresponding to the identified region from the captured image. An information processing device according to item 1 or item 2, characterized in that it is an information processing device according to item 1 or item 2. (Item 4) A portion of the captured image that is received includes a portion of the captured image that corresponds to the identified region. An information processing device according to any one of items 1 to 3, characterized by the features described in item 1 to 3. (Item 5) A portion of the captured image received includes a region corresponding to the three-dimensional model of the foreground. An information processing device according to any one of items 1 to 4, characterized in that it is the same as described in item 1 to 4. (Item 6) The system further includes a requesting means for requesting the external device to transmit a portion of the captured image based on the specified region if the amount of data transmitted to and from the external device meets the conditions. An information processing device according to any one of items 1 to 5, characterized by the features described in item 1 to 5. (Item 7) The request means requests the transmission of a portion of the captured image based on the identified region at a specified timing. The information processing device described in item 6, characterized by the features described herein. (Item 8) A setting means for setting importance levels for the aforementioned three-dimensional model of the foreground, A request means for requesting an external device to transmit a portion of the captured image based on the identified region, according to the importance of the aforementioned region. An information processing device according to any one of items 1 to 7, characterized by having the following features. (Item 9) The setting means sets the importance level according to the position of the three-dimensional model in the foreground. The information processing device described in item 8, characterized by the features described herein. (Item 10) The setting means sets the importance level according to the distance between a specific object and the three-dimensional model of the foreground. An information processing device according to item 8 or item 9, characterized in that it is an information processing device. (Item 11) The image showing the foreground region is an image showing the foreground extracted from the captured image, excluding a specific region. An information processing device according to any one of items 1 to 10, characterized by the features described in item 1 to 10. (Item 12) The aforementioned specific area is the spectator seating area. The information processing device described in item 11, characterized by the features described herein. (Item 13) A method performed by an information processing device, The system receives an image showing the foreground region included in the captured image. A three-dimensional model of the foreground is generated based on an image showing the foreground region. In the captured image, identify the region corresponding to the three-dimensional model of the foreground that is not included in the received image. Based on the identified region, a portion of the captured image is received. A method characterized by the following: (Item 14) A program that causes a computer to perform the actions described in item 13.
[0155] This disclosure is not limited to the embodiments described above, and various modifications and variations are possible. [Explanation of Symbols]
[0156] 12... Three-dimensional model generation device, 31, 41, 42... Foreground, 50... Unreceived region, 62... Foreground mask image, 63... Foreground texture image, 103... Model generation unit, 105... Unreceived region calculation unit.
Claims
1. A receiving means for receiving an image showing the foreground region included in the captured image, A generation means for generating a three-dimensional model of the foreground based on an image showing the foreground, The captured image includes a means for identifying regions that are not included in the received image, among the regions corresponding to the three-dimensional model of the foreground. The receiving means receives a portion of the captured image based on the identified region. An information processing device characterized by the following:
2. The identifying means identifies a region corresponding to the three-dimensional model of the foreground that is not included in the received image by projecting the three-dimensional model of the foreground. The information processing apparatus according to feature 1.
3. A portion of the captured image that is received is generated by cutting out the portion corresponding to the identified region from the captured image. The information processing apparatus according to feature 1.
4. A portion of the captured image that is received includes a portion of the captured image that corresponds to the identified region. The information processing apparatus according to feature 1.
5. A portion of the captured image received includes a region corresponding to the three-dimensional model of the foreground. The information processing apparatus according to feature 1.
6. The system further includes a requesting means for requesting the external device to transmit a portion of the captured image based on the specified region if the amount of data transmitted to and from the external device meets the conditions. The information processing apparatus according to feature 1.
7. The request means requests the transmission of a portion of the captured image based on the identified region at a specified timing. The information processing apparatus according to feature 6.
8. A setting means for setting importance levels for the aforementioned three-dimensional model of the foreground, A request means for requesting an external device to transmit a portion of the captured image based on the identified region, according to the importance of the aforementioned region. The information processing apparatus according to claim 1, characterized by having the following features.
9. The setting means sets the importance level according to the position of the three-dimensional model in the foreground. The information processing apparatus according to feature 8.
10. The setting means sets the importance level according to the distance between a specific object and the three-dimensional model of the foreground. The information processing apparatus according to feature 8.
11. The image showing the foreground region is an image showing the foreground extracted from the captured image, excluding a specific region. The information processing apparatus according to feature 1.
12. The aforementioned specific area is the spectator seating area. The information processing apparatus according to feature 11.
13. A method performed by an information processing device, The system receives an image showing the foreground region included in the captured image. A three-dimensional model of the foreground is generated based on an image showing the foreground region. In the captured image, identify the region corresponding to the three-dimensional model of the foreground that is not included in the received image. Based on the identified region, a portion of the captured image is received. A method characterized by the following:
14. A program for causing a computer to perform the method described in claim 13.