Information processing device, information processing method, and program
The information processing system addresses the issue of out-of-focus objects in captured images by using multiple imaging devices to determine depth of field and color information, ensuring high-quality virtual viewpoint image generation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2021-01-19
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies struggle to appropriately process shape data representing the three-dimensional shape of an object when parts of the object, such as a player's expression, are outside the depth of field in captured images obtained by multiple imaging devices.
An information processing system that acquires data from multiple imaging devices with different viewpoints, determines the shooting conditions, generates shape data within the depth of field, and determines the existence and color information of objects based on captured images to generate high-quality virtual viewpoint images.
The system ensures that only images within the depth of field are used for generating and coloring the three-dimensional model, resulting in higher-quality virtual viewpoint images.
Smart Images

Figure 0007853000000001 
Figure 0007853000000002 
Figure 0007853000000003
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to a technology for performing processing related to a three-dimensional model based on captured images of a plurality of imaging devices.
Background Art
[0002] Recently, a technology has been attracting attention in which a plurality of imaging devices are installed at different positions and synchronized shooting is performed from multiple viewpoints, and a virtual viewpoint image representing the view from a specified viewpoint (virtual viewpoint) is generated using the plurality of images obtained by the shooting. When generating a virtual viewpoint image, a video of the shooting target area as seen from the virtual viewpoint is created by obtaining shape data representing the three-dimensional shape of an object (object) such as a person existing in the shooting target area. There are various targets for generating virtual viewpoint images, and for example, a sports event held in a stadium can be mentioned. For example, in the case of soccer, a plurality of imaging devices are arranged so as to surround the periphery of the field. Patent Document 1 discloses a technology for controlling the focus position so that each of a plurality of cameras covers the entire range on the field where players and the like can move within its depth of field.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, for example, there may be a case where a part of an object such as a player's expression is desired to be enlarged and photographed with high resolution. In this case, among the captured images obtained by a plurality of imaging devices, there are also those in which an object appears outside the depth of field. When such a captured image in which an object is out of the depth of field is used for generating shape data representing a three-dimensional shape or the like, appropriate processing cannot be performed.
[0005] The technology disclosed herein was developed in view of the above-mentioned problems, and its purpose is to appropriately process shape data representing the three-dimensional shape of an object. [Means for solving the problem]
[0006] The information processing system related to this disclosure is a system that acquires data from multiple imaging devices with different viewpoints. movie An acquisition means for acquiring parameters that define the shooting conditions for each of the multiple shooting devices, and the multiple Corresponding to a specified time in the video Based on the captured images, The aforementioned captured image includes A generation means for generating shape data that shows the three-dimensional shape of an object, and after the generation of the shape data, based on the shape data and the parameters, within the depth of field of each of the plurality of imaging devices. , corresponding to the shape data generated based on the captured image corresponding to the predetermined time A determination means for determining whether or not the object exists, and among the captured images corresponding to predetermined times in the plurality of videos, obtained from the shooting device in which the determination means has determined that the object exists within the depth of field. Corresponding to a specified time in the video The system is characterized by having a determination means that determines color information corresponding to the object in a virtual viewpoint image based on a captured image. [Effects of the Invention]
[0007] Information processing related to this disclosure system This includes an acquisition means for acquiring multiple captured images obtained from multiple imaging devices with different viewpoints, and parameters defining the shooting conditions for each of the multiple imaging devices, A generation means for generating shape data that shows the three-dimensional shape of an object based on the aforementioned plurality of captured images. , After generating the shape data, based on the shape data and the parameters, Within the depth of field of each of the aforementioned multiple imaging devices The aforementioned object A determination means for determining whether or not the object exists, and based on the captured images obtained from the shooting device in which the determination means has determined that the object exists within the depth of field among the plurality of captured images, the virtual viewpoint image object The invention is characterized by having a determination means for determining color information corresponding to . [Brief explanation of the drawing]
[0008] [Figure 1] A block diagram showing an example of the configuration of an image processing system that generates virtual viewpoint images. [Figure 2] A schematic diagram showing an overhead view of the arrangement of cameras (filming equipment). [Figure 3] A schematic diagram of a player's object viewed from the side. [Figure 4] Block diagram showing the hardware configuration of the information processing device. [Figure 5] Block diagram showing the software configuration of the server that generates virtual viewpoint images. [Figure 6] (a) and (b) are diagrams illustrating visibility determination. [Figure 7] A table showing the results of the visibility assessment. [Figure 8] A flowchart illustrating the process of generating virtual viewpoint images. [Figure 9] A schematic diagram showing the relationship between the camera's angle of view and depth of field. [Figure 10] A table summarizing the results of visibility assessment and depth of field assessment. [Modes for carrying out the invention]
[0009] The embodiments of this disclosure will be described below with reference to the drawings. Note that the following embodiments are not intended to limit the invention, and not all combinations of features described in these embodiments are necessarily essential to the solution. The same components will be denoted by the same reference numerals.
[0010] [Embodiment 1] <Basic System Configuration> Figure 1 is a block diagram showing an example of the configuration of an image processing system for generating virtual viewpoint images according to this embodiment. The image processing system 100 includes a plurality of camera systems 110a to 110t, a switching hub (HUB) 115, a server 116, a database (DB) 117, a control device 118, and a display device 119. In Figure 1, each camera system 110 contains cameras 111a to 111t and camera adapters 112a to 112t, each connected by internal wiring. Adjacent camera systems are interconnected by network cables 113a to 113s. That is, each camera system 110 transmits data via daisy-chain connection using network cables. The switching hub 115 performs routing between each network device. The server 116 is an information processing device that performs processing of captured images transmitted from the camera systems 110, generation of three-dimensional models of objects (subjects), and colorization (rendering) of the three-dimensional models. The server 116 also has a time server function that generates a time synchronization signal for synchronizing the time of this system. Each camera 111 synchronizes with each other with high precision based on a synchronization signal and takes pictures frame by frame. The database 117 is an information processing device that stores the captured images and generated 3D model data processed by the server 116, and sends the stored data to the server 116. The control device 118 is an information processing device that controls each camera system 110 and the server 116. The control device 118 is also used to set up virtual cameras (virtual viewpoints). The display device 119 displays a user interface screen (UI screen) for the user to specify a virtual viewpoint in the control device 118, and displays a UI screen for viewing generated virtual viewpoint images. The display device 119 can be, for example, a television, a computer monitor, or the LCD display unit of a tablet or smartphone; the type of device is not limited.
[0011] The switching hub 115, the camera systems 110a and 110t are connected by network cables 114a and 114b respectively. Similarly, the switching hub 115 and the server 116 are connected by a network cable 114c, and further, the server 116 and the database 117 are connected by a network cable 114d. And, the switching hub 115 and the control device 118 are connected by a network cable 114e, and further, the control device 118 and the display device 119 are connected by a video cable 114f.
[0012] In the example of FIG. 1, the camera systems 110a to 110t are configured in a daisy chain connection, but a star connection in which the switching hub 115 and each camera system 110 are directly connected may also be used. Also, in the example of FIG. 1, the number of camera systems 110 is set to 20, but this is only an example. The actual number of camera systems is determined in consideration of the size of the shooting space, the content of the target event, the number of assumed objects, the desired image quality, etc.
[0013] Here, a rough flow of virtual viewpoint image generation in the image processing system 100 will be described. The image captured by the camera 111a is subjected to image processing such as separating the foreground object (subject) and the background in the camera adapter 112a, and then transmitted through the network cable 113a to the camera adapter 112b of the camera system 110b. Similarly, the camera system 110b combines the image captured by the camera 111b with the captured image received from the camera system 110a and transmits it to the camera system 110c. By continuing such operations, the captured images acquired by each of the cameras 111a to 111t from the camera systems 110a to 110t are transmitted from the camera system 110t to the switching hub 115 via the network cable 114b, and then transmitted to the server 116.
[0014] In this embodiment, the server 116 generates both the three-dimensional model and the virtual viewpoint image, but the system configuration is not limited to this. For example, there may be separate servers for generating the three-dimensional model and for generating the virtual viewpoint image.
[0015] <Camera Arrangement> FIG. 2 is a schematic diagram showing an overhead view of the state in which 20 cameras 111a to 111t in the above-described image processing system 100 are arranged around a soccer field. In this embodiment, the 20 cameras 111a to 111t are divided into a first camera group (first imaging device group) and a second camera group (second imaging device group). The first camera group consists of 10 cameras 111k to 111t that photograph the entire field from a relatively distant position. On the other hand, the second camera group consists of 10 cameras 111a to 111j that photograph a specific area within the field from a relatively close position. And it is assumed that the 10 cameras 111k to 111t belonging to the first camera group with a long shooting distance face the center of the field. Also, it is assumed that the 10 cameras 111a to 111j belonging to the second camera group with a short shooting distance face five different directions each. That is, cameras 111a, 111b, 111c, 111i, 111j face the vicinity of the goal on the left side of the field, and cameras 111d, 111e, 111f, 111g, 111h face the vicinity of the goal on the right side of the field. Generally, there is a correlation between the shooting distance and the depth of the depth of field, and the depth of the depth of field becomes longer as the shooting distance becomes farther. Therefore, the cameras 111k to 111t belonging to the first camera group have a wider (deeper) depth of field than the cameras 111a to 111j belonging to the second camera group. And by ensuring that the shooting areas of the cameras 111a to 111j belonging to the second camera group are photographed from many directions by a sufficient number of cameras, it becomes possible to generate a higher-quality virtual viewpoint image. In this case, it is assumed that the positions of each camera belonging to both camera groups are defined by the coordinate values of three-dimensional coordinates with an arbitrary point on the field as the origin. Note that all the cameras belonging to each camera group may have the same height in terms of the camera group, or may have different heights.
[0016] Each camera 111k to 111t in the first camera group is assumed to have its camera parameters set so that the entire field (the entire three-dimensional space including the height direction) is within its depth of field. In other words, in the images captured by each camera 111k to 111t, players and other players on the field will always be in focus. On the other hand, each camera 111a to 111j in the second camera group is assumed to have its camera parameters set so that a certain range centered around the area in front of the goal on one side of the field assigned to it is within its depth of field. In other words, for each camera 111a to 111j, if players or other players are within a certain range, a higher resolution and higher quality image can be obtained, but the depth of field is narrower (shallower). Therefore, for each camera 111a to 111j, there are areas on the field that are outside the depth of field even within the field of view, and the players in those areas may be out of focus. Figure 3 is a schematic diagram of player 201, located within the center circle in Figure 2, viewed from the direction of arrow 202 (side view). To explain the difference in depth of field between the two camera groups, only camera 111q, belonging to the first camera group, and camera 111f, belonging to the second camera group, are shown. For camera 111f, the area indicated by trapezoid ABCD represents the depth of field when shooting with this camera. In other words, in the case of camera 111f, the area in front of line segment AD and behind line segment BC, as viewed from the camera, is outside the depth of field. As a result, player 201, which is behind line segment BC, will be out of focus or its original color may not be captured in the captured image. Similarly, the depth of field of camera 111q is shown by pentagon EFGHI. In the case of camera 111q, the area in front of line segment EI and behind line segment GH, as viewed from the camera, is outside the depth of field. Note that in Figure 3, line segments AD and EI are called "forward depth of field," and line segments BC and GH are called "backward depth of field." From Figure 3, it can be seen that player 201 is outside the depth of field of camera 111f and within the depth of field of camera 111q.
[0017] In this embodiment, for the sake of explanation, all 10 cameras in the first camera group are assumed to capture images at the same point of focus, and the second camera group is assumed to be divided into groups of 5 cameras each, each capturing images at a different point of focus. However, this is not limited to this configuration. For example, the first camera group, which has a long shooting distance, may also be divided into multiple groups, and each group may capture images at a different point of focus. Alternatively, the second camera group, which has a short shooting distance, may be divided into three or more groups, each capturing images at a different point of focus. Furthermore, the 10 cameras 111a to 111j belonging to the second camera group may be pointed at different positions or different areas, or several of the 10 cameras may be pointed at the same position or the same area. Similarly, the 10 cameras 111k to 111t belonging to the first camera group may be pointed at different positions or different areas, or several of the 10 cameras may be pointed at the same position or the same area. Moreover, three or more camera groups with different shooting distances (i.e., different depths of field) may be provided.
[0018] <Hardware Configuration> Figure 4 is a block diagram showing the hardware configuration of an information processing device, such as a server 116 and a control device 118. The information processing device includes a CPU 211, ROM 212, RAM 213, auxiliary storage device 214, operation unit 215, communication interface 216, and bus 217.
[0019] The CPU 211 controls the entire information processing device using computer programs and data stored in the ROM 212 or RAM 213, thereby realizing each function of the information processing device. The information processing device may also have one or more dedicated hardware components or a GPU (Graphics Processing Unit) separate from the CPU 211. Furthermore, at least a portion of the processing performed by the CPU 211 may be performed by the GPU or dedicated hardware. Examples of dedicated hardware include ASICs (Application-Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), and DSPs (Digital Signal Processors). The ROM 212 stores programs and other data that do not require modification. The RAM 213 temporarily stores programs and data supplied from the auxiliary storage device 214, as well as data supplied externally via the communication interface 217. The auxiliary storage device 214 is, for example, a hard disk drive and stores various data such as image data and volume data. The operation unit 215 has a display device consisting of a liquid crystal display or LEDs, and an input device consisting of a keyboard or mouse, and inputs various user instructions to the CPU 211 via a graphical user interface (GUI). Alternatively, a touch panel that combines the functions of both a display device and an input device may be used. The communication interface 216 is used for communication between the information processing device and external devices. For example, if the information processing device is connected to an external device by a wire, a communication cable is connected to the communication interface 216. If the information processing device has the function of wirelessly communicating with an external device, the communication interface 216 is equipped with an antenna. Bus 217 connects the various parts of the information processing device and transmits information.
[0020] <Software Configuration> Figure 5 is a block diagram showing the software configuration of server 116, which generates virtual viewpoint images based on images captured by the first and second camera groups. Server 116 includes a data acquisition unit 501, a three-dimensional model generation unit 502, a distance estimation unit 503, a visibility determination unit 504, a depth of field determination unit 505, and a virtual viewpoint image generation unit 506. The functions of each unit will be described below.
[0021] The data acquisition unit 501 acquires, via the switching hub 115, parameters (camera parameters) that define the shooting conditions of each camera belonging to the first camera group and the second camera group, such as position, orientation, field of view, focal length, and aperture value, as well as the captured image data obtained from each camera. The data acquisition unit 501 also acquires, via the switching hub 115, information related to the virtual viewpoint set by the control device 118 (specifically, the position and orientation of the virtual viewpoint, field of view, etc. (hereinafter referred to as "virtual viewpoint information")).
[0022] The three-dimensional model generation unit 502 generates three-dimensional models of objects such as players and balls on the field based on the camera parameters of each camera and captured image data received from the data input unit 501. Here, the specific generation procedure will be briefly explained. First, foreground-background separation processing is performed on each captured image to generate a foreground image representing the silhouette of the object. Here, the background subtraction method will be used as the foreground-background separation method. In the background subtraction method, first, frames from different time series in the captured image are compared, and areas with small differences in pixel values are identified as motionless pixels, and a background image is generated using these identified motionless pixels. Then, by comparing the obtained background image with the frame of interest in the captured image, pixels with large differences from the background in that frame are identified as foreground pixels, and a foreground image is generated. The above processing is performed for each captured image. Next, using the multiple foreground images corresponding to each captured image, a three-dimensional model is extracted using the viewing volume cross-eyed method (for example, the Shape from silhouette method). In the visual volume cross-eyed method, the target three-dimensional space is divided into small unit cubes (voxels), and the pixel position of each voxel when it appears in multiple captured images is determined by three-dimensional calculation, and it is determined whether each voxel corresponds to a foreground pixel or not. If a voxel is determined to be a foreground pixel in all captured images, that voxel is identified as a voxel that constitutes an object in the target three-dimensional space. Only the voxels thus identified are kept, and the other voxels are deleted. The final remaining group of voxels (a set of points with 3D coordinates) becomes a model (3D shape data) that represents the 3D shape of the object existing in the target three-dimensional space.
[0023] The distance estimation unit 503 estimates the distance between each point (voxel) constituting the three-dimensional model generated by the three-dimensional model generation unit 502 and the imaging plane of each camera, using camera parameters input from the data acquisition unit 501. Here, the specific estimation procedure is briefly explained. First, the coordinates (world coordinates) of any point in the target three-dimensional model are multiplied by an external matrix representing the position and orientation of the target camera to convert it to the camera coordinate system. The z-value when the position of the target camera is taken as the origin and the direction in which its lens points is taken as positive on the z-axis of the camera coordinate system is the distance when the arbitrary point is viewed from the target camera. This is done for all points in the target three-dimensional model to obtain the distance from each point to the target camera. By performing this process for each camera, distance information indicating the distance from the target three-dimensional model to each camera is obtained.
[0024] The visibility determination unit 504 uses the distance information obtained by the distance estimation unit 503 to determine, on a voxel-by-voxel basis, whether the model representing the three-dimensional shape of each object is visible from each camera. The specific procedure for visibility determination will now be explained with reference to the figures. Figure 6(a) is a schematic view from directly above showing how eight cameras 611 to 618, placed at equal intervals, photograph an object 600, which is a regular octagonal prism with a regular octagonal base and simulates a person. The eight cameras 611 to 618 are equidistant from the center point of the object 600 and at the same height. The eight points 601 to 608 indicate the points (representative points) where the heights of the cameras 611 to 618 are the same at the joints between the sides that make up the object 600. Figure 6(b) is a schematic view from the side of the relationship between camera 613, one of the eight cameras 611 to 618, and the object 600. Object 600 is assumed to be tall enough so that a representative point exists in the line of sight of every camera. Furthermore, the line of sight of each camera 611-618 is towards the center point of object 600, and all cameras are assumed to be positioned horizontally to the ground. Figure 7 is a table showing whether representative points 601-608 are visible (captured) in the images captured by each of the eight cameras 611-618 when photographing object 600. In the table in Figure 7, the intersection of rows and columns contains values of "1" or "0". "1" means that a specific representative point is visible from a specific camera (visibility exists), and "0" means that it is not visible (invisibility does not exist). In actual visibility determination, first, the distance to the point of interest (voxel) of the 3D model of the object to be determined is compared with the distance to the known center coordinates of the object, using the target camera as a reference. Furthermore, if the distance to the point of interest is shorter than the distance to the object's center coordinates (i.e., the point of interest is closer to the target camera), it is visible; otherwise, it is invisible.
[0025] The depth-of-field determination unit 505 determines whether each object in the generated three-dimensional model is within the depth of field of each camera belonging to the first camera group and the second camera group. This determination uses the captured image data and camera parameters of each camera provided by the data acquisition unit 501, and the distance information obtained by the distance estimation unit 503. Specifically, it is as follows: First, the distance from each object to each camera (coordinate values representing the position of the object in the shooting space) is known for each object from the aforementioned distance information. Then, the forward depth of field and backward depth of field are calculated from the camera parameters of the target camera, and it is determined whether the position of the target object in the captured image (e.g., center coordinates) falls between the calculated forward depth of field and backward depth of field. By performing this for each camera for all objects in the captured image, it is possible to determine whether each object is included in the depth of field of each camera.
[0026] The virtual viewpoint image generation unit 506 generates an image representing the view from the virtual viewpoint by coloring (rendering) the three-dimensional model of each object based on the virtual viewpoint information input from the data acquisition unit 501. In doing so, the determination results of the depth of field determination unit 605 and the visibility determination unit 504 are referenced.
[0027] <Generation process of virtual viewpoint images> Figure 8 is a flowchart showing the flow of virtual viewpoint image generation processing in server 116. The series of processes shown in the flowchart of Figure 8 are initiated in response to receiving a user instruction (a signal instructing the generation of a virtual viewpoint image) containing virtual viewpoint information from the control device 118, and are executed frame by frame (if the captured image is a video). In the following explanation, the symbol "S" means step.
[0028] In S801, the data acquisition unit 501 acquires captured image data obtained by the first camera group and the second camera group through synchronized shooting, as well as camera parameters for each camera 111a to 111t belonging to both camera groups.
[0029] In S802, the three-dimensional model generation unit 502 generates three-dimensional models of objects such as players and balls based on the image data captured by the first group of cameras at a distant shooting distance and the camera parameters of each camera belonging to the first group of cameras. At this stage, since images captured by cameras at a distant shooting distance are used, the accuracy of the resulting three-dimensional models is relatively low.
[0030] In S803, the distance estimation unit 503 estimates the distance from each point constituting the three-dimensional model of each object generated in S802 to each camera 111k to 111t belonging to the first camera group, and generates the distance information described above. The subsequent processing from S804 onwards is performed on an object-by-object basis.
[0031] In S804, the depth-of-field determination unit 505 determines, based on the distance estimation results obtained in S803, whether or not an object of interest exists within the depth of field of each camera 111a to 111t belonging to the first and second camera groups. Figure 9 is a schematic diagram showing the relationship between the camera's angle of view and depth of field. For example, in a multi-player sport such as soccer or rugby, as shown in Figure 9, it is highly likely that multiple players (objects A to D in this example) will be within the angle of view of each camera. However, even if multiple players are captured by a camera, it does not necessarily mean that all of them are captured with high image quality and accuracy. In other words, as shown in Figure 9, when multiple players are within the camera's angle of view, there may be a mix of players (objects B and C) that are within the depth of field and players (objects A and D) that are outside the depth of field. This type of image is particularly likely to be obtained with cameras that have a short shooting distance (i.e., a narrow depth of field), such as the second camera group. Therefore, in this step, it is determined on an object-by-object basis whether or not it is included in the depth of field of each camera belonging to the first and second camera groups, based on the forward and backward depth of field obtained from the camera parameters and the result of distance estimation in S803. Now, each camera 111a to 111j in the first camera group has a wide depth of field that covers the entire field, so players and balls on the field will always be within its depth of field. Also, players and balls near either goal will also be within the depth of field of each camera 111k to 111t in the second camera group. And most of the objects outside the field, such as coaches, substitute players, and spectators, will be determined not to be within the depth of field. Then, according to the determination result, the processing is distributed as follows. First, if the object of interest is not within the depth of field of any camera, the process proceeds to S814 to process the next object. If the object of interest is within the depth of field of the first camera group, the process proceeds to S805; if it is within the depth of field of both the first and second camera groups, the process proceeds to S807.
[0032] In S805, the visibility determination unit 504 determines the visibility of the three-dimensional model of the object of interest from the perspective of each camera belonging to the first camera group. If the three-dimensional model of the object of interest is visible, the process proceeds to S806. On the other hand, if it is not visible, rendering of that object is not required, and the process proceeds to S814 to process the next object.
[0033] In S806, the virtual viewpoint image generation unit 506 performs rendering processing to color the three-dimensional model of the object of interest generated in S802, based on the virtual viewpoint information provided by the control device 118. Specifically, for each point in the point cloud constituting the three-dimensional model based on the first camera group that is determined to be visible, rendering is performed to color it using images captured by cameras in the first camera group that include the object of interest within its depth of field. Note that in S806, processing is performed to determine which camera's image will be used to determine the color of the object of interest, and rendering may be performed later. In other words, it is decided that the image used to determine the color of the object of interest will be the image captured by a camera in the first camera group that includes the object of interest within its depth of field.
[0034] In S807, processing is distributed to cameras belonging to the second camera group, which were determined in S804 to have the object of interest within the depth of field, depending on whether certain conditions are met. These conditions include, for example, that the total number of cameras belonging to the second camera group that were determined to have the object of interest within the depth of field is above a certain number. This certain number (threshold) can be set in advance by the user to be the number of captured images sufficient to generate a three-dimensional model in the next S810. In addition to the total number of cameras, conditions such as being able to capture from a certain number of shooting directions (for example, the front, back, right side, and left side of the object) may also be added. If, as a result of the determination, a certain number of cameras in the second camera group were determined to have the object of interest within the depth of field, the process proceeds to S810; otherwise, it proceeds to S808.
[0035] In S808, the visibility determination unit 504 determines the visibility of the three-dimensional model of the object of interest from each camera belonging to the first and second camera groups. If the three-dimensional model of the object of interest is visible, the process proceeds to S809. On the other hand, if it is not visible, rendering is not required for that object, and the process proceeds to S814 to process the next object. Figure 10 is a table summarizing the results of the visibility determination and depth of field determination, assuming that the shape of player 201, located in the center of the field in Figure 2, is the regular prism object 600 shown in Figure 6. Now, assume that the plane containing the line segment connecting representative points 606 and 607 in object 600, which is the three-dimensional model of player 201, is visible from the front of camera 111a. In the table in Figure 10, a value of "1" for the item "Visibility" indicates that it is visible, and a value of "0" indicates that it is not visible. Furthermore, a value of "1" for the "depth of field" item indicates that the object is within the depth of field, while a value of "0" indicates that it is outside the depth of field. By summarizing the visibility and depth of field information from each camera for each point that makes up the three-dimensional model, it becomes possible to easily determine whether each point in the three-dimensional model is visible or not, and whether it is inside or outside the depth of field.
[0036] In S809, the virtual viewpoint image generation unit 506 performs rendering processing to color the three-dimensional model of the object of interest generated in S802, based on the virtual viewpoint information provided by the control device 118. In this step, as in S806, the target is the three-dimensional model based on the first camera group. The difference from S806 is that, for coloring each point determined to be visible, images captured by the cameras of the second camera group, which were determined in S804 to be within the depth of field of the object of interest, are used preferentially. When deciding on the priority use, the table in Figure 10 above can be referred to. Points (voxels) that were determined to be visible only from the cameras of the first camera group will be colored using images captured by the first camera group. By prioritizing the use of images captured by the second camera group, which takes pictures from a position closer to the object and in which the object of interest is within its depth of field, a virtual viewpoint image that more accurately represents the color of the object can be obtained.
[0037] In S810, the three-dimensional model generation unit 502 generates a three-dimensional model of the player or other object based on the image data captured by the second camera group, which photographs the player or other object from a close distance, and the camera parameters of each camera belonging to the second camera group. However, the image data used for generation is the image data from the second camera group that was determined in S804 to be within the depth of field of the object of interest. By using images captured by the second camera group, which photographs the object from a closer position, a more precise three-dimensional model can be obtained.
[0038] In S811, the distance estimation unit 503 estimates the distance from each point constituting the three-dimensional model of the object of interest generated in S810 to each camera belonging to the second camera group, and generates distance information. However, the cameras used for distance estimation are those from among the cameras belonging to the second camera group that were determined in S804 to be within the depth of field of the object of interest.
[0039] In S812, the visibility determination unit 504 determines the visibility of the three-dimensional model of the object of interest generated in S810 from the perspective of each camera belonging to the second camera group. However, as in S811, the cameras used for visibility determination are those belonging to the second camera group that were determined in S804 to be within the depth of field of the object of interest. If the three-dimensional model of the object of interest is visible as a result of the determination, the process proceeds to S813. On the other hand, if it is not visible, rendering of that object is not required, and the process proceeds to S814 to process the next object.
[0040] In S813, the virtual viewpoint image generation unit 506 performs rendering processing to colorize the three-dimensional model of the object of interest generated in S810, based on the virtual viewpoint information provided by the control device 118. Unlike S806 and S809 described above, this step targets a three-dimensional model based on the second camera group. And, as with S809, the images captured by the cameras of the second camera group, which were determined in S804 to be within the depth of field, are preferentially used as the captured images for coloring each point that has been determined to be visible. By prioritizing the use of images captured by the second camera group for coloring, a high-quality virtual viewpoint image is obtained that accurately represents the detailed three-dimensional model with precise colors.
[0041] In S814, it is determined whether processing is complete for all objects. If there are any unprocessed objects, the process returns to S804, the next object of focus is determined, and processing continues. On the other hand, if processing is complete for all objects, this process ends.
[0042] The above describes the flow of the virtual viewpoint image generation process according to this embodiment. In the rendering processes of S806, S809, and S813, the points (voxels) that are actually colored are the points that remain after object occlusion determination from the points that were determined to be visible. In this embodiment, the camera parameters of cameras 111k to 111l, which capture the entire field within their shooting range, were used to determine the position of objects in the target space, but this is not limited to this. For example, the position of objects may be determined by methods other than distance estimation, such as having athletes carry devices equipped with GPS functionality.
[0043] As described above, according to this embodiment, it becomes possible to use only the images captured by the camera in which the object is within its depth of field for the generation and colorization of the three-dimensional model, thereby generating higher-quality virtual viewpoint images.
[0044] (Other examples) This disclosure can also be implemented by supplying a program that implements one or more of the functions of the embodiments described above to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions. [Explanation of Symbols]
[0045] 116 servers 501 Data Acquisition Unit 502 Three-dimensional model generation unit 505 Depth of field determination section 506 Virtual viewpoint image generation unit
Claims
1. An acquisition means for acquiring multiple videos obtained from multiple shooting devices with different viewpoints, and parameters defining the shooting conditions for each of the multiple shooting devices, A generation means for generating shape data that shows the three-dimensional shape of an object included in a captured image based on a captured image corresponding to a predetermined time in the plurality of videos, After generating the shape data, a determination means determines, based on the shape data and the parameters, whether or not the object corresponding to the shape data generated based on the captured image corresponding to the predetermined time exists within the depth of field of each of the plurality of imaging devices. A determination means that determines color information corresponding to the object in a virtual viewpoint image based on the captured image corresponding to the predetermined time in the video acquired from the shooting device in which the determination means has determined that the object is within the depth of field, among the captured images corresponding to predetermined times in the plurality of videos, An information processing system characterized by having the following features.
2. The determination means determines, based on the shape data and the parameters, whether the object is visible from the imaging device in which it has been determined that the object is within the depth of field, The determination means determines the color information corresponding to the object in the virtual viewpoint image based on the captured image corresponding to the predetermined time in the video acquired from the camera that determined the object was visible, among the captured images corresponding to the predetermined time in the plurality of videos. The information processing system according to feature 1.
3. The information processing system according to claim 2, characterized in that the determination means does not perform a process to determine whether or not the object is visible from the imaging device when it is determined that the object does not exist within the depth of field of the imaging device.
4. An acquisition step of acquiring multiple videos obtained from multiple shooting devices with different viewpoints, and parameters that define the shooting conditions for each of the multiple shooting devices, A generation step of generating shape data that shows the three-dimensional shape of an object included in the captured image based on the captured image corresponding to a predetermined time in the plurality of videos, After generating the shape data, a determination step is made to determine whether or not the object corresponding to the shape data generated based on the captured image corresponding to the predetermined time exists within the depth of field of each of the plurality of imaging devices, based on the shape data and the parameters. A determination step in which, based on the captured image corresponding to a predetermined time in the video acquired from the camera that was determined in the determination step to be within the depth of field, among the captured images corresponding to a predetermined time in the plurality of videos, the color information corresponding to the object in the virtual viewpoint image is determined; An information processing method characterized by having the following features.
5. A program for causing a computer to function as an information processing system according to any one of claims 1 to 3.
Citation Information
Patent Citations
Inside tube wall shape measuring device
JP2006064589A
Information processing apparatus, imaging device, imaging system, information processing method and program
JP2016005027A
Imaging system, image processing device, image processing method, and program
JP2018056971A
Control device, image processing system, control method, and program
JP2019161462A
Image processing apparatus, method for controlling image processing apparatus, and program
JP2019197279A