Information processing apparatus, information processing method, and program

The information processing device simplifies 3D reconstruction by estimating camera parameters from detected object parts, reducing the effort needed for calibration and enabling accurate 3D reconstruction with a small number of cameras.

JP2026012023APending Publication Date: 2026-01-23CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025014953
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2025-01-31
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Conventional 3D shape estimation methods require precise calibration of multiple cameras and fail to handle concave shapes or textures that do not match between stereo images, necessitating laborious camera calibration processes, especially when using a small number of cameras.

Method used

An information processing device that acquires images from multiple cameras, detects a predetermined object part, estimates camera parameters using detected positions, updates these parameters, and determines camera parameters through 3D reconstruction, eliminating the need for traditional camera calibration methods.

Benefits of technology

Reduces the effort required for camera calibration, enabling accurate 3D reconstruction using a small number of cameras without the need for fixed patterns or precise initial camera positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012023000001_ABST
    Figure 2026012023000001_ABST
Patent Text Reader

Abstract

To reduce time and effort when performing camera calibration to obtain a three dimensional reconfiguration result of an object.SOLUTION: An information processing apparatus comprising: an acquisition unit configured to acquire a captured image obtained by capturing an image of an object by each of a plurality of image capturing apparatuses; a detection unit configured to detect a position of a predetermined part in the object from the captured image of each of the plurality of image capturing apparatuses; The information processing apparatus includes an estimation unit configured to estimate camera parameters indicating a position and an orientation of each of the plurality of image capturing devices, an update unit configured to update the camera parameters of each of the plurality of image capturing devices by using the estimated camera parameters as initial values, and a determination unit configured to determine the camera parameters of each of the plurality of image capturing devices based on a result of three dimensional reconfiguration of the object based on the updated camera parameters.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to processing based on captured images. [Background technology]

[0002] There are methods for estimating the 3D shape of an object. The estimated 3D shape is used to generate a virtual viewpoint image, which is a 2D image of the object viewed from a virtual viewpoint. Conventionally, the 3D shape of an object has been estimated from a geometric perspective using multi-view stereo or shape-from-X techniques such as visual hull. What can be said about these conventional methods is that they have been proposed as solutions to the inverse problem of the ill-posed problem by mathematically formalizing the projection process from 3D to 2D.

[0003] To estimate high-quality 3D shapes using these conventional methods, a large number of captured images are required as input, necessitating the use of multiple cameras, which must be precisely calibrated. Furthermore, visual hulls cannot handle concave shapes, and multi-viewpoint stereo techniques cannot achieve high accuracy when textures that fail to match between stereo images are input. As such, all conventional methods face the challenges of estimating 3D shapes due to the existence of certain shapes and textures.

[0004] Therefore, with the development of deep learning technology, methods have been proposed for obtaining 3D reconstruction results from an input image in order to output an image of an object viewed from a desired viewpoint.

[0005] Non-Patent Document 1 describes a method for estimating the three-dimensional shape of a trained object contained in an input two-dimensional image by utilizing a model trained in advance by deep learning, limited to a person. However, the method in Non-Patent Document 1 is premised on the assumption that an image of an object captured by a camera at a position where the distance between the camera and the object is the same as that during training is input to the model.

[0006] Non-Patent Document 2 proposes a method for simultaneously optimizing the camera posture and the radiance field. However, the method in Non-Patent Document 2 requires synchronized shooting with many cameras when performing 3D reconstruction in a changing scene with the presence of a moving object.

[0007] Non-Patent Document 3 describes a method called NeRF. According to Non-Patent Document 3, a fully connected deep network is used to represent a scene. According to Non-Patent Document 3, it is possible to reconstruct the three-dimensional shape of an object from images captured by a sparse set of cameras, and then render the object as a two-dimensional image viewed from a specified viewpoint. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] Saito, S, Huang, Z, Natsume, R, Morishima, S, Li, H, Kanazawa, A. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. In: 2019 IEEE / CVF International Conference on Computer Vision (ICCV). 2019. [Non-patent document 2] Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey. BARF: Bundle-Adjusting Neural Radiance Fields. In: 2021 IEEE / CVF International Conference on Computer Vision (ICCV). 2021,

Outdoor Tools3

Outdoor Tools 4

Direct Environment 5

Outdoor Configuration6

[0009] However, the method of Non-Patent Document 3 cannot obtain a proper 3D reconstruction result of the object unless the position and orientation of the camera are known. Therefore, it is necessary to perform camera calibration in advance to obtain the camera parameters.

[0010] The method in Non-Patent Document 4 is a method for determining the pose of unknown cameras in the NeRF learning process even when the number of cameras is around 4 to 6. However, the method in Non-Patent Document 4 requires that all cameras are placed under the condition that they are facing the center of the scene from a known distance.

[0011] In this way, when the camera positions are not known and 3D reconstruction is performed using deep learning from images captured by a small number of cameras, it is difficult to obtain accurate 3D reconstruction results unless the camera calibration is performed appropriately. For this reason, the user must perform camera calibration, such as preparing a fixed pattern and taking images, which requires a lot of effort on the part of the user. [Means for solving the problem]

[0012] The information processing device disclosed herein is characterized by having an acquisition means for acquiring captured images obtained by each of a plurality of imaging devices capturing an object; a detection means for detecting the position of a predetermined part of the object from the captured images of each of the plurality of imaging devices; an estimation means for estimating camera parameters indicating the position and attitude of each of the plurality of imaging devices using the detected positions of the predetermined part; an update means for updating the camera parameters of each of the plurality of imaging devices using the estimated camera parameters as initial values; and a determination means for determining the camera parameters of each of the plurality of imaging devices based on the results of three-dimensional reconstruction of the object based on the updated camera parameters. [Effects of the Invention]

[0013] The techniques disclosed herein can reduce the effort required to perform camera calibration in order to obtain a 3D reconstruction result of an object. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a diagram illustrating a shooting environment. [Figure 2] A diagram showing a camera arrangement where there is no common field of view. [Figure 3] FIG. 1 illustrates an example of a system configuration. [Figure 4] 10 is a flowchart illustrating a process for determining camera parameters. [Figure 5] FIG. 10 is a diagram illustrating a flow of processing for determining camera parameters. [Figure 6] FIG. 10 is a diagram showing an example of skeleton information. [Figure 7] FIG. 1 is a diagram illustrating a shooting environment. [Figure 8] 10 is a flowchart illustrating a process for determining camera parameters. [Figure 9] FIG. 10 is a diagram for explaining a mask. [Figure 10] FIG. 10 is a diagram illustrating a flow of processing for determining camera parameters. [Figure 11] FIG. 10 is a diagram showing an example of an image output from a three-dimensional reconstruction result. [Figure 12] FIG. 1 illustrates an example of a system configuration. [Figure 13] 10 is a flowchart illustrating a process for determining camera parameters. [Figure 14] FIG. 10 is a diagram for explaining the shooting times of a plurality of cameras. [Figure 15] FIG. 10 is a diagram showing estimated joint trajectories. [Figure 16] FIG. 10 is a diagram for explaining displacement of a joint. [Figure 17] FIG. 1 is a diagram illustrating a shooting environment. [Figure 18] 10 is a flowchart illustrating a process for determining camera parameters. [Figure 19] FIG. 1 is a diagram illustrating a shooting environment. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, embodiments of the technology of the present disclosure will be described with reference to the drawings. The following embodiments do not limit the technology of the present disclosure, and not all combinations of features described in the present embodiments are necessarily essential to the solution of the technology of the present disclosure. The configurations of the embodiments may be modified or changed as appropriate depending on the specifications, usage conditions, usage environment, etc. of the device to which the technology of the present disclosure is applied. Furthermore, in the following embodiments, the same or similar configurations are designated by the same reference symbols, and redundant explanations will be omitted.

[0016] First Embodiment In this embodiment, a method is described in which, in a shooting environment where a single person is present, a small number of shooting devices are used in synchronization to shoot the person, and the respective images (images) obtained by the synchronized shooting are used as inputs to perform three-dimensional reconstruction and determine camera parameters.

[0017] 3D reconstruction is a computational process to obtain a 3D reconstruction result.

[0018] The 3D reconstruction result cannot be directly loaded into an external CG tool, but is a model such as NeRF that stores the learning results of the 3D shape of the target object as MLP weights. NeRF is a model that can input arbitrary viewpoint information and extract an image of the target object observed from that viewpoint. The 3D reconstruction result also includes models that, when an arbitrary viewpoint is input, output a 2D image seen from that viewpoint. Some such models may not actually have parameters representing the 3D shape, but because they produce output similar to models that have the learning results of the 3D shape as weights, such models are also included in the 3D reconstruction result. Alternatively, the 3D reconstruction result may be an image (3D shape data) of the 3D shape of the target object in a format that can be viewed by loading it into an external CG tool, such as voxel or mesh data format. In the following, the 3D reconstruction result will be described as a model such as NeRF. Furthermore, in this embodiment, the object to be 3D reconstructed will be described as a single person in the shooting environment.

[0019] FIG. 1(a) is a diagram illustrating a shooting environment in this embodiment. The shooting environment assumed in this embodiment is a location where a single person is performing a performance such as dance, ballet, or gymnastics. For example, a target person 11, who is the object whose three-dimensional shape is to be learned in 3D reconstruction, is a child participating in a recital such as ballet, rhythmic gymnastics, or martial arts. The scenario is assumed to involve family members 10 and 12 of the child taking photographs for recording. In this case, the number of cameras that the family members 10 and 12 are able to prepare and take photographs is likely to be two or three, as shown by cameras 13 and 16 in FIG. 1(a). For this reason, in a shooting environment such as that shown in FIG. 1(a), 3D reconstruction cannot be performed using a method based on images captured by a large number of cameras, as in professional 3D reconstruction. Therefore, this embodiment provides a technology that enables 3D reconstruction of a person's performance itself based on images captured synchronously by a small number of cameras, such as two or three, and then outputs images rendered from various viewpoints.

[0020] The technology required to perform 3D reconstruction using images captured by a small number of cameras, such as two or three, is difficult to achieve from the captured image data alone compared to performing the reconstruction using a sufficiently large number of cameras. For this reason, in this embodiment, 3D reconstruction is performed using deep learning while utilizing prior knowledge about the target person 11. A method of performing 3D reconstruction from a small number of camera viewpoints while utilizing prior knowledge performs 3D reconstruction as faithfully as possible to the observed area for the captured area. On the other hand, 3D reconstruction is performed by inferring the image for the uncaptured area using prior knowledge or information from other frames. Therefore, to obtain good 3D reconstruction results, it is desirable for each camera to capture images in a way that reduces overlapping information so as to capture as much information about the target person 11 as possible, even with a small number of camera viewpoints. In other words, a camera installation environment advantageous for 3D reconstruction is one in which each camera captures images from a position and orientation that reduces the common field of view between the cameras.

[0021] In Fig. 1(a), cameras 13 and 16 capture images of target person 11 on either side. Therefore, camera 16 captures the left side of target person 11, and camera 13 captures the right side of target person 11. In this way, when there are two cameras, the cameras 13 and 16 are positioned so that there is almost no common field of view between them.

[0022] FIG. 2 is a top view of the shooting environment, showing how a target person 20 is being shot with three cameras. FIG. 2 shows an example in which cameras 21, 22, and 23 are positioned to surround the entire periphery of target person 20 so that no area of ​​target person 20 is left unshot. Even when there are three cameras, if the cameras are positioned to be advantageous for 3D reconstruction, it can be seen that there is almost no common field of view. Thus, to ensure proper 3D reconstruction based on images from two or three cameras, the cameras should be positioned so that there is almost no common field of view.

[0023] However, it is difficult to perform camera calibration to obtain camera parameters from images captured by cameras at positions and orientations with little common field of view.

[0024] In camera calibration using natural features, natural features such as RootSIFT features are detected, and the relative positional relationship of the cameras is calculated based on prominent feature points in the scene using SfM (Structure from Motion). It is known that this method fails unless the images are somewhat similar to obtain corresponding points. For this reason, camera calibration fails when there is almost no common field of view between the cameras, as in the shooting environment of this embodiment.

[0025] Another method for performing camera calibration involves photographing a known reference object and determining camera parameters based on the specific positions photographed. The simplest example of a known reference object is one that uses a two-dimensional plane, typically using a known fixed pattern that can be stably detected and printed on the two-dimensional plane. Specifically, the camera to be calibrated photographs a chessboard, a representative example of a fixed pattern. Then, one method performs camera calibration by utilizing the corner points of the chessboard. When performing such camera calibration in a system including multiple cameras, it is necessary to determine the relative positional relationship between the cameras, so it is necessary to simultaneously photograph the corner points of the same chessboard. Therefore, even with this camera calibration method, in order to obtain appropriate camera parameters, it is necessary to photograph the images so that there are many common fields of view between the cameras. As such, performing general camera calibration is difficult in the camera installation environment assumed in this embodiment, which is advantageous for 3D reconstruction.

[0026] Furthermore, the shooting environment in this embodiment is not a studio-like environment where camera calibration can be performed sufficiently in advance, but rather an environment where families are filming a child's ballet recital or the like for recording. Even if camera calibration could be performed using the above-mentioned chessboard method, this would still impose a heavy burden on the family members 10 and 12 who are filming, as they are the photographers, to stop the other performers' performances before filming and have them hold up a chessboard pattern for camera calibration.

[0027] Therefore, in this embodiment, we propose a method that can perform appropriate camera calibration without relying on natural features or capturing fixed patterns. In this embodiment, we estimate the skeleton of the target person 11 from images obtained from each camera, and perform camera calibration by using the position information (skeleton information) of the target person 11's joints obtained as a result. We will refer to this camera calibration as camera calibration 1.

[0028] [Skeleton information errors] FIG. 1(b) shows the theoretical positions of skeleton information 14 estimated from images captured by camera 13 and skeleton information 15 estimated from images captured by camera 16. Ideally, the joint positions of a person estimated from images captured synchronously by each of cameras 13 and 16 should be completely consistent in the same coordinate system (world coordinates). In this way, it is assumed that the skeleton information 14 and 15 estimated from the images captured by cameras 13 and 16 are completely consistent in world coordinates. In this case, it is possible to appropriately calculate camera parameters indicating the relative positional relationship between each of cameras 13 and 16, based on the position coordinates of each joint indicated by this skeleton information, under the assumption that different cameras are observing the same points.

[0029] FIG. 1(c) is a diagram showing actual skeleton information 17 estimated from an image captured by camera 13 and actual skeleton information 18 estimated from an image captured by camera 16. As shown in FIG. 1(c), in reality, it is extremely rare for the skeleton information estimated from each camera 13, 16 to completely match, and the positions and scales of each joint indicated by the skeleton information estimated from each camera 13, 16 are different. The usage scenario assumed in this embodiment is not an environment in which the positions and orientations of multiple cameras are determined in advance, but rather, the camera positions and orientations shown in FIG. 1(c) are determined only in a test environment. Therefore, when estimating a skeleton from images captured by such cameras 13, 16, it is almost impossible to obtain skeleton information at completely matching coordinates as shown in FIG. 1(b).

[0030] Therefore, in this embodiment, camera calibration 2, which will be described later, is further performed. This method allows camera parameters to be appropriately determined even with a small number of shooting viewpoints, eliminates the need for the photographer to shoot a fixed pattern just for camera calibration, and allows camera calibration to be performed using only the shot scene.

[0031] [System Configuration] FIG. 3 is a diagram for explaining the devices constituting the system according to this embodiment and the hardware configuration of each device.

[0032] The system of this embodiment includes an information processing device 300, three capture groups 310, 320, and 330, and a clock generator 340.

[0033] The information processing device 300 is a device that receives the captured images obtained by the image capturing units 312, 322, and 332 of the capture groups 310, 320, and 330 capturing images in synchronization with each other, and performs camera calibration and three-dimensional reconstruction.

[0034] As a simple example, each of the capturing groups 310, 320, and 330 is realized by a photographing device such as a digital camera. In the case of a digital camera, the storage units 311, 321, and 331 are storage units such as memory cards. Although the number of capturing groups 310, 320, and 330 is three, this is just an example, and in a photographing environment such as that shown in FIG. 1, there would be, for example, two capturing groups 310 and 320. Hereinafter, when simply referring to a camera in the embodiments, this refers to a capturing group.

[0035] The clock generator 340 is a device that assigns a shooting time, such as a time code, to the images (frames) captured by the image capturing units 312, 322, and 332 in the capture groups 310, 320, and 330, respectively.

[0036] By referring to the shooting time stamp attached to the received photographed image, the information processing device 300 can synchronize with the received photographed image and record the photographed image. Furthermore, the information processing device 300 can perform camera calibration and three-dimensional reconstruction using the recorded photographed image.

[0037] 3, the clock generator 340 is shown as a single device connected by wire or wirelessly to all of the capturing groups 310, 320, and 330. Alternatively, for example, the clock generator 340 may be multiple clock generators that are synchronized in advance. In this case, the three clock generators 340 are each included in the capturing groups 310, 320, and 330, and the same effect can be achieved by each clock generator 340 embedding the capture time in the captured image.

[0038] Furthermore, shooting by the shooting units 312, 322, and 332 is performed, for example, by one user issuing synchronized shooting commands to each of the capture groups 310, 320, and 330 via a smartphone or the like. Alternatively, as shown in FIG. 1(a), in an environment where one photographer can control each of the cameras 13 and 16 from nearby locations, synchronization can be achieved by issuing a time code, and it is not necessary to strictly match the start and end times of shooting. For this reason, the photographers who manage the capture groups 310, 320, and 330 may issue shooting commands.

[0039] [Hardware configuration of information processing device] The information processing device 300 includes a CPU 301 , a RAM 302 , a ROM 303 , a storage device 304 , an operation unit 305 , and a display unit 306 .

[0040] The CPU 301 executes various processes using computer programs and data stored in the RAM 302 and the ROM 303. As a result, the CPU 301 controls the overall operation of the information processing device 300 and executes or controls various processes.

[0041] RAM 302 has an area for storing computer programs and data loaded from ROM 303 or storage device 304, and an area for storing data received from capture groups 310, 320, and 330. RAM 302 also has a work area used by CPU 301 when executing various processes. In this way, RAM 302 can provide various areas as needed.

[0042] The ROM 303 is a storage unit that stores setting data for the information processing device 300, computer programs and data related to startup, computer programs and data related to basic operations, and the like.

[0043] The storage device 304 is realized by a hard disk drive device or the like. The storage device 304 stores an OS (operating system), computer programs for causing the CPU 301 to execute or control various processes performed by the information processing device 300, or data. The data stored in the storage device 304 also includes data related to a DNN model for performing three-dimensional reconstruction. The computer programs and data stored in the storage device 304 are loaded into the RAM 302 as appropriate under the control of the CPU 301, and become targets for processing by the CPU 301.

[0044] The operation unit 305 is a user interface such as a keyboard, a mouse, or a touch panel, and allows the user to input various instructions to the CPU 301 by operating it.

[0045] The display unit 306 has a screen such as a liquid crystal screen or a touch panel screen, and can display the processing results by the CPU 301 as images, text, etc. The display unit 306 may be a projection device such as a projector that projects images and text. At least one of the display unit 306 and the operation unit 305 may exist as a separate device outside the information processing device 300. The CPU 301 operates as a display control unit that controls the screen display by the display unit 306, and as an operation control unit that controls the operation unit 305.

[0046] The CPU 301, RAM 302, ROM 303, storage device 304, operation unit 305, and display unit 306 are connected to a system bus 307. Note that the configuration of the information processing device 300 is not limited to the configuration shown in FIG.

[0047] The information processing device 300 is a computer device such as a PC (personal computer) including a set of input / output devices, a smartphone, or a tablet terminal device. Alternatively, the information processing device 300 in this embodiment may be an information processing system configured with multiple information processing devices. In other words, the information processing device 300 is considered to include an information processing system.

[0048] [Camera calibration and 3D reconstruction] Fig. 4 is a flowchart for explaining the flow of processing for camera calibration and 3D reconstruction according to this embodiment. The series of steps shown in the flowchart in Fig. 4 are performed by the CPU 301 of the information processing device 300 expanding program code stored in the ROM 303 into the RAM 302 and executing it. In addition, some or all of the functions of the steps in Fig. 4 may be realized by hardware such as an ASIC or electronic circuit. The symbol "S" in the description of each process indicates a step in the flowchart, and this also applies to subsequent flowcharts.

[0049] The flowchart in Figure 4 assumes that a small group of two or three cameras each captures a single person in a synchronized time series over a period of several minutes. To simplify the explanation, in this embodiment, the internal camera parameters of the camera are assumed to be fixed and acquired in advance. Therefore, in the explanation of Figure 4, the camera parameters obtained by camera calibration are assumed to be external camera parameters.

[0050] The flowchart in Figure 4 will be explained assuming that two cameras 13 and 16, which are two capture groups as in Figure 1, capture images of the target person 11. In this embodiment, the minimum number of cameras is two, but three or more cameras may be used. In addition, the explanation will be given assuming that the cameras 13 and 16 capture images while maintaining fixed positions and orientations.

[0051] In S401, the CPU 301 receives a video (image) including the target person 11, which is obtained by synchronously capturing the target person 11 with the cameras 13 and 16. The video is an image made up of a plurality of frames. After the start of capturing, F frames, which are captured images consecutively in the time series, are received.

[0052] In S402, CPU 301 refers to the time code embedded in each received frame. Then, CPU 301 acquires, as input images, the frames captured by camera 13 and the frames captured by camera 16 that have the same time code assigned. Time codes whose difference is within a predetermined value may be processed as the same time code.

[0053] Fig. 5 is an architecture diagram for explaining the processing of the flowchart in Fig. 4. Input image 500, Image 0, is an image captured by camera 16. Input image 501, Image 1, is an image captured by camera 13. Input images 500 and 501 are assumed to be frames received in S401 and assigned the same time code in S402.

[0054] Then, in S402, a person's posture is estimated (skeleton estimation) for each input image that has the same time code. The human body skeleton estimation is performed by estimating the position of each joint in three-dimensional camera coordinates as the position of the target person 11's body parts.

[0055] The skeleton estimation performed in S402 may be a method of directly estimating the three-dimensional coordinate position of each joint of a person from the input image, but in this embodiment, a method of detecting the two-dimensional coordinate position of each joint in the input image will first be described. If the position of each joint in the input image (two-dimensional coordinate plane) can be detected with high accuracy, it is easy to convert the position information of each joint from two-dimensional coordinate to three-dimensional coordinate information based on the keypoint feature values ​​of the detected two-dimensional coordinate position of each joint.

[0056] As a method for detecting the position of each joint in two-dimensional coordinates, for example, there is a Cascaded Pyramid Network (CPN) described in Non-Patent Document 5.

[0057] The method of detecting joint positions by CPN is a method of detecting a person's area by an object detection algorithm and then estimating the posture for each detected person's area. When using a method for estimating the position of each joint in two-dimensional coordinates such as CPN, the likelihood regarding the estimated position of each joint can be calculated. The likelihood can be calculated by any method. For example, if it is a method of integrating multiple likelihood maps that output the position coordinates of multiple joints to obtain the final joint position, the likelihood can be obtained by taking the cumulative maximum value when all the likelihood maps are aligned to the resolution of the input image and overlaid, etc.

[0058] By the above method, input images consecutive in the time series direction of a moving image are received, and the estimation of the position of each joint in two-dimensional coordinates is performed for each of them.

[0059] Assuming that the internal camera parameters of each camera are known and fixed, and that shooting is performed using P cameras, the camera ID of each camera is denoted as p (0 ≤ p < P), and the camera with camera ID p is denoted as camera p. For example, when there are two cameras (P = 2), camera 0 is camera 16 in FIG. 1, and camera 1 is camera 13 in FIG. 1. Also, when the total number of joints of one human body is J, the joint ID of each joint is denoted as j (0 ≤ j < J), and the joint with joint ID j is denoted as joint j. Also, when the same number of F frames of consecutive images synchronized in the time series direction are input from each camera, the frame ID indicating each frame is denoted as f (0 ≤ f < F), and the frame (captured image) with frame ID f is denoted as frame f.

[0060] FIG. 6 is a diagram for explaining the definition of each joint indicated by human body skeleton information. As with the example joints in skeleton information 600, a method for stably detecting each position defined with a meaning in advance will result in a good three-dimensional reconstruction result, as will be described later. An example of the definition of joint j with joint ID j is shown in list 610 in FIG. 6. The number of joints in the human body skeleton shown in FIG. 6 is 17, but this is just an example, and in this embodiment, there is no limitation on the definition of the number of joints. Even if each joint is detected using a human body skeleton definition different from that in FIG. 6, the following process can be performed in the same way.

[0061] Next, the CPU 301 receives the time-series sequence of the estimation results of the position of each joint in two-dimensional coordinates, and converts the position information in two-dimensional coordinates into position information in three-dimensional camera coordinates by performing temporal convolution of fixed-length frames. The method of converting the position information in two-dimensional coordinates into position information in three-dimensional camera coordinates can be performed, for example, by a method such as that described in Non-Patent Document 6.

[0062] In S402, for all frames of the moving image sequence to be reconstructed in three dimensions, the three-dimensional position coordinates of all joints of the target person 11 in each camera coordinate system are determined by the above-described processing flow.

[0063] 5, skeleton estimation 502 indicates skeleton estimation performed in S402 on input image 500, and skeleton information 504 indicates the shape of the human body skeleton indicating the joint positions of the target person 11 obtained as a result of skeleton estimation 502. Skeleton estimation 503 indicates skeleton estimation performed in S402 on input image 501, and skeleton information 505 indicates the shape of the human body skeleton indicating the joint positions of the target person obtained as a result of skeleton estimation 503.

[0064] In S403, CPU 301 executes camera calibration 1, which estimates the camera parameters of each camera p with coarse accuracy. In this embodiment, a small number of cameras are used for camera calibration. Therefore, by executing camera calibration 1, it is possible to estimate the camera positions and orientations with a certain degree of accuracy, albeit with coarse accuracy, and thereby appropriately execute subsequent processing. Camera calibration 1 is executed using images captured by a small number of cameras that have almost no common field of view. Therefore, camera calibration 1 is executed based on the positions of each joint indicated by human body skeleton information 504, 505 obtained by skeleton estimation in S402.

[0065] As already mentioned, in S402, joint positions are not estimated for a trained person in a trained environment, but rather the joint positions of the target person 11 are estimated in a test environment where photography is performed for the first time. For this reason, the position information of each joint obtained by skeleton estimation in S402 contains a certain amount of error.

[0066] In this embodiment, time codes for synchronized shooting are assigned to the images (frames) captured by each of the multiple cameras 13 and 16. Therefore, optimization calculations are performed with the constraint that the skeleton information estimated from images with synchronized time codes must be identical in the world coordinate system. This is expected to update and converge the positions of each joint indicated by the estimated skeleton information to positions closer to the correct answer.

[0067] When there is only one input image frame, there is little information about the target person 11, making it difficult to perform camera calibration with high accuracy. Therefore, the accuracy of camera calibration can be improved by using multiple frames obtained by capturing images over a time series of several minutes. In addition, since the target person 11 performs an action and moves around within the field of view over multiple frames, it is possible to obtain joint positions observed synchronously across multiple cameras in many areas of the captured image.

[0068] Using VideoPose3D of Non-Patent Document 6, if the human body region is three-dimensionally estimated for an input image and joint positions are obtained, the initial skeleton information is individually input and estimated for each camera and each frame. In this case, originally, the overall scale such as the length between each joint of the skeleton of the same person should be constant for all the skeleton information. However, as described above, the skeleton information 504 and 505 obtained from the respective input images 500 and 501 have slightly different estimation results. Here, first, the CPU 301 performs optimization to minimize the variance of the lengths of all joints estimated from the same person photographed in synchronization by each camera. Now, the intrinsic internal camera parameters calibrated in advance for each known camera p (0 ≤ p < P) are represented as K p and denoted. Also, the position and orientation of the camera are represented using R p and t p and the purpose is to estimate this parameter set <R p , t p >.

[0069] Let the position of joint j in the three-dimensional world coordinates estimated from frame f among the images captured by camera p be represented as X j,f . Also, let the position of joint j in the three-dimensional coordinates in the camera coordinate system of camera p be represented as X p j,f . Also, let the likelihood of the position of each corresponding joint j be represented as L p j,f . In this case, the position X p j,f of joint j in the camera coordinate system of camera p when camera p captures an image can be expressed as Equation 1.

[0070]

Equation

[0071] Let the position of joint j on the two-dimensional image captured by camera p be x p j,fIn this case, the position x of the joint j on the 2D image of the camera p is p j,f and the position X of joint j in the original 3D coordinate system p j,f The relationship between p Using Equation 1, Equation 2 can be expressed.

[0072]

number

[0073] As in Equation 1, the direction from a specific joint to a predefined joint point in the world coordinate system is defined as v j,f In the camera coordinate system when photographed by camera p, the direction from a specific joint to a predefined joint point is denoted as v p j,f can be expressed as in Equation 3.

[0074]

number

[0075] Note that the direction v defined here j,f is a convenient orientation vector for defining the orientation of each joint point, and it is assumed that one vector is defined for each joint. For example, if the predefined reference joint point is HEAD (j=10) shown in Figure 6, then v for j=9 (NECK) is p 9,f is a vector that represents the direction from the neck to the head and the length from the neck to the head. If the vector is defined as a vector from each joint to the adjacent joint connected to it, including overlaps, 17 vectors can be obtained from each frame. In this embodiment, 17 vectors v are obtained from frame f of camera p. p j,f The following description will be given assuming that

[0076] The number of input frames from each camera input simultaneously is F, and the number of joints of the target person 11 is J. Assuming that the target person 11 is photographed in all frames, the total number of orientation vectors v j,f that can be used for camera calibration is J×F. Therefore, by transposing both sides of Equation 3 for J×F three-dimensional directions, Equation 4 is obtained.

[0077]

Equation

[0078] In Equation 4, [v p 0,0 …v p J-1,F-1 is denoted as V T , and similarly, [v p …v 0,0 …v J-1,F-1 is denoted as V. As a result, Equation 4 is expressed as Equation 5. T Here, since the number of cameras used in the current shooting is P and the camera ID is p (0≦p<P), it can be expressed as Equation 6.

[0079]

Equation

[0080] At this time, [V

[0081]

Equation

[0082] …V 0 …V P-1 has a rank of 3 and can be expressed as Equation 7 according to singular value decomposition.

[0083]

Equation

[0084] This allows any invertible 3x3 matrix M to be factorized as shown in Equation 8.

[0085]

number

[0086] Here, by selecting M, the camera pose matrix is ​​made into an orthogonal representation and the recovered orientation is normalized. p can be obtained.

[0087] Furthermore, once the camera rotation is obtained, the translation can be estimated by the collinearity and coplanarity constraints, as shown in Non-Patent Document 7. Therefore, t p can be obtained.

[0088] For simplicity, we have described the method as using all J joint points detected in all F frames captured synchronously by all P cameras. As mentioned above, the detected points of each joint are calculated based on the likelihood L p j,f Since there is no need to use unreliable points at the time of detection, for example, the likelihood L p j,f If the joint position is equal to or smaller than a certain threshold, it may be possible not to use the joint position in the above calculation.

[0089] Up to this point, the parameter set for camera p is obtained by performing linear camera calibration in S403. <R p ,t p In Fig. 5, camera calibration 508 is camera calibration 1 in S403.

[0090] Compared to camera calibration using a known fixed pattern, this camera calibration 1 is based on the joint positions of a person obtained by skeleton estimation, so there is a high possibility that each estimated camera parameter contains errors.

[0091] Therefore, camera calibration is performed by further updating the camera parameters through bundle adjustment using the joint positions obtained by skeleton estimation. Furthermore, the camera parameters estimated in S403 are used as initial values ​​to perform 3D reconstruction, and camera calibration is performed including the rendering loss that can be calculated by comparing the resulting rendered image with the actual image. This camera calibration is called camera calibration 2.

[0092] In this way, in this embodiment, after estimating the camera parameters indicating the position and orientation of the camera with coarse accuracy in S403, it is possible to improve the estimation accuracy of the camera parameters by effectively utilizing image information other than the human body skeleton.

[0093] In S404, optimization processing of parameters related to camera calibration 2 and optimization processing of parameters related to 3D reconstruction are simultaneously performed. The processing when performing these is an iterative process for reducing the accumulated error according to the calculation formula described later to a desired value or less.

[0094] The loss for the optimization process is a value based on two types of loss: the loss related to bundle adjustment and the rendering loss related to 3D reconstruction. In other words, the bundle adjustment process and the optimization process for the rendering loss criterion are performed simultaneously. Then, the overall loss is reduced by repeating parameter updates and rendering.

[0095] Fig. 4(b) is a flowchart showing the details of S404. Details of S404 will be explained using the flowchart of Fig. 4(b) and Fig. 5.

[0096] In S411, the CPU 301 acquires, as initial values, the camera parameters derived as a result of the camera calibration 1 in S403 and the parameters indicating the joint positions of the human body derived as a result of S402.

[0097] Alternatively, if the parameters are updated in S417 because the total loss, which will be described later, exceeds a threshold τ, the CPU 301 acquires the updated camera parameters and joint position parameters in S411.

[0098] In S412, CPU 301 converts the position information of each joint in each camera coordinate estimated from each input image of the multiple cameras 13, 16 into position information in world coordinates using the external camera parameters acquired in S411. Then, CPU 301 integrates the skeleton information. That is, a single piece of skeleton information is generated in which the positions of each joint of the object are indicated in three-dimensional world coordinates. This process corresponds to process 509 in FIG. 5.

[0099] For example, for the position of a joint j in world coordinates, the X 0 j,f From Equation 1, X j,f =A is calculated, and X of camera 1 1 j,f From now on, X j,f =B is obtained. In this way, the camera parameter R p and t p The position of joint j in world coordinates, X j,f will result in different values ​​being obtained as A or B.

[0100] For this reason, in this embodiment, skeletal information is integrated by suppressing the phenomenon that the lengths between joints become unstable for each frame. A network is defined that suppresses all of these complex factors at once. Then, X estimated from the images of each camera is j,f is ΔX relative to the true value j,f If the error is included, ΔX j,f The skeleton information is integrated by estimating and correcting the estimated error. j,fis passed to Skeletal Transformation 510 in Fig. 5 in the next step S413 and S414. One way to suppress this error using rule-based processing is to impose a penalty when the lengths of the skeleton joints of the same person differ depending on the observation time or camera.

[0101] In S413, the CPU 301 calculates a skeleton-based loss ε1 that indicates an error related to the joint position on the two-dimensional coordinate system, based on the integrated skeleton information and the camera parameters acquired in S411.

[0102] In this step, we define the skeleton-based loss among the losses to be considered for optimization. We formulate the minimization of the reprojection error as a maximum likelihood estimation problem, and assume that the estimated results of the joint positions on two-dimensional coordinates can be approximated by a normal distribution with a standard deviation σ. By evaluating the difference from the reprojected joint as a loss, the quality of the camera position and posture is reflected in the score. Therefore, the joint positions obtained by reprojecting the joint positions shown in world coordinates onto two-dimensional coordinates are used as the optimization target parameter K p , R p , t p , and we express it as a vector with a hat representing reprojection with respect to X. X is the joint position in world coordinates after integration in S412.

[0103]

number

[0104] In this case, the skeleton-based loss ε1, which indicates the error of the joint position on the two-dimensional coordinate, can be expressed as in Equation 9.

[0105]

number

[0106] Equation 9 converts the likelihood of each joint position into the negative log-likelihood for each detected joint. pj,f It is designed as a cumulative score, which is a multiplication score. p j,f Since joint positions with low values ​​cannot be trusted as detection points for calibration, this has the effect of reducing the influence on the loss accumulation that is the optimization target.

[0107] The joint position X indicated by the integrated skeleton information is position information on a three-dimensional coordinate system. Therefore, when the position of each joint j in three-dimensional coordinates is projected onto the same two-dimensional plane as the input image and compared with the position x of joint j in the original input image, an error occurs if the estimated parameters do not match the true values. For this reason, the loss ε1 resulting from the accuracy of skeleton estimation can be calculated from Equation 9. If R p and t p If values ​​that completely match the true values ​​are estimated as ∈{square root over (x)} and the positions X of the skeleton joints are also estimated as true values, the error will be 0. Therefore, high-quality camera calibration can be achieved by optimizing to minimize the loss ε1 in Equation 9.

[0108] In S414, the CPU 301 performs three-dimensional reconstruction using the camera parameters acquired in S411.

[0109] First, CPU 301 estimates a three-dimensional reconstruction result 506 in a space where voxels in a fixed three-dimensional space are expressed as an observation space having color information and density from input image 500 obtained from camera 16 (camera 0) as shown in Fig. 5. Furthermore, CPU 301 estimates a three-dimensional reconstruction result 507 in a space where voxels in a fixed three-dimensional space are expressed as an observation space having color information and density from input image 501 obtained from camera 13 (camera 1).

[0110] The format of the 3D reconstruction results 506, 507 may be similar to the method disclosed in Non-Patent Document 3. That is, with regard to the 3D reconstruction result of the target person 11 in the Observation Space, once the viewpoint and line of sight direction during learning are determined, the R, G, B, and density of the sampling points on the ray at that time are accumulated in order from the direction closest to the viewpoint. Then, the result accumulated up to the point where the cumulative density becomes 1 is calculated as the rendering result. Once the 3D reconstruction result is obtained using this method, a rendering result can be obtained by providing an arbitrary test viewpoint.

[0111] In this embodiment, unlike Non-Patent Document 3, in order to simplify the problem, when acquiring and rendering 3D reconstruction results, the appearance information of each voxel is defined as completely diffusely reflective, without being retained as a parameter that varies depending on the viewing direction. Therefore, once a specific viewpoint is determined, the 2D image observed at that time is automatically determined regardless of the viewing direction once the voxel position coordinates in the 3D space to be projected onto the 2D plane are determined. Therefore, the color information and density information returned by each voxel are simplified and modeled so that they remain constant. Therefore, knowing only the spatial position X of each 3D reconstruction allows the RGB value and density returned by the corresponding voxel to be obtained and rendering can be performed.

[0112] Note that while we've explained that Observation Space is "a fixed three-dimensional space in which voxels have color information and density," the NeRF representation of the model containing three-dimensional information about the target person from the input image does not need to be expressed in voxel grid format. It can be expressed in a general MLP format such as NeRF. Extracting shape information is easy with voxel grid format. However, if the image seen from any viewpoint can be extracted by querying the observation viewpoint, rendering loss can be calculated and optimized, so it is not necessary to assign R, G, B, and density to the voxel grid.

[0113] Naturally, because 3D reconstruction is performed for a single camera viewpoint, there will be large errors in the 3D reconstruction results 506 and 507 of the Observation Space at the beginning of learning. Therefore, next, Skeletal Transformation 510 is performed to transform the posture of the target person 11 in the 3D reconstruction results 506 and 507 and integrate the 3D reconstruction results 506 and 507.

[0114] Because the 3D reconstruction results 506 and 507 are based on images captured synchronously by the cameras, the postures of the target person 11 in the 3D reconstruction results 506 and 507 are consistent in world coordinates. Therefore, the CPU 301 integrates the postures of the target person 11 in the 3D reconstruction results 506 and 507 by transforming them into a standard pose using Skeletal Transformation 510, which references the single skeleton information integrated in S412. By using the integrated skeleton information, the 3D reconstruction result of the human body in the learning process is converted into a 3D reconstruction result in canonical space using rig weights determined according to the distance from each joint position. As a result, a 3D reconstruction result 511 of the human body in the standard pose can be acquired in canonical space.

[0115] The target person 11 assumes a free pose when photographed, so the pose of the target person 11 differs for each frame. For this reason, by transforming the target person 11 contained in all frames input from all cameras into a standard pose, which is a common pose, it is possible to stably integrate the observation results for the point of the observation target. The standard pose defined in Canonical Space is defined as the pose for this integration. The standard pose can be any pose as long as it can be transformed into a common pose, but the standard pose state of a three-dimensional human shape, called the canonical T-pose, A-pose, or Y-pose, is generally used. The Skeletal Transformation transformation for converting into a standard pose state is called T skelThen, Skeletal Transformation 510 can be expressed as the following Equation 10.

[0116]

number

[0117] In Equation 10, U p is a position coordinate expression that represents the area in which the 3D reconstructed model of the target person is defined. Specifically, it is an expression for the entire area included in the Observation Space exemplified by the 3D reconstruction results 506 and 507 that are the estimation target. By performing skeletal transformation using Equation 10, it is calculated as inverse linear blend skinning that maps points in the observation space to canonical space. W p j represents the blend weight at joint j in the observation space when observed by camera p. This is based on the general method used in human body animation in computer graphics, which is implemented by the surface that represents the human body and the human skeleton that is defined in correspondence with the surface. p It is used as a weight associated with R p j is the rotation matrix at joint j observed by camera p, t p j represents the movement vector. p j , t p j is the world coordinate position of joint j in the integrated skeleton information and the camera parameter R obtained in S411. p , t p are parameters obtained from

[0118] Now, let the blend weight in Canonical Space be W c j Then W p j and Wc j The relationship can be defined by Equation 11.

[0119]

number

[0120] Skeletal Transformation 510 is an image of a voxel grid-type matrix with the same resolution as the X, Y, and Z directions of the 3D reconstruction target space, or a lower resolution version. These store the rig weight parameters (blend weights) used during animation, which are usually weighted by the distance from the skeleton joint position in the Canonical Pose. This holds a parameter space of 17 joints, and these parameter sets are optimized during the learning process. This parameter set converts the 3D reconstruction results of the human body between Observation Space and Canonical Space.

[0121] Specifically, suppose the resolution of the space containing the target person 11 during 3D reconstruction is, for example, 640 x 640 x 640. The weight parameter matrix for the skeleton rig is a lower-resolution version of the space, e.g., 32 x 32 x 32, with 17 matrices, corresponding to the number of joints. When the resolution is restored to the actual resolution, the weight parameters for skeletal transformation of each human body region are simply obtained by trilinear interpolation. Even if the 3D reconstruction of the human body is not written in voxel grid format, it is possible to obtain each RGB and density when a 3D position is input as a query. Therefore, the parameter set for skeletal transformation can be stored in a voxel grid format table.

[0122] Then, as a result of skeletal transformation 510, the two 3D reconstruction results 506 and 507 are integrated. For example, if two cameras are used to capture a single shot synchronously, the two 3D reconstruction results 506 and 507 estimated from the two images by the transformation using Equation 10 become the 3D reconstruction result 511 for the standard pose. In this case, NeRF estimates RGB and density in the X, Y, and Z directions sampled by rays corresponding to pixels for each image. However, the rays are distorted and do not travel in a straight line due to the transformation using Equation 10. However, since each ray only determines the sampling point, they are treated the same. The sampling points defined from each input image are observed in the same world coordinate system, and the estimation values ​​are integrated with each other, allowing optimization to proceed. Here, the word "integration" refers to the process of combining observation results from multiple viewpoints to reconstruct a single scene, as in NeRF, which inputs a large number of images after the original camera calibration.

[0123] In S415, the CPU 301 calculates the rendering loss ε2. As described above, the three-dimensional reconstruction result 511 has been obtained by transformation into canonical space in S414. The CPU 301 performs the inverse transformation of the Skeletal Transformation 510 on this three-dimensional reconstruction result 511 to return it to a three-dimensional reconstruction result in observation space. Specifically, when the correct image is frame f, the inverse transformation transforms the three-dimensional shape information of the target person 11 in canonical space into the posture of the target person 11 in frame f. As a result, a three-dimensional reconstruction result 550 in observation space is obtained, which holds information about the three-dimensional shape of the target person 11 posing as when frame f was captured. The CPU 301 adds the camera parameters acquired in S411 as information indicating the viewpoints of the respective cameras to this three-dimensional reconstruction result 550. <R p ,t p Then, the CPU 301 compares the output image output as a result with the frame f of the camera p, which is the correct image, to calculate the rendering loss ε2.

[0124] For example, in the case of FIG. 5, the posture of the target person 11 is converted into the posture of the input images 500 and 501, and the viewpoint information of camera 0 is <R 0 ,t 0 Then, an output image 551, which is the observation space viewed from the viewpoint of camera 0, is output, and the output image 551 is compared with the input image 500 as the correct image. Similarly, the viewpoint information of camera 1 is added to the three-dimensional reconstruction result 550. <R 1 ,t 1 > is input, and an output image 552, which is the observation space viewed from the viewpoint of camera 1, is output. Then, the output image 552 is compared with the input image 501, which is the ground truth image. This is performed for all F frames, and the rendering loss ε2 is calculated based on the comparison.

[0125] Specifically, if the rendering loss is ε2, ε2 can be defined as Equation 12.

[0126]

number

[0127] In Equation 12, I p f is the correct image (frame f) of input camera p with frame ID f. c represents an MLP that outputs R, G, B, and density when given an input point. F c The points given to are as defined in Equation 10. Since it is assumed that all points are diffusely reflective, Fc returns the same value for points at the same position coordinates regardless of the view given. Γ is a volume renderer. Γ is used to inversely transform from Canonical Space to Observation Space to the orientation of frame f, and to obtain the position and orientation of camera p. <R p ,t p The output image shows the observation space from the viewpoint shown in the figure. <R p ,tp The two-dimensional image (output image) obtained when the view is given a value of > and the two-dimensional image I, which is the correct image for learning. p f The difference between the pixel values ​​of and is calculated. The difference between the estimated volume rendering and the actual observed image is accumulated for all F frames and all P cameras to become ε2.

[0128] In this way, three-dimensional reconstruction of the human body is performed, and the results are converted into regular expressions, and the rendering loss is calculated based on the differential loss between the rendered image and the actual image by integrating each estimation result.

[0129] In S416, the CPU 301 determines whether the loss ε1 and the loss ε2 satisfy the formula 13. That is, it determines whether the loss based on the loss ε1 and the loss ε2 is equal to or less than the threshold τ. Each parameter to be determined by minimization is solved as satisfying the formula 13.

[0130]

number

[0131] λ is a weight that adjusts which of the loss ε1 calculated based on the skeleton of the human body and the loss ε2 calculated based on the rendering result is given more importance. λ may be a fixed value. For example, λ may be set to 0 so that learning is performed using only the loss ε2, or λ may be set to 1 so that learning is performed using only the loss ε1. λ may be used for scheduling that changes the weighting between the early and later stages of learning. For example, in the early stages of learning, learning may be performed based on the estimated skeleton of the human body by starting with a value of λ close to 1, and as learning progresses, learning may be performed so that the priority of error minimization based on the rendering result is increased.

[0132] If the CPU 301 determines in S416 that Expression 13 is not satisfied, the process proceeds to S417. In S417, the CPU 301 updates the parameters. S417 corresponds to camera parameter update processing 541 in FIG. 5. After updating the parameters, the CPU 301 returns the process to S411. Then, the CPU 301 acquires the parameters updated in S411 and executes S412 to S417 using the updated parameters. Then, S411 to S417 are repeated until Expression 13 is satisfied, that is, until the termination condition is reached.

[0133] The updated parameters are, for example, the camera parameters of camera p. <R p ,t p For example, the CPU 301 may calculate the camera parameters determined in the camera calibration 1. <R p ,t p > to update the camera parameters. Also, the position X in world coordinates of joint j in the integrated skeleton information in S412 is updated. Also, when the camera parameters and joint positions are updated, the parameters used in Skeletal Transformation 510 are updated. Therefore, the parameter set (weights) representing information on the three-dimensional shape of the target person 11 in the three-dimensional reconstruction result 511 in canonical space is also updated. Furthermore, since the updated three-dimensional reconstruction result 511 is inversely transformed, the parameter set representing information on the three-dimensional shape of the target person 11 in the three-dimensional reconstruction result 550 in observation space obtained as a result of the inverse transformation is also updated.

[0134] If it is determined in S416 that Expression 13 is satisfied, the CPU 301 ends the learning. Then, the CPU 301 outputs the camera parameters, skeleton information, and 3D reconstruction results 511, 550 when Expression 13 is satisfied. In this way, the CPU 301 can determine the camera parameters of each camera p. In FIG. 5, this corresponds to parameter determination process 542.

[0135] To summarize the processing of S404, for example, if each camera p shoots at 30 fps for one minute, 1800 images will be input from each camera p. In the case of two cameras as shown in Fig. 5, 1800 input images 500 and 1800 input images 501 will be input.

[0136] If three-dimensional reconstructions 516 and 517 are performed using NeRF, which is the method described in Non-Patent Document 3, three-dimensional reconstruction cannot be performed at all immediately after the start of learning, and the three-dimensional reconstruction results 506 and 507 will be something like random, light-colored point clouds.

[0137] At the beginning of learning, 3D reconstruction results 506 and 507 are obtained for 3600 (1800 + 1800) images that are uninterpretable by humans. These 3D reconstruction results are passed through an immature skeletal transformation 510 network. In this way, in canonical space, a 3D reconstruction result 511 is obtained, which contains information on the 3D shape of a person in a standard pose by integrating the appearance information obtained from all training images. Note that in this embodiment, the 3D reconstruction result 511 itself does not need to be in a state that can be viewed by loading it into a CG tool. In this embodiment, the 3D reconstruction result is expressed as a set of weight parameters for an MLP, since it is a model such as NeRF.

[0138] The parameters to be updated during training are the parameter set of the MLP that represents the target person in canonical space, and this is performed based on the results of error minimization. The error calculation is taken from the MLP in observation space and compared with the input image whose true value is known as the training image. The sum of the errors calculated here is the rendering loss.

[0139] For example, the images output from the 3D reconstruction results 506 and 507 in Observation Space do not contain information about the back side of the target person at the beginning of learning, so they cannot render an image that looks like a person. However, when learning is complete, the person in Canonical Space can be accurately acquired as a result of integrating the 360° images of the person. Therefore, by performing an inverse transformation from these using the MLP of Skeletal Transformation, the 3D reconstruction result 550 in Observation Space also reconstructs the back side of the target person, which was not visible in the input image.

[0140] In general NeRF learning, complex learning cannot be performed with a single input image. For this reason, in this embodiment, parameters are first roughly estimated from a single input image. Then, information from a large number of input images is integrated to obtain a 3D reconstruction result 511 in the canonical pose. Then, by referring to the 3D reconstruction result in the canonical pose, reconstruction is performed in the observation space at a desired pose using information such as the pose parameters of the person in the single input image.

[0141] In this embodiment, it is assumed that 3D reconstruction will be performed only on the number of videos of target people that are shot for normal use, so the goal will be achieved if optimization can be performed on all input images and 3D reconstruction can be performed.

[0142] Here, we will also explain a method for obtaining 3D reconstruction results for a pose not included in the input image after training is complete. In this case, instead of extracting a skeleton from the input image in the normal training process, a skeleton with the same length between all joints is given the desired pose and input in the same way. By doing this, it is possible to inversely transform the trained MLP obtained in canonical space and obtain an MLP parameter set for the target person in the desired pose that was not included in the input image in observation space.

[0143] As described above, in this embodiment, Equation 13, which takes into account the sum of losses representing errors defined by Equation 9 and Equation 12, has been described as being used as an equation for optimization. The condition shown in Equation 13 is satisfied by searching for optimal values ​​for parameters related to both the camera parameters and the human body skeleton. Therefore, the accuracy of camera calibration is improved by optimizing the parameters to be optimized in S411 to S417. Furthermore, by performing learning to minimize the sum of losses representing errors defined by Equation 9 and Equation 12, it is possible to ultimately obtain good 3D reconstruction results.

[0144] In other words, when Equation 13 is satisfied, both the 3D reconstruction result 511 in canonical space and the 3D reconstruction result 550 in observation space are in good condition. In skeletal transformation 510, parameters that allow both transformations are learned. Therefore, once learning is complete, it becomes possible to freely transform between them, similar to the posture deformation caused by a skeleton in CG animation. Therefore, the 3D reconstruction results of both observation space and canonical space are good. The 3D reconstruction result 550 in observation space is used to calculate the difference with the 2D input image, and the canonical space is actually the space for updating parameters.

[0145] Furthermore, even if skeleton information representing postures not included in learning is entered into the human body in canonical space, it is still possible to some extent to reconstruct three-dimensional postures that were not input during learning.

[0146] As described above, according to this embodiment, camera parameters can be appropriately determined for a small number of cameras used for 3D reconstruction without the need for a process of photographing the chessboard for camera calibration. Furthermore, according to this embodiment, it is possible to achieve 3D reconstruction using images input from a small number of cameras, and to improve the quality of images rendered from new test viewpoints.

[0147] In the description of this embodiment, the input image has been described as the captured image itself obtained by capturing an image with a camera p. For example, the input image may be an image in which semantic region segmentation is performed on the captured image using a method such as Mask-RCNN to extract only regions related to the human body. Furthermore, although the camera calibration has been described as determining external camera parameters, internal camera parameters may also be determined simultaneously.

[0148] In addition, in the description of this embodiment, a case where F frames obtained by synchronous shooting in a time series direction are input has been described. These F frames may be frames of a moving image input offline after shooting is completed, or may be frames input as part of real-time processing while shooting the moving image. Furthermore, if the F frames are part of a moving image, learning and estimation processing may be performed for each frame.

[0149] Although the present embodiment has been described with respect to 3D reconstruction of a person, the generality is not lost even if the subject is other than a person. For example, if the subject is an animal other than a human, by implementing a method for estimating the skeleton of the animal instead of estimating the skeleton of a person, the same effect can be obtained for various animals such as dogs and cats.

[0150] <Second embodiment> In the first embodiment, the object to be photographed is described as a single person. In this embodiment, a method for camera calibration and 3D reconstruction will be described when photographing multiple people, non-human animals, non-living objects, etc. using multiple synchronized cameras. The number of photographing cameras is assumed to be a small number of 2 to 3, as in the first embodiment. In addition, a method will be described in which camera calibration and 3D reconstruction are simultaneously performed by using photographed images obtained by synchronously photographing a small number of cameras, without requiring the photographer to perform camera calibration work in addition to photographing. This embodiment will be described mainly focusing on the differences from the first embodiment. Portions not specifically mentioned have the same configuration and processing as the first embodiment.

[0151] 7(a) is a diagram showing an example of a shooting environment assumed in this embodiment. For example, it is assumed that a family uses two cameras 700 and 705 to shoot a scene with activity, such as two people 701 and 703 playing with a dog 704 and a ball 702, for recording purposes.

[0152] 8 is a flowchart for explaining the flow of camera calibration and three-dimensional reconstruction processing according to this embodiment, which corresponds to the flowchart of FIG.

[0153] The flowchart in Fig. 8 explains the processing when two cameras 700 and 705, which are two capture groups, capture target objects in sync over a period of several minutes in the time series direction, as shown in Fig. 7(a). In addition, in the explanation of Fig. 8, the camera parameters obtained by camera calibration are assumed to be external camera parameters.

[0154] In S801, the CPU 301 receives video images (video) obtained by cameras 700 and 705 synchronously capturing images of an object to be reconstructed in three dimensions. After the start of capturing, multiple frames, which are consecutive captured images in chronological order, are received as input images containing the object. As described above, the captured scene assumed in this embodiment contains multiple objects. The input images received in S801 are assumed to include people 701 and 703, a dog 704, and a ball 702. In S801, the time codes embedded in each captured image (frame) are referenced, and a set of images captured by different cameras and assigned the same time code is acquired, enabling subsequent processing.

[0155] In S802, the CPU 301 detects the area of ​​each object from an input image containing multiple objects. Specifically, the CPU 301 detects the area indicating the object in the input image by instance segmentation, tracks the same object in consecutive images (frames) in the time series, and assigns the same ID (identifier). This ID is called an instance ID. An object assigned an instance ID is also called an instance. Instance segmentation and tracking can be achieved by a method such as that described in Non-Patent Document 8.

[0156] FIG. 9(a) shows an example of the results of instance segmentation and tracking performed on a specific input image. Generally, instance segmentation distinguishes the regions of individual instances, making it possible to assign individual instance IDs to the regions of each object in an image containing multiple objects. As shown in FIG. 9, person regions 901 and 903 are not only classified with a class name of Person, but are also identified and assigned instance IDs for tracking. That is, person region 901 is assigned the instance ID of Person1, and person region 903 is assigned the instance ID of Person2. Object region 902 is assigned the instance ID of Football1. Dog region 904 is assigned the instance ID of Dog1. In this way, the regions of the instances are distinguished, and objects included in consecutive images in the time series are tracked. Furthermore, tracking allows the boundaries between objects to be distinguished even when the objects overlap and one of them is occluded.

[0157] Therefore, in S802, the CPU 301 can generate a mask indicating the area of ​​each object in the input image, as shown in Fig. 9. Therefore, even if objects overlap, masks are generated with the overlapping objects as separate areas, as shown in Fig. 9(b). Unlike the first embodiment, in this embodiment, multiple objects overlap each other and occlusion occurs. Therefore, by separating each object and using the mask information shown in Fig. 9 for the overlapping areas, it is possible to obtain accurate results.

[0158] The next steps S803 to S805 are loop processes. In S803, an instance ID to be processed is selected from the instance IDs that are consistent in the time series direction obtained in S802, and in S804, skeleton estimation is performed on the instance indicated by the instance ID to be processed in the same manner as in the first embodiment.

[0159] S804 is a step corresponding to the skeleton estimation in S402 in the first embodiment. In S804, the CPU 301 performs skeleton estimation that can be used for camera calibration, using a detection model corresponding to the class indicated by the instance ID of the processing target.

[0160] When the instance indicated by the instance ID to be processed is a person 701 or 703, similar to the first embodiment, a trained model for estimating the skeleton of a human body is used as a detection model to perform skeleton estimation.

[0161] Fig. 7(b) is a diagram for explaining the results of skeleton estimation. Skeleton information 711 and 713 of the person is obtained by estimating the skeleton from the images obtained by capturing Fig. 7(a) using cameras 700 and 705.

[0162] If the instance indicated by the instance ID to be processed is a non-human animal such as a dog 704, CPU 301 performs animal skeleton estimation using the method described in Non-Patent Document 9. When animal skeleton estimation is performed, the resulting skeleton information 714 can be treated in the same way as human skeleton information. In other words, skeleton information 714 can be used for camera calibration 1. Due to differences in skeleton definition resulting from differences in skeleton estimation methods, subsequent processing can be performed even if the number of joints in a human skeleton and the number of joints in an animal skeleton differ.

[0163] For non-living objects such as ball 702 with an instance ID of "Football," similar processing is performed when a non-living rigid body model is determined. For example, if the class is sphere, a sphere model is used. Therefore, if the class is sphere, CPU 301 performs general sphere fitting to determine the position coordinates of the sphere center. Then, based on the camera position and orientation, sphere fitting obtains the ellipse center by performing two-dimensional ellipse fitting on the ball area of ​​multiple images captured synchronously. Rays are then extended from the sensor center toward the ellipse centers of the multiple cameras, and a sphere center is assumed to be at the center of the closest point of each ray. The sphere radius is then fitted so as to best overlap with the two-dimensional ellipse. This allows the position of sphere center 712 as seen from each camera to be estimated. In this way, for non-living objects, a center-like position can be determined as a body part. All sphere centers moving in the time series direction can be used for camera calibration 1 in S805.

[0164] Alternatively, if the subject being photographed includes a person or an animal, skeleton estimation is performed on the person or animal. Therefore, the estimated joint positions can be used to estimate the camera position and orientation, as in the first embodiment. Therefore, if an instance ID indicating a non-living object, such as ball 702, is selected as the processing target in S803, CPU 301 may skip S804, assuming that there is no detection model corresponding to the target class.

[0165] In S806, the CPU 301 estimates initial values ​​of external camera parameters indicating the position and orientation of each camera by camera calibration 1 similar to that in the first embodiment, using the joint positions indicated by the skeleton information.

[0166] In this embodiment, since the input image contains multiple objects, objects frequently occlude other objects. Therefore, as shown in FIG. 9(b), the object regions, such as person regions 905 and 906, dog region 907, and ball region 908, may be detected as overlapping. When an object is occluded in this way, the accuracy of the joint positions related to the occluded region in the skeleton information of the human body and the skeleton information of the animal deteriorates.

[0167] For this reason, it is preferable to perform camera calibration 1 using position information of only those joints of the target instance that are included in the area indicated by the mask of the target instance obtained in instance segmentation in S802. In other words, it is preferable to perform camera calibration 1 so that position information of joints included in occluded areas is not used. For this reason, as described in the first embodiment, estimated likelihoods regarding joint positions may be used.

[0168] S807 is a step corresponding to S404 in FIG. 4. In S807, the CPU 301 performs camera calibration 2, corrects joint positions, and performs three-dimensional reconstruction. The flowchart in FIG. 8(b) is a flowchart for explaining the details of the processing of S807. Steps S811 to S817 are the same processing as steps S411 to S417 in the first embodiment. That is, bundle adjustment is performed using the initial values ​​of the camera parameters of each camera estimated in camera calibration 1, the joint positions of each instance (each object), and the like, and learning is performed to minimize rendering loss in the learning process.

[0169] Fig. 10 is an image diagram of optimization calculation using a space corresponding to the number of instances detected in S802. Unlike Fig. 5, Fig. 10 shows up to the process of integrating the 3D reconstruction results into Canonical Space.

[0170] Unlike the first embodiment, in this embodiment, if there are multiple people being photographed, calculations must be performed for each person. Similar calculations must also be performed for animals and non-living objects other than people. Therefore, the results of skeleton estimation for each instance are integrated to perform camera calibration 1, and the camera parameters obtained as a result become the initial values ​​for camera calibration 2.

[0171] Furthermore, in three-dimensional reconstruction, since an observation space and a canonical space are defined and calculations are performed, a space corresponding to the number of instances detected in S802, that is, the number of objects in the shooting environment, is required.

[0172] Processes 1001 to 1004 indicate processes performed for each instance (object). Note that the process content of process 1003 is omitted from the illustration due to space limitations. In principle, the explanation will be given assuming that there are no differences between processes 1001 to 1004 for each instance. As shown in FIG. 10, pose correction performed in S812 and conversion to canonical space by executing skeletal transformation performed in S814 are performed for each instance.

[0173] In FIG. 10, lines connect the skeleton information obtained from each of the four instances with camera calibration 1010, which indicates camera calibration 1. This indicates that camera calibration 1 is performed by referencing the position information of the joints of each instance. The initial values ​​of the camera parameters obtained as a result of camera calibration 1 affect pose correction and other processes in all processes 1001 to 1004. As a result, the 3D shape result expressed in the canonical space of each instance is updated. Furthermore, although not shown in FIG. 10, in the processing of each instance, an inverse transformation from canonical space to observation space is performed, and a 3D reconstruction result expressed in observation space is obtained. A rendering loss ε2, which indicates the error between the output image output from the 3D reconstruction result expressed in observation space and the input image that is the ground truth image, is calculated for each of the four instances.

[0174] As described above, in this embodiment, it is necessary to calculate the rendering loss ε2 for each instance. When calculating the rendering loss ε2 in one of processes 1001 to 1004, the rendering loss ε2 may be calculated using a mask indicating the region of the instance detected in S802. Specifically, assume that the target instance is a Person 1 instance and the ground truth image is frame f. In this case, the rendering loss ε2 may be calculated by comparing the pixel value difference between only the region corresponding to the Person 1 mask detected in frame f, between the output image from the 3D reconstruction result of the observation space and frame f. In this way, by comparing only the region of the target instance's mask in processes 1001 to 1004 for each instance, the rendering loss ε2 can be calculated without using information about the region where the target instance is occluded.

[0175] Similarly, in the first embodiment, a mask indicating the object area may be generated by excluding the background area from the captured image, and the rendering loss ε2 may be calculated using the mask.

[0176] Furthermore, when comparing the total loss with the threshold value τ in S816, the total loss may be calculated by summing up the losses ε1 and ε2 calculated for each instance in steps 1001 to 1004. In this case, a weight based on the instance may be determined, and the sum of the losses ε1 and ε2 taking the weight into consideration may be calculated.

[0177] As described above, in this embodiment, camera calibration is performed from images captured synchronously using a small number of cameras, such as two or three, and therefore skeleton estimation of people and animals and position estimation of general objects are described. In this way, this embodiment stably determines the posture of people and the position of objects, individually recognizes and tracks each object, and independently performs 3D reconstruction. As a result, it is possible to provide applications that can display images that could not be achieved by methods such as NeRF in Non-Patent Document 3, which only reference information from multiple cameras observed at the same time and reconstruct the entire scene in 3D.

[0178] FIG. 11A is a diagram illustrating an example of a screen 1100 displayed on the display unit 306 by the CPU 301. The screen 1100 displays an image obtained by rendering the 3D reconstruction results for each instance (object) obtained in S807 to create an image viewed from an arbitrary test viewpoint. The screen 1100 in FIG. 11A depicts a dog 1101 and a person 1102, which are the targets of 3D reconstruction, with the dog 1101 and the person 1102 closely spaced apart. If a user wants to view an image of only a particular object, in this embodiment, the user can select the object to be rendered using a pointer 1103 shown in FIG. 11B. In this embodiment, an instance ID is assigned to the object to be 3D reconstructed in S802, and 3D reconstruction is performed for each instance in S807, resulting in a 3D reconstruction result for each object. This makes it possible to accept the selection of the object to be rendered from the user.

[0179] FIG. 11(c) shows an example of a screen 1104 when a dog 1101 is selected and rendering is performed using the 3D reconstruction results for the dog 1101. FIG. 11(d) shows an example of a screen 1106 when a person 1102 is selected and rendering is performed. As shown in screen 1100 of FIG. 11(a), when the dog 1101 and person 1102 are in close contact, the rendering results for regions 1105 and 1107 are indeterminate when using the NeRF method of Non-Patent Document 3 or the classical Visual Hull method. Therefore, when attempting to generate the image shown in the screen of FIG. 11(c) or 11(d), the NeRF method of Non-Patent Document 3 or the classical method generates an image of degraded quality. On the other hand, according to this embodiment, 3D reconstruction results obtained by observation across multiple frames are integrated. Therefore, even areas not observed by the camera when selecting the object to be rendered can be reproduced using information observed at other times. Therefore, according to this embodiment, it is possible to display an image of an object with minimal degradation. This type of use is useful, for example, when shooting a sports scene such as soccer or rugby, and you want to observe the movements of a specific person in detail.

[0180] In the example of Figure 7(a), the joint positions indicated by the skeleton information of each object are determined with high accuracy. Therefore, after performing 3D reconstruction, a virtual viewpoint can be set at the joint position of the dog's 704's top head so that the viewpoint is from the dog 704, and virtual viewpoint video obtained by rendering based on that virtual viewpoint can be output. This eliminates the need to attach a real camera to the pet as a real object or to attach a marker such as a chessboard to the pet, thereby reducing the user's effort required for shooting. Similarly, it is also possible to output virtual viewpoint video from the viewpoints of two children, persons 701 and 703, and virtual viewpoint video looking down from a ball.

[0181] As described above, according to this embodiment, even if an image is captured in an environment where an object is frequently occluded by another object, it is possible to correctly extract an object from the image and reconstruct it in 3D, thereby providing images rendered from various viewpoints.

[0182] <Third embodiment> In the first and second embodiments, the subject is described as being photographed using multiple synchronized cameras. In this embodiment, a method for processing camera calibration and 3D reconstruction is described, even if a small number of cameras, such as two or three, photographs are taken without the photographing times being strictly synchronized. This embodiment will be described focusing on the differences from the first embodiment. Portions not specifically mentioned have the same configuration and processing as the first embodiment.

[0183] As in the second embodiment, there may be multiple subjects to be photographed, but for simplicity, as in the first embodiment, a single subject 11 will be photographed. Therefore, the settings of the photographing environment and the like in this embodiment will be described as being the same as those in the first embodiment, except that the two cameras do not photograph synchronously.

[0184] [System Configuration] FIG. 12 is a diagram illustrating the hardware configuration of each device and the devices that make up the system according to this embodiment. The same components as those in FIG. 3 are designated by the same reference numerals. The difference from FIG. 3 is that the clock generator 340 is not included in FIG. 12. Therefore, the image capturing units 312, 322, and 332 in the capture groups 310, 320, and 330 of this embodiment capture images asynchronously. The information processing device 300 receives each captured image obtained as a result of the capture. The information processing device 300 of this embodiment synchronizes with the captured images after receiving them by referencing the capture time stamps assigned to the received images, and stores the captured images. Furthermore, the information processing device 300 of this embodiment performs camera calibration and 3D reconstruction using the stored captured images.

[0185] While synchronized shooting by the camera units 312, 322, and 332 is not required, the shooting times of the camera units 312, 322, and 332 must overlap in terms of the time they are shooting the same subject. This assumes a case in which three camera operators operating capture groups 310, 320, and 330 press the shooting start switches on capture groups 310, 320, and 330 at a signal, causing the camera units 312, 322, and 332 to start shooting, respectively. In addition, in the case shown in FIG. 1(a), it is assumed that camera operator 10 of camera 13 and camera operator 12 of camera 16 start shooting when target person 11 starts performing and end shooting when the performance ends. Therefore, although the shooting start and end times are roughly the same across the multiple cameras, the posture of target person 11 cannot be captured at exactly the same time as he or she moves around.

[0186] [Camera calibration and 3D reconstruction] Fig. 13 is a flowchart for explaining the flow of processing for camera calibration and 3D reconstruction according to this embodiment. A series of steps shown in the flowchart in Fig. 13 is performed by the CPU 301 of the information processing device 300 of this embodiment by loading program code stored in the ROM 303 into the RAM 302 and executing it. Furthermore, some or all of the functions of the steps in Fig. 13 may be realized by hardware such as an ASIC or an electronic circuit.

[0187] For simplicity of explanation, in this embodiment, the internal camera parameters of the camera are assumed to be fixed and acquired in advance. Therefore, in the explanation of Fig. 13, the camera parameters obtained by camera calibration are assumed to be external camera parameters.

[0188] The flowchart in FIG. 13 assumes that two or three small capturing groups each capture one person over a period of several minutes in the chronological order. Therefore, the flowchart in FIG. 13 will be described assuming that two capturing groups, consisting of two cameras 13 and 16, capture the target person 11, as in FIG. 1. While the minimum number of cameras in this embodiment is two, it may be three or more. The description will also assume that the cameras 13 and 16 capture images while maintaining a fixed position and orientation. Furthermore, in this embodiment, as described above, strict synchronization of the capture times between the cameras 13 and 16 is not required.

[0189] In S1301, the CPU 301 receives a video (image) including the target person 11 obtained by the cameras 13 and 16 capturing the target person 11. The video is an image (image sequence) made up of a plurality of frames. When capturing images of a target such as the target person 11 continuously in a chronological order, the multiple cameras 13 and 16 must capture the images at the same time. However, unlike the first embodiment, the times at which the multiple cameras 13 and 16 capture the target, such as the timing of the shutters for continuous capture, do not need to be the same.

[0190] Fig. 14 is a diagram that schematically illustrates a situation in which two cameras are simultaneously capturing images but the timing at which they capture the images does not match. The upper row 1403 of Fig. 14 shows images (frames) captured by camera 13 of Fig. 1 at each time, and the lower row 1404 shows images (frames) captured by camera 16 of Fig. 1 at each time. The frame rates (fps) of the cameras 13 and 16 are different, and the shutters of cameras 13 and 16 are released at different times. For this reason, Fig. 14 illustrates that the posture of target person 11 when captured by camera 13 is different from the posture of target person 11 when captured by camera 16.

[0191] Similar to the first embodiment, the camera ID when shooting with the P cameras is distinguished as p (0 ≤ p < P), and for the frame f (0 ≤ f < F) of the captured image, basically, it is determined which camera p captured the frame. In FIG. 14, the correspondence between the camera p and the frame f is shown. The camera ID of camera 16 is p = 0, and the target person 11 captured by camera 16 is at time f 0,1 , f 0,2 , f 0,3 . It is shown that the shooting is being performed. The camera ID of camera 13 is p = 1, and the target person 11 captured by camera 13 is at time f 1,1 , f 1,2 , f 1,3 , f 1,4 . It is shown that the shooting is being performed. Thus, since cameras 13 and 16 are not shooting at the same time, when representing the frame f, the camera ID is given before the frame ID representing the shooting order, and it is represented as f p,f . Even if the shooting timings of different cameras accidentally coincide in the actual environment, there is no need to change the processing flow of this embodiment.

[0192] In S1301, after the start of shooting by cameras 13 and 16, the CPU 301 receives a group of continuously captured images in the time series direction from each of the cameras 13 and 16. In this embodiment, unlike the first embodiment, the shooting timings and the frame rates (fps) of the shooting speeds of each of the cameras 13 and 16 are different. Therefore, the number of captured images acquired by the CPU 301 may also be different for each of the cameras 13 and 16.

[0193] It is preferable that the groups of captured images captured by each of the cameras 13 and 16 are groups of captured images whose shooting start times and shooting end times generally coincide. Therefore, the group of captured images acquired by the CPU 301 in S1301 may also be a group of captured images that are determined and associated only by determining that they were captured in a time zone that is approximately close from the group of captured images given an inaccurate time code including a timing deviation.

[0194] Alternatively, a user who knows that the same target person 11 was photographed at approximately the same time without referring to the time code may select a set of image sequences obtained by photographing with the multiple cameras 13, 16 after photographing and input the sets into the information processing device 300. In this flowchart, the two cameras 13, 16 are assumed to have photographed the performance of the same target person 11 from start to finish. In this case, it is easy to find correspondence between image sequences of the target person in a three-dimensional reconstruction photographed by the multiple cameras at similar times.

[0195] In S1302, the CPU 301 performs skeleton estimation on a series of image sequences captured by the cameras 13 and 16. The skeleton estimation method can be implemented in the same manner as in the first embodiment, and therefore a description thereof will be omitted. In this embodiment, as in the first embodiment, each joint position has an estimated likelihood.

[0196] An architecture diagram for explaining the processing of the flowchart in FIG. 13 is shown in FIG. 5, as in the first embodiment. Input image 500, which is Image 0, is an image captured by camera 16. Input image 501, which is Image 1, is an image captured by camera 13. Input images 500 and 501 are received in S1301. In the first embodiment, person pose estimation (skeleton estimation) was performed for each input image having the same time code or a time code difference within a predetermined value, and camera calibration 1 was performed by associating the images on the assumption that the same point was observed. The skeleton to be estimated in this embodiment may be defined in any way, but the human body skeleton definition exemplified in FIG. 6 is used, as in the first embodiment.

[0197] In this embodiment, since the cameras 13 and 16 do not capture images of the target person 11 at the same time, it is not possible to associate the images captured by the camera 13 with the images captured by the camera 16 by referring only to the time code. Therefore, it is not possible to perform camera calibration based only on a skeleton estimated for a single frame, as was performed in the first embodiment. For this reason, the image sequence for which the CPU 301 performs skeleton estimation may be the entire image sequence captured by the user with the cameras 13 and 16 within a predetermined time period. Alternatively, it may be an image sequence between the start and end times of the movement of the target for 3D reconstruction specified and input by the user via a GUI or the like.

[0198] In S1303, the CPU 301 identifies joint points that can be used for camera calibration 1 using the joint positions of the estimated skeleton. Joint points that are suitable for use in camera calibration 1 are those at joint positions that are almost stationary within a specific period of time, and therefore, such joint points are identified in S1303.

[0199] For example, the joint position coordinates of the right foot joint position (R_FOOT) in FIG. 6 estimated from a walking person can be treated as being substantially stationary from the time the right foot places on the ground until the foot kicks off the ground in the direction of travel and then leaves the ground. Even if there is a difference in the shooting times of the multiple cameras 13 and 16, there will be a joint position that remains stationary for a period longer than the difference in the shooting times. Therefore, in S1303, such a joint position may be identified. For example, from a group of captured images obtained by capturing images with p different cameras, a frame set whose assigned time codes are separated by a predetermined threshold ζ or less (e.g., ζ = 1.0 sec) is extracted. Then, from the extracted frame set, only joint points whose estimated joint positions have a movement distance of a predetermined threshold η or less (e.g., η = 3 pixels or less) between previous and next frames in the chronological order are identified.

[0200] In S1304, the CPU 301 performs camera calibration 1 in the same manner as in S403 in the first embodiment. The difference from the first embodiment is that the joint points used in camera calibration 1 are the joint points identified in S1303, and that there is a possibility that some error may have increased depending on the selection criteria. However, because camera calibration 1 only requires that a rough level of camera calibration be performed, calibration 1 may also be performed in S1304 using the same procedure as in S403.

[0201] In S1305, camera calibration 2 is performed with the aim of achieving the same effect as in S404 in the first embodiment.

[0202] FIG. 13B is a flowchart showing the details of S1305.

[0203] In S1311, the CPU 301 acquires, as initial values, the camera parameters derived as a result of camera calibration 1 in S1304 and the parameters indicating the joint positions of the human body derived as a result of S1302. Alternatively, if the total loss, which will be described later, exceeds the threshold value τ in S1317 and the parameters are updated in S1318, the CPU 301 acquires the updated camera parameters and joint position parameters in S1311.

[0204] In S1312, the CPU 301 converts the position information of each joint in each camera coordinate estimated from each input image into position information in world coordinates using the external camera parameters acquired in S1311, in the same way as in S412. Then, the CPU 301 integrates the skeleton information. That is, one piece of skeleton information is generated in which the positions of each joint of the object are indicated in three-dimensional world coordinates. This integration of the skeleton information is also performed in the same way as in S412. Therefore, the X corrected as a result of the integration is j,f is passed to Skeletal Transformation 510 in FIG. 5 in S1315 and S1316.

[0205] In the first embodiment, camera calibration 2 was performed assuming that multiple cameras captured images of the target person 11 synchronized with each other and that the estimated joint positions ideally corresponded to the same point in world coordinates. In this embodiment, camera calibration 2 cannot be performed using the estimated joint positions as they are. For this reason, in this embodiment, a loss ε1, among the losses considered as optimization targets, is defined as a loss based on a trajectory indicating the movement of the estimated joint point positions. In the first embodiment, a method was described in which minimizing the reprojection error was formulated as a maximum likelihood estimation problem, and the estimated joint position in two-dimensional coordinates was assumed to be approximated by a normal distribution with a standard deviation σ. The difference from the reprojected joint position was defined as the loss ε1, and the quality of the camera position and orientation was reflected in the score. In this embodiment, identical three-dimensional joint positions do not generally exist. For this reason, in this embodiment, trajectories of the movement of joint points, which should essentially be the same across multiple cameras, are derived, and the distance between these trajectories is defined, and the difference distance between the trajectories is evaluated as the loss ε1.

[0206] For this reason, first, in step S1313, the CPU 301 estimates, for each camera, a trajectory indicating the continuous movement of the positions of the joint points of the human skeleton in the input time-series image sequence. Since 17 joints are estimated for each frame of a human body to be reconstructed in 3D, the trajectories of 17 joint points are estimated for each camera in the sequence that continues in the time series direction.

[0207] The camera 13 continuously captures images of the target person 11 in a time series direction, performs human skeleton estimation for each of a group of captured images, which are multiple frames obtained by capturing images, and plots the estimated position coordinates of only the positions of the joint points PELVIS in three-dimensional space. Then, from the plotted discrete positions of the joint points PELVIS, a trajectory (track) that is the transition of the positions of the joint points PELVIS in a time series is estimated.

[0208] FIG. 15(a) shows, with a dashed line, a trajectory 1505 of the movement of the joint point PELVIS estimated from the images captured by the camera 13. The starting point 1501 of the dashed line is a point estimated as the position of the joint point PELVIS of the target person 11 in the captured image at the start of capture. The ending point 1504 of the dashed line is a point estimated as the position of the joint point PELVIS of the target person 11 in the captured image at the end of capture. Furthermore, points 1502 and 1503 are points estimated as the positions of the joint point PELVIS of the target person 11 in the captured images between the starting point 1501 and the ending point 1504. For ease of viewing and understanding, FIG. 15(a) illustrates an example in which human body skeleton estimation was performed four times from the start to the end of capture. Naturally, a better estimation result of the trajectory of the joint point can be obtained by performing human body skeleton estimation more frequently with shorter capture time intervals.

[0209] FIG. 15(b) is a diagram showing a trajectory 1505 of the movement of the joint point PELVIS estimated from the image captured by the camera 13 and a trajectory 1507 of the movement of the joint point PELVIS estimated from the image captured by the other camera 16. Even if the start and end times of the captures of the two cameras 13 and 16 are different, as described above, the two cameras 13 and 16 capture the performance of the same target person 11 from start to finish. Therefore, most of the two trajectories 1505 and 1507 of the same joint point should actually match in world coordinates. However, in the initial stage, trajectory estimation is performed independently by each camera 13 and 16. Therefore, as shown in FIG. 15(b), a difference occurs between the trajectory 1505 and the trajectory 1507 of the same joint point PELVIS.

[0210] The trajectory of the joint point estimated from the movement of the same joint position (e.g., PELVIS in FIG. 6) estimated from multiple cameras within a predetermined time should originally completely coincide in the world coordinates. Therefore, it is difficult to perform camera calibration 2 using the displaced positions when the shooting times do not match based on the joint position criteria in a single frame. However, if it is possible to accurately estimate the movement trajectory of the three-dimensional points by tracking the discretely observed joint positions in the time series direction, camera calibration 2 can be performed based on the trajectory of the joint point. The difference between the trajectories 1505 and 1507 of the same joint point estimated in the initial S1313 can be reduced by estimating the camera parameters representing the position and orientation of the camera to approach the correct values.

[0211] Next, as a specific method for estimating the trajectory of the joint point, a method for estimating the trajectory from a plurality of discrete joint positions of the skeleton that are continuously displaced in the time series direction estimated from the captured images with different shooting times by each camera will be described. The method for estimating the trajectory of the joint point is performed, for example, by an algorithm that reconstructs the three-dimensional trajectory from the moving points in the two-dimensional perspective projection. In this case, each trajectory is represented by a linear combination of compact trajectory basis functions based on the positions of the discrete estimated joint points. At this time, the trajectory coefficient vector by the linear least squares method is solved.

[0212] Similar to the first embodiment, now, the intrinsic internal camera parameters calibrated in advance for each known camera p (0 ≤ p < P) are represented as K p and denoted. Also, the position and orientation of the camera are represented by R p and t p and this parameter set <R p , t p > is estimated. In addition, the notation of the symbols used in the mathematical formulas also conforms to that of the first embodiment.

[0213] Also, similar to the first embodiment, the optimization target parameters K p , R p , t p、Calculate the error with respect to X, and solve the problem of estimating a continuously changing trajectory from the discrete skeleton joint positions estimated under the conditions of each optimization target parameter. Estimating this trajectory is a process performed to define the error as a distance, and it can be performed by any method as long as an optimizable distance can be calculated.

[0214] Among the captured images obtained by camera p, the position of joint j in the three-dimensional world coordinates estimated from frame f and the position of joint j in the three-dimensional coordinates in the camera coordinate system of camera p are as defined in Equation 1. f in Equation 1 is a parameter related to p as described above. Similarly, when the position of joint j on the two-dimensional image is represented as x p j,f When represented as, at camera p, the position x p j,f of joint j on the two-dimensional image and the position X p j,f of joint j on the original three-dimensional coordinates can be represented by Equation 2.

[0215] There is a point cloud obtained by tracking the estimated joint point j in the time series direction, and a set of three-dimensional trajectories is derived from these estimated points by the method described below. Here, the three-dimensional trajectory to be derived is represented as G(j), and this structure is defined as Equation 14. Equation 14 is not a definition formula for three-dimensional trajectories that varies depending on the camera, but a definition formula for the three-dimensional trajectory to be obtained that is consistent for all cameras. Therefore, Equation 14 is expressed without including p which is the camera ID. [[ID=asc="19"]]

[0216]

Number

[0217] X 0,j,f represents the position coordinate in the x coordinate of the three-dimensional coordinates when the joint ID is j and the frame is f (0 ≤ f < F). X 1,j,f represents the position coordinate in the y coordinate of the three-dimensional coordinates when the joint ID is j and the frame is f. X 2,j,frepresents the position coordinates when the joint ID in the Z coordinate of the three-dimensional coordinate system is j and the frame is f.

[0218] Next, the points on the 3D coordinates estimated using Equation 14 are divided into a point cloud set for each camera p, and an approximation calculation is performed on the trajectories of the joint points estimated from each camera p. The trajectories of the joint points corresponding to each camera are linear combinations of the base trajectories used for the approximation calculation, and the trajectories of each joint point for the number P of cameras are found using Equation 15.

[0219]

number

[0220] a 0,i (p,j), a 1,i (p,j), a 2,i (p,j) represents the coefficients of the basis vector. Equation 15 expresses G0^(p,j), G1^(p,j), and G2^(p,j) as linear combinations of k predefined basis orbitals. When defined as in Equation 15, calculations can be performed using, for example, a method using a predefined DCT (Discrete Fourier Transform) basis, a DWT (Discrete Wavelet Transform) basis, or a Hadamard transform basis as the orbital basis vector.

[0221] By minimizing the estimation error of the trajectory of each joint point estimated for each camera p using the trajectory basis vectors exemplified above, errors in the three-dimensional trajectories inferred from the joint points estimated from the images captured by each camera are corrected. In the case of two cameras, the goal can be achieved by an optimization method that minimizes the estimation error of the trajectory of each joint point estimated from two cameras, p=0 and P=1. Therefore, by expressing the trajectory of each joint point as in Equation 15, it can be expressed with k parameters per coordinate, and the error between the trajectories estimated by each camera can be calculated. In this case, the total number of bases k can be achieved by determining a predetermined number in advance.

[0222] In step S1314, the CPU 301 calculates the loss ε1 calculated based on the trajectory reference of the moved joint position. p , R p , t p , calculate the loss ε1 of the trajectory criterion for optimizing X.

[0223] The time direction component f listed in Equation 15 p Optimization will be performed on three types of trajectories for the displacement in the three axes X, Y, and Z relative to the object. In the case of two cameras, the trajectory of the joint point observed by camera p=0 and the trajectory of the joint point observed by camera p=1 are essentially trajectories that track the position of the same joint point. Therefore, the trajectories of the joint points estimated from the images captured by each camera must match in world coordinates, so by optimizing to minimize the difference between the estimated trajectories, we can approach the desired trajectory.

[0224] FIG. 16(a) is a diagram showing the estimated displacement of a certain joint j in the X-axis direction relative to the time direction component f. FIG. 16(a) depicts a displacement 1600 of the trajectory in the X-axis direction relative to f estimated from the camera 13 where p=1, and a displacement 1601 of the trajectory in the X-axis direction relative to f estimated from the camera 16 where p=0. The displacement 1600 is a diagram showing the estimated trajectory 1505 in the three-dimensional space in FIG. 15(b) based on the displacement in the X-axis direction relative to the time direction component f. The displacement 1601 is a diagram showing the estimated trajectory 1507 in the three-dimensional space in FIG. 15(b) based on the displacement in the X-axis direction relative to the time direction component f. FIG. 16(a) depicts the positions of the trajectories estimated from each camera relative to the time direction component f, but the positions of the plots do not match because the timing of the images captured by each camera does not match. Therefore, for example, suppose that points included in the displacement 1600 and the displacement 1601 at adjacent f are compared. The time of point 1602 photographed by camera 13 with p=1 is f0, and the time of point 1603 photographed by camera 16 with p=0 is f1. Therefore, Fig. 16(a) shows the trajectories of the joint points estimated from the photographed images when the photographing timings of the cameras are different.

[0225] Naturally, just like the trajectory of the three-dimensional joint point described above, even if the graph is a displacement graph in the X-axis direction against the time component f, it should match because it is a displacement on the X-axis in world coordinates estimated by photographing the joint j of the exact same person with each camera. In the initial state, these two trajectories do not match. Therefore, the distance between the trajectories estimated by each camera for joint j is δ j,d This is calculated as an error by defining it as Equation 16, and the distance δ j,d The camera position and orientation can be roughly estimated by performing an optimization calculation to minimize the distance δ. The subscript d indicates the X axis where d=0, the Y axis where d=1, and the Z axis where d=2. That is, the distance δ on the X axis j,0 , distance δ on the Y axis j,1 , distance δ on the Z axis j,2 Optimization calculations are performed to reduce

[0226] Equation 16 is the simplest example of a definition of loss, and is the squared loss of the difference between trajectories. Additionally, the X-axis, Y-axis, and Z-axis are simply expressed as d = 0, 1, 2. By calculating the difference in distance for all joints j using Equation 16, loss ε1 was defined in S1314 based on the current camera position and orientation parameters and the joint position coordinates estimated at that time.

[0227]

number

[0228] By using the loss calculation results for the estimated joint point trajectories, it is possible to obtain better camera position and orientation estimation results while updating the camera position and orientation through iterative processing, similar to the case where estimated joint position coordinates are given as described in the first embodiment. As a result, for example, it can be confirmed that the trajectory displacements 1600 and 1601 of the multiple cameras 13 and 16, which are in the initial state as shown in Figure 16(a), have converged into multiple close curves, as shown in Figure 16(b) for the trajectory displacements 1604 and 1605.

[0229] In S1315, the CPU 301 uses the camera position and orientation as initial values ​​and improves the accuracy of the camera position and orientation while referring to the 3D reconstruction result for the person. S1315 and S1316 follow the same flow as S414 and S415, so detailed explanations will be omitted. In S1315, the CPU 301 performs 3D reconstruction using the camera parameters acquired in S1311. Then, in S1316, the CPU 301 calculates the rendering loss ε2 defined in Equation 12. The final camera calibration result and 3D reconstruction result can be obtained by optimizing based on the loss ε2. Thereafter, in S1317, a comparison with a threshold τ, which is the termination condition for this optimization process, is made, and if the condition is met, the processing of S1305 is terminated.

[0230] As described above, according to this embodiment, even when a target person is photographed asynchronously with a small number of cameras, accurate camera calibration can be performed by updating the estimated results of the camera positions and orientations while performing 3D reconstruction, thereby making it possible to obtain high-quality 3D reconstruction results.

[0231] Furthermore, in this embodiment, the example has been described in which the subject of photography is a single person, but it goes without saying that this can be achieved even when there are multiple subjects, or when the subject is an animal or non-living object other than a person, by introducing the loss calculation method implemented in this embodiment into the processing flow of the second embodiment.

[0232] <Fourth embodiment> This embodiment is a modified example of the third embodiment. The following description will focus on the differences from the third embodiment. Unless otherwise specified, the configuration and processing are the same as those of the third embodiment.

[0233] In the third embodiment, the trajectories of consecutive joint points of a human skeleton in a time-series image sequence input for each camera in S1313 are estimated based on the definition of Equation 15. In the third embodiment, the trajectory basis vectors are assumed to be calculated using predefined DCT (Discrete Fourier Transform) bases, DWT (Discrete Wavelet Transform) bases, and Hadamard Transform bases. In this embodiment, a method will be described in which Gaussian basis functions are used instead of the trajectory basis vectors by redefining Equation 15 as Equation 17. Using Equation 17 makes the method even easier to implement and can eliminate the influence of estimated joint points with large estimation errors.

[0234] In Equation 17, the trajectory of a three-dimensional point is estimated by dividing the three-dimensional points that change in the time series direction into X, Y, and Z axis components, and simplifying it as a graph when the horizontal axis is taken as the time direction component. Note that the horizontal axis can also be the time code of the time direction component, but for the purposes of notation it is taken as f, the frame ID.

[0235]

number

[0236] The change from Equation 15 to Equation 17 is that the basis function is related to the horizontal axis f, so the predefined orbital basis vector in Equation 15 is θ i In Equation 17, θ i (f) is a function related to f. In Equation 17, the coefficients a i p,j,0 Each trajectory is calculated by superposing k bases. In Equation 17, k must be defined in advance, but the trajectory basis function for f can be calculated only around the observed point, so the number of basis functions used to estimate the trajectory for camera p is the sum of all f p (0≦f p <F p ) is a subset of F p is the total number of images taken by camera p. <F p Therefore, k is 0≦k <F p However, the estimation accuracy improves if a large value is set. Therefore, in this example, all the photographed numbers, that is, the f corresponding to the joint position coordinates photographed by camera p and estimated, are used. p A Gaussian function is given to all of them on the horizontal axis and optimized for all of them. i p,j,0 , a i p,j,1 , a i p,j,2 The problem is to find the following. Trajectory estimation is performed using the above method for all joints to be estimated.

[0237] Next, the estimated trajectory is corrected in S1313 using the estimated joint point trajectory. p , R p , t p , X. As in the third embodiment, the distance between the trajectories estimated for each joint j is defined as δ j,dThen, the displacement on the X axis, the displacement on the Y axis, and the displacement on the Z axis are combined to define ε1 as equation 18. ε1 defined in equation 18 is calculated as the loss, and an optimization calculation is performed to reduce the loss ε1. This makes it possible to roughly estimate the camera position and orientation. As an example of the simplest definition of loss, equation 18 uses the squared loss of the difference between trajectories. Additionally, the X axis, Y axis, and Z axis are simply expressed as d=0, 1, 2.

[0238] By calculating the distance difference for all joints j using Equation 18, the loss ε1 is defined in S1314 based on the current camera position and orientation parameters and the joint position coordinates estimated at that time.

[0239]

number

[0240] The loss ε1 defined in Equation 18 has the same format as in the third embodiment. However, unlike the third embodiment, as defined in Equation 17, G p,j,0 ^(f p ) is composed of k basis functions. In reality, some joint positions estimated at each moment of photography are very close to the true value, while others are not. As mentioned in the explanation of skeleton estimation, each joint position has an estimated likelihood, so it is possible to determine whether a joint point is reliable based on the estimated likelihood.

[0241] Unlike the third embodiment, this embodiment uses directly estimated joint positions for trajectory estimation. Therefore, by performing basis function trajectory estimation only for positions with high estimated likelihood, without using joint positions with low estimated likelihood, it becomes easier to avoid estimating erroneous trajectories. In addition, joint positions may be estimated at positions with high estimated likelihood but far from the true value. For example, a joint position at a hidden position not observed by the camera capturing the image may be estimated at a statistically likely position, but the true pose may not be at that position. In this case, a method is required to eliminate erroneous estimated joint positions using an index other than the likelihood output by the skeleton estimator. In this case, by reducing and bringing together the global differences between trajectories estimated using k trajectory basis functions from different cameras, and then searching for areas where the differences are large locally, outlier estimated joint points can be accurately detected.

[0242] Assume that a part of the trajectory moving in the time series direction of a given joint position has a large noise on the original estimated joint position, which causes a deterioration in the accuracy of trajectory estimation. In order to identify the joint position that has a large noise mixed in, the following process is performed. The score for determining the estimation error is W(f p,j,d ) is defined as W(f p,j,d ) is defined by Equation 19.

[0243]

number

[0244] When searching for a point where the error does not decrease even if the optimization to minimize the loss ε1 is carried out, the score W(f p,j,d ) can be used to identify points where joint position estimation is incorrect. p,j,d ) is considered to be an incorrect point, and the basis function f p,j,d is excluded as an outlier. For example, the score W(f p,j,d) exceeds a predetermined threshold, are removed so as not to be used in trajectory estimation, and the loss ε1 is recalculated. By determining that the joint points on the trajectory estimated after removing the erroneous points are the correct joint positions, it becomes possible to estimate closer to the true value. By processing in this manner, it is possible to estimate the trajectories of joint points with higher reliability in S1314, and as a result, it is possible to determine not to use erroneous poses during 3D reconstruction. By using the loss calculation results for the estimated trajectories of joint points, it is possible to obtain better camera position and orientation estimation results by repeatedly updating the camera position and orientation in the same way as when estimated joint position coordinates are given as described in the first embodiment. The processing flow from S1315 onwards is the same as in the third embodiment.

[0245] As described above, according to this embodiment, it is possible to easily estimate trajectories indicating the time-series directions of joint positions of a person by asynchronous imaging. Also, it is possible to obtain high-quality 3D reconstruction results while performing accurate camera calibration.

[0246] <Fifth embodiment> This embodiment is a modification of the third and fourth embodiments. In the above-described embodiments, a method for simultaneously performing camera calibration and 3D reconstruction using only captured images without performing a camera calibration process using a fixed pattern when capturing images using a small number of cameras with a small common field of view has been described. Also, in the third and fourth embodiments, a method for performing camera calibration with high accuracy when capturing images using two or three cameras has been described, even when the capture timing is shifted and images are captured continuously in a chronological order. In this embodiment, a method for performing good 3D reconstruction even in sports scenes where the object for 3D reconstruction moves quickly will be described.

[0247] FIG. 17 is a diagram showing a soccer stadium. There may be cases where 3D reconstruction of sports play content is desired in a location such as that shown in FIG. 17. In this case, for example, a visual hull may be used to perform 3D reconstruction of objects such as people in the stadium. When reconstructing the 3D shape of an object using a visual hull, the 3D shape is estimated based on the common area of ​​the view volume indicated by the silhouette of the object to be reconstructed in 3D shape by each camera. To perform accurate 3D estimation using a visual hull, it was necessary for all cameras 1701 installed in the stadium to take photos at the same time. Therefore, even when 3D reconstruction of sports play content is desired in a location such as that shown in FIG. 17, by adopting the methods according to the third and fourth embodiments, good 3D reconstruction results can be obtained even if the images were taken at different times.

[0248] Furthermore, methods like Visual Hull require meticulous camera calibration before the start of the game, and the camera's position and orientation must remain constant throughout the entire game. Therefore, if an interesting scene occurs during a game that is to be replayed on television, it may be difficult to accurately reconstruct the scene in 3D if there are only a few cameras capturing the area where the scene occurred at high resolution. For this reason, it may be possible to predetermine the number of locations within a large space where 3D reconstruction is possible, and only the plays that occurred within those locations may be used for 3D reconstruction. There is also a need to move the camera as the game progresses, rather than using a fixed camera, so that various spaces can be used for 3D reconstruction.

[0249] When capturing images with a panning camera while constantly tracking the subject, it is necessary to dynamically calibrate the camera for each captured scene. When only the soccer field is captured at high resolution, it is often difficult to calibrate the camera using natural features derived from stationary objects. Therefore, by performing camera calibration using the human skeleton estimation results according to the above-described embodiment, camera calibration and 3D reconstruction can be performed as soon as the capture of an interesting scene begins.

[0250] 18 is a flowchart for explaining the processing flow of camera calibration and 3D reconstruction in this embodiment. It is assumed that multiple cameras 1701 installed in the stadium detect the soccer ball and track the ball position as needed, while constantly moving their heads to capture the game content.

[0251] In S1801, the CPU 301 instructs the multiple cameras 1701 shown in FIG. 17 to start shooting with a predetermined time difference in order to increase the shooting resolution in the time series direction. Each of the multiple cameras starts shooting with the shooting start time shifted by the predetermined time difference. The predetermined time difference is, for example, the time for one frame (1 / M f ) is less than the time difference. d Then, for example, based on Equation 20, the time difference T d is determined.

[0252] [Formula 20] T d =1.0 / (M f ×N) N is the number of cameras 1701 capturing the stadium, and M f (fps) is the shooting frame rate of N cameras.

[0253] In S1802, CPU 301 determines a scene to be reconstructed in three dimensions. For example, the space to be reconstructed in three dimensions is determined automatically or by manual instruction from an image captured by camera 1701 capturing an image of a stadium where a game to be captured is being played, as shown in FIG. 17. CPU 301 can automatically select a scene to be reconstructed in three dimensions by automatically detecting impactful times in the game content, triggered by the moment the ball enters the goal or the moment the crowd's voices get louder, and then determining the scene to be reconstructed in three dimensions. Selecting salient scenes from a series of scenes can be done using a rule-based method or by detecting salient scenes using an existing machine learning method.

[0254] In S1803, the CPU 301 acquires an image sequence captured by the multiple cameras 1701 in a target time range from a time several to several tens of seconds before the time indicating the scene determined in S1802 to the time of the determined scene. In S1803, it is not necessary to acquire images captured in the target time range from all cameras 1701 installed in the stadium. For example, images captured in the target time range may be acquired only from cameras capturing the entire body of the target person to be 3D reconstructed in all frames corresponding to the target time range.

[0255] S1804 is a step similar to S802 in FIG. 8 of the second embodiment, in which area segmentation and tracking are performed to detect the area of ​​each object.

[0256] The loop processing of S1805 to S1807 is the same as the loop processing of S803 to S805 in Fig. 8 of the second embodiment. That is, an object to be processed is selected from objects in the target time range, and joint position estimation is performed using an appropriate model for the object to be processed. As a result, joint position estimation is performed for all objects in all frames in the target time range.

[0257] In S1808, camera calibration 1 is performed using a method similar to the method described in S806 of the second embodiment in Fig. 8. In S1808, the camera parameters are roughly determined.

[0258] In step S1809, calibration 2 is executed. Details of the internal processing in step S1809 are equivalent to the methods described in the third or fourth embodiment. That is, in step S1809, the same processing as in steps S1311 to S1318 in FIG. 13(b) is executed.

[0259] In step S1809, the 3D reconstruction shown in the third and fourth embodiments is performed using the image sequences from each camera obtained by time-shifted shooting. This enables 3D reconstruction with human body movement at a high frame rate, which is not possible with conventional shooting using a small number of cameras. For example, assume that each camera shoots at 60 fps, and after each camera moves its shooting direction to capture the target person, image sequences from 20 cameras are used for 3D reconstruction. Because the 20 cameras shoot with a time lag of one frame at 60 fps, the virtual viewpoint video generated from the 3D reconstruction results can be a 60 x 20 = 1200 fps video. Therefore, this embodiment allows for the capture of scenes of athletes quickly moving around the field at an ultra-high frame rate and for later viewing from any viewpoint.

[0260] <Other embodiments> In the above-described embodiment, the description was given assuming that the image was taken with a fixed camera. However, if there are a sufficient number of key points that can be obtained from the object, the limitation of taking images with a fixed camera is not necessary. Even when synchronous images are taken with a handheld camera, accurate 3D reconstruction is possible.

[0261] FIG. 19 is a diagram showing how handheld cameras 1901 and 1902 capture images of a target 1900 for three-dimensional reconstruction. In this case, the images captured by each handheld camera 1901 and 1902 may include objects such as backgrounds 1903 and 1904 in addition to the target person 1900 for three-dimensional reconstruction. In this case, it is easy to estimate the position and orientation of the handheld camera at each time in the camera coordinate system using static points in the scene such as backgrounds 1903 and 1904 and SfM (Structure from Motion) in Non-Patent Document 10. Similarly, when three-dimensional reconstruction is performed including backgrounds 1903 and 1904, it is easy to estimate the position and orientation of the handheld camera using backgrounds 1903 and 1904. Therefore, all of the above-described embodiments described as methods implemented in the case of fixed cameras can also be implemented as embodiments using handheld cameras.

[0262] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0263] The techniques disclosed in the above-described embodiments include the following configurations.

[0264] (Configuration 1) an acquisition means for acquiring photographed images obtained by each of the plurality of photographing devices photographing an object; a detection means for detecting a position of a predetermined part of the object from each of the images captured by the plurality of image capturing devices; an estimation means for estimating camera parameters indicating the position and orientation of each of the plurality of image capturing devices using the detected positions of the predetermined portion; an updating means for updating the camera parameters of each of the plurality of image capturing devices using the estimated camera parameters as initial values; a determination means for determining camera parameters for each of the plurality of image capturing devices based on a result of three-dimensional reconstruction of the object based on the updated camera parameters; An information processing device comprising: (Configuration 2) the result of the three-dimensional reconstruction of the object is a model that, when viewpoint information is input, outputs an image of the object observed from the input viewpoint information; The determining means determines the camera parameters of each of the plurality of image capturing devices based on a result of comparing an output image obtained by inputting viewpoint information indicated by the updated camera parameters into the model with the captured images of each of the plurality of image capturing devices. 2. The information processing device according to configuration 1, (Configuration 3) The determining means When a value indicating a difference between the output image and each of the images captured by the plurality of image capturing devices becomes equal to or smaller than a threshold, the updated camera parameters are determined as the camera parameters for each of the plurality of image capturing devices. 3. The information processing device according to configuration 2. (Configuration 4) a processing means for performing three-dimensional reconstruction of the object for each of a plurality of frames captured by each of the plurality of image capturing devices; and an integration means for integrating the result of the three-dimensional reconstruction performed by the processing means with the updated camera parameters. 4. The information processing device according to configuration 2 or 3. (Configuration 5) the integration means performs the integration by converting the posture of the object represented by the result of the three-dimensional reconstruction performed by the processing means into a predetermined posture; and an inverse transformation means for inversely transforming the pose of the object represented by the integrated result of the three-dimensional reconstruction; The model is a model obtained by the inverse transformation. 5. The information processing device according to configuration 4. (Configuration 6) The update means and updating camera parameters of each of the plurality of image capturing devices further based on the position of the predetermined part. 6. The information processing device according to any one of configurations 1 to 5. (Configuration 7) the detection means detects the position of the predetermined part in the captured image of each of the plurality of imaging devices; The method further includes integrating means for integrating the positions of the predetermined parts detected from each of the photographed images using the updated camera parameters, The update means and updating the camera parameters of each of the plurality of photographing devices based on a result of comparing the position of the predetermined part in the photographed image detected by the detection means with the position of the integrated predetermined part converted into a position in the photographed image. 7. The information processing device according to configuration 6. (Configuration 8) The determining means calculating a first value indicating an error in the updated camera parameters based on the position of the predetermined part based on the updated camera parameters; calculating a second value indicating an error in the updated camera parameters based on a result of three-dimensionally reconstructing the object based on the updated camera parameters; When a value based on the first value and the second value becomes equal to or smaller than a threshold, the updated camera parameters are determined as camera parameters for each of the plurality of image capturing devices. 8. The information processing device according to configuration 6 or 7. (Configuration 9) The estimation means determining a likelihood of each of the positions of the predetermined body part detected from the images captured by each of the plurality of imaging devices; The first value is a value further based on the likelihood. 9. The information processing device according to configuration 8. (Configuration 10) The estimation means A likelihood of the position of each of the predetermined parts detected from the captured images of each of the plurality of image capturing devices is determined, and camera parameters indicating the position and orientation of each of the plurality of image capturing devices are estimated without using the positions of the predetermined parts whose likelihood is equal to or less than a threshold value. 10. The information processing device according to any one of configurations 1 to 9. (Configuration 11) the object photographed by each of the plurality of photographing devices is a plurality of objects, a generation unit that generates a mask representing a region of each of the plurality of objects in the captured image, the mask being obtained by tracking each of the plurality of objects in a time series direction; The determining means determines camera parameters for each of the plurality of image capturing devices based on a result of comparing the mask area between the output image and the captured image. 3. The information processing device according to configuration 2. (Configuration 12) The image capturing device further includes an assigning unit that assigns an identifier to each of the plurality of objects included in the captured image of each of the plurality of image capturing devices by instance segmentation. 12. The information processing device according to configuration 11. (Configuration 13) the object photographed by each of the plurality of photographing devices is a plurality of objects, the models correspond to a plurality of objects, The display control means further includes a display control means for displaying an image in which the object selected by the user is rendered using the model of the object selected by the user. 3. The information processing device according to configuration 2. (Configuration 14) further comprising a trajectory estimation means for estimating a trajectory indicating a change in the position of the predetermined part for each of the plurality of imaging devices; The update means updating the camera parameters of each of the plurality of image capturing devices so that the estimated trajectories for each of the plurality of image capturing devices coincide with each other; 6. The information processing device according to any one of configurations 1 to 5. (Configuration 15) The trajectory estimation means Estimating the trajectory for each of the plurality of image capture devices using predefined trajectory basis vectors or basis functions. 15. The information processing device according to configuration 14. (Configuration 16) The update means The trajectories estimated for each of the plurality of imaging devices are compared for each directional component in three-dimensional space. 16. The information processing device according to configuration 14 or 15. (Configuration 17) The camera further includes an instruction means for instructing the plurality of image capturing devices to start capturing images by giving a predetermined time difference that is less than the time for one frame. 17. The information processing device according to any one of configurations 14 to 16. (Configuration 18) the object is a person or an animal; The predetermined parts are each joint of a person or an animal. 18. The information processing device according to any one of configurations 1 to 17. (Configuration 19) the object is an inanimate object; The position of the predetermined portion is the position of the center of the object. 19. The information processing device according to any one of configurations 1 to 18. (Configuration 20) The plurality of imaging devices may be two or three imaging devices. 20. The information processing device according to any one of configurations 1 to 19. (Configuration 21) an acquisition step of acquiring photographed images obtained by each of the plurality of photographing devices photographing an object; a detecting step of detecting a position of a predetermined part of the object from each of the images captured by the plurality of image capturing devices; an estimation step of estimating camera parameters indicating the position and orientation of each of the plurality of image capturing devices using the detected positions of the predetermined part; an updating step of updating the camera parameters of each of the plurality of image capturing devices using the estimated camera parameters as initial values; a determining step of determining camera parameters for each of the plurality of image capturing devices based on a result of three-dimensional reconstruction of the object based on the updated camera parameters; An information processing method comprising: (Configuration 22) 21. A program for causing a computer to execute each means of the information processing device according to any one of configurations 1 to 20. [Explanation of symbols]

[0265] 300 Information processing device 301 CPU

Claims

1. an acquisition means for acquiring photographed images obtained by each of the plurality of photographing devices photographing an object; a detection means for detecting a position of a predetermined part of the object from each of the images captured by the plurality of image capturing devices; an estimation means for estimating camera parameters indicating the position and orientation of each of the plurality of image capturing devices using the detected positions of the predetermined portion; an updating means for updating the camera parameters of each of the plurality of image capturing devices using the estimated camera parameters as initial values; a determination means for determining camera parameters for each of the plurality of image capturing devices based on a result of three-dimensional reconstruction of the object based on the updated camera parameters; An information processing device comprising:

2. the result of the three-dimensional reconstruction of the object is a model that, when viewpoint information is input, outputs an image of the object observed from the input viewpoint information; The determining means determines the camera parameters of each of the plurality of image capturing devices based on a result of comparing an output image obtained by inputting viewpoint information indicated by the updated camera parameters into the model with the captured images of each of the plurality of image capturing devices.

2. The information processing apparatus according to claim 1, wherein:

3. The determining means When a value indicating a difference between the output image and each of the images captured by the plurality of image capturing devices becomes equal to or smaller than a threshold, the updated camera parameters are determined as the camera parameters for each of the plurality of image capturing devices.

3. The information processing apparatus according to claim 2, wherein:

4. a processing means for performing three-dimensional reconstruction of the object for each of a plurality of frames captured by each of the plurality of image capturing devices; and an integration means for integrating the result of the three-dimensional reconstruction performed by the processing means with the updated camera parameters.

3. The information processing apparatus according to claim 2, wherein:

5. the integration means performs the integration by converting the posture of the object represented by the result of the three-dimensional reconstruction performed by the processing means into a predetermined posture; and an inverse transformation means for inversely transforming the pose of the object represented by the integrated result of the three-dimensional reconstruction; The model is a model obtained by the inverse transformation.

5. The information processing apparatus according to claim 4,

6. The update means and updating camera parameters of each of the plurality of image capturing devices further based on the position of the predetermined part.

2. The information processing apparatus according to claim 1, wherein:

7. the detection means detects the position of the predetermined part in the captured image of each of the plurality of imaging devices; The method further includes integrating means for integrating the positions of the predetermined parts detected from each of the photographed images using the updated camera parameters, The update means and updating the camera parameters of each of the plurality of photographing devices based on a result of comparing the position of the predetermined part in the photographed image detected by the detection means with the position of the integrated predetermined part when converted into a position in the photographed image.

7. The information processing apparatus according to claim 6,

8. The determining means calculating a first value indicating an error in the updated camera parameters based on the position of the predetermined part based on the updated camera parameters; calculating a second value indicating an error in the updated camera parameters based on a result of three-dimensionally reconstructing the object based on the updated camera parameters; When a value based on the first value and the second value becomes equal to or smaller than a threshold, the updated camera parameters are determined as camera parameters for each of the plurality of image capturing devices.

7. The information processing apparatus according to claim 6,

9. The estimation means determining a likelihood of each of the positions of the predetermined body part detected from the images captured by each of the plurality of imaging devices; The first value is a value further based on the likelihood.

9. The information processing apparatus according to claim 8,

10. The estimation means A likelihood of the position of each of the predetermined parts detected from the captured images of each of the plurality of image capturing devices is determined, and camera parameters indicating the position and orientation of each of the plurality of image capturing devices are estimated without using the positions of the predetermined parts whose likelihood is equal to or less than a threshold value.

2. The information processing apparatus according to claim 1, wherein:

11. the object photographed by each of the plurality of photographing devices is a plurality of objects, a generation unit that generates a mask representing a region of each of the plurality of objects in the captured image, the mask being obtained by tracking each of the plurality of objects in a time series direction; The determining means determines camera parameters for each of the plurality of image capturing devices based on a result of comparing the mask area between the output image and the captured image.

3. The information processing apparatus according to claim 2, wherein:

12. The image capturing device further includes an assigning unit that assigns an identifier to each of the plurality of objects included in the captured image of each of the plurality of image capturing devices by instance segmentation.

12. The information processing apparatus according to claim 11,

13. the object photographed by each of the plurality of photographing devices is a plurality of objects, the models correspond to a plurality of objects, The display control means further includes a display control means for displaying an image in which the object selected by the user is rendered using the model of the object selected by the user.

3. The information processing apparatus according to claim 2, wherein:

14. further comprising a trajectory estimation means for estimating a trajectory indicating a change in the position of the predetermined part for each of the plurality of imaging devices; The update means updating the camera parameters of each of the plurality of image capturing devices so that the estimated trajectories for each of the plurality of image capturing devices coincide with each other; 2. The information processing apparatus according to claim 1, wherein:

15. The trajectory estimation means Estimating the trajectory for each of the plurality of image capture devices using predefined trajectory basis vectors or basis functions.

15. The information processing apparatus according to claim 14,

16. The update means The trajectories estimated for each of the plurality of imaging devices are compared for each directional component in three-dimensional space.

15. The information processing apparatus according to claim 14,

17. The camera further includes an instruction means for instructing the plurality of camera devices to start taking pictures by giving a predetermined time difference that is less than the time for one frame.

15. The information processing apparatus according to claim 14,

18. the object is a person or an animal; The predetermined parts are each joint of a person or an animal.

2. The information processing apparatus according to claim 1, wherein:

19. the object is an inanimate object; The position of the predetermined portion is the position of the center of the object.

2. The information processing apparatus according to claim 1, wherein:

20. The plurality of image capturing devices may be two or three image capturing devices.

2. The information processing apparatus according to claim 1, wherein:

21. an acquisition step of acquiring photographed images obtained by each of the plurality of photographing devices photographing an object; a detecting step of detecting a position of a predetermined part of the object from each of the images captured by the plurality of image capturing devices; an estimation step of estimating camera parameters indicating the position and orientation of each of the plurality of image capturing devices using the detected positions of the predetermined part; an updating step of updating the camera parameters of each of the plurality of image capturing devices using the estimated camera parameters as initial values; a determining step of determining camera parameters for each of the plurality of image capturing devices based on a result of three-dimensional reconstruction of the object based on the updated camera parameters; An information processing method comprising:

22. A program for causing a computer to execute each means of the information processing device according to any one of claims 1 to 20.