Location estimation method, location estimation program, and information processing device.
By estimating camera position and orientation and correcting it based on joint trajectories, the method addresses the challenge of accurately determining the three-dimensional position of a person in moving images, improving estimation accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FUJITSU LTD
- Filing Date
- 2023-05-02
- Publication Date
- 2026-05-20
AI Technical Summary
Existing methods struggle to accurately estimate the three-dimensional position of a person in a moving image captured by a moving camera, as it is difficult to distinguish between camera movement and the person's movement, leading to inaccurate position estimation.
A method that estimates the camera's position and orientation for each frame, calculates joint torques, and corrects the camera's position and orientation based on joint trajectories to improve estimation accuracy, using a computer to perform coordinate estimation processes and simulations.
Enables accurate estimation of the three-dimensional position of a person from moving images captured by a moving camera, enhancing the precision of position estimation.
Smart Images

Figure 0007862758000002 
Figure 0007862758000003 
Figure 0007862758000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a location estimation method, a location estimation program, and an information processing device. [Background technology]
[0002] Techniques are known for estimating the state (position and orientation) of objects and people in three-dimensional space based on captured video footage. For example, one proposed device estimates the initial orientation of a multi-jointed object based on joint feature data and a motion model, generates motion vectors based on the estimated information, and estimates the most likely orientation among the orientations included in the state sample for each selected viewpoint candidate based on the motion vectors.
[0003] In addition to moving images, depth information is sometimes used to estimate the state of an object. For example, a device has been proposed that extracts three-dimensional mesh information from depth images and compares the three-dimensional coordinates of the mesh information with the object recognition results from color images to extract the three-dimensional coordinates of an object. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2013-20578 [Patent Document 2] U.S. Patent Application Publication No. 2016 / 364912 [Overview of the project] [Problems that the invention aims to solve]
[0005] However, when a person is filmed with a moving camera, there is a problem in that it is difficult to accurately estimate the person's three-dimensional position based on the moving image. For example, if the person's position in the image moves upward, it is difficult to determine whether the person's position in space moved upward due to an action such as jumping, or whether the camera's position moved downward, making it difficult to accurately estimate the person's three-dimensional position.
[0006] In one aspect, the present invention aims to provide a position estimation method, a position estimation program, and an information processing device that can accurately estimate the three-dimensional position of a person from moving images captured by a moving camera. [Means for solving the problem]
[0007] One proposed method provides a position estimation method in which a computer performs the following processes: In this method, the computer estimates a first position and orientation of the camera for each of several frames contained in a video image of a person captured by the camera, and performs a coordinate estimation process to estimate a first three-dimensional absolute coordinate for each of several joints in the person based on the first position and orientation. Furthermore, the coordinate estimation process for the first frame among the several frames includes the following first to third processes: In the first process, the computer estimates a second position and orientation of the camera corresponding to the first frame based on the positional relationship of several joints between the first frame and a second frame preceding the first frame, and the first position and orientation estimated from the second frame. In the second process, the computer estimates the joint torque acting on each of several joints in the first frame based on the positional relationship, and estimates the trajectory of each of several joints from the second frame based on the joint torque. In the third process, the computer estimates a first position and orientation corresponding to the first frame by correcting the second position and orientation based on the trajectory estimation result, and then estimates a first three-dimensional absolute coordinate for each of the multiple joints corresponding to the first frame based on the estimated first position and orientation. [Effects of the Invention]
[0008] On one side, the three-dimensional position of a person can be estimated with high accuracy from a moving image captured by a moving camera. The above and other objects, features, and advantages of the present invention will become apparent from the following description in connection with the accompanying drawings showing preferred embodiments of the present invention as examples.
Brief Description of the Drawings
[0009] [Figure 1] It is a diagram showing a configuration example and a processing example of a first information processing system. [Figure 2] It is a diagram showing a configuration example of an information processing system according to a second embodiment. [Figure 3] It is a diagram for explaining problems when a person moves within a frame. [Figure 4] It is a diagram showing a configuration example of a processing function provided in an information processing device. [Figure 5] It is a diagram showing an outline of a process for estimating global coordinates of each joint point. [Figure 6] It is a diagram for explaining a contact surface detection process. [Figure 7] It is a flowchart (Part 1) showing an example of a three-dimensional position estimation process by an information processing device. [Figure 8] It is a flowchart (Part 2) showing an example of a three-dimensional position estimation process by an information processing device.
Modes for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. 〔First Embodiment〕 FIG. 1 is a diagram showing a configuration example and a processing example of a first information processing system. The information processing system shown in FIG. 1 includes a camera 1 and an information processing device 2.
[0011] Camera 1 captures a moving image in which person 3 is shown. Also, camera 1 is movable, for example, by being carried by a photographer. The data of the captured moving image is input to information processing device 2. The captured moving image data may be input to information processing device 2 in real time at the time of shooting, or may be input to information processing device 2 at a timing asynchronous with shooting after being once stored in a storage device.
[0012] Information processing device 2 has a processing unit 2a. Processing unit 2a is, for example, a processor. Processing unit 2a executes a coordinate estimation process for estimating the first three-dimensional absolute coordinates of each of a plurality of joints included in person 3 for each of frames 4_1 to 4_N included in the input moving image. In this coordinate estimation process, processing unit 2a estimates the first position and orientation of camera 1 and estimates the first three-dimensional absolute coordinates based on the estimated first position and orientation.
[0013] The coordinate estimation process for each of frames 4_1 to 4_N includes the processes of the following steps S1 to S3. Here, as an example, frame 4_M (M ≤ N) will be described. [[ID=IO]]
[0014] [[ID=I2]]
[0015]
[0016] [Step S1] Processing unit 2a estimates the second position and orientation of camera 1 corresponding to frame 4_M based on the positional relationship of each joint between frame 4_M and the frame before it (here, the immediately preceding frame 4_M-1) and the first position and orientation estimated from frame 4_M-1. [Step S2] Processing unit 2a estimates the joint torque acting on each joint in frame 4_M based on the positional relationship of each joint between frame 4_M and frame 4_M-1. Processing unit 2a estimates the trajectory of each joint from frame 4_M-1 based on the estimated joint torque. This trajectory estimation is performed, for example, by motion simulation.[Step S3] The processing unit 2a estimates a first position and orientation corresponding to frame 4_M by correcting the second position and orientation estimated in step S1 based on the estimated trajectory of each joint in step S2. Based on the estimated first position and orientation, the processing unit 2a estimates a first three-dimensional absolute coordinate for each joint corresponding to frame 4_M.
[0017] In step S1, the second position and orientation of camera 1 is estimated based on the video footage captured by the movable camera 1. At this time, for example, three-dimensional positional information about the surrounding environment in which a person is present, or the positional information of camera 1 while it is moving, is not used. For this reason, the accuracy of the estimation of the second position and orientation cannot be said to be high.
[0018] For example, if the position of person 3 in frame 4_M moves upward, the processing unit 2a cannot determine whether the upward movement was due to an action such as jumping, or whether the downward movement was due to camera 1. Therefore, it is difficult to accurately estimate the second position and orientation. Consequently, it is difficult to accurately estimate the three-dimensional absolute coordinates of each joint using such a second position and orientation.
[0019] In contrast, the second position and posture are corrected based on the trajectory estimation of each joint in step S2. This corrects the second position and posture according to the movement of the person in space, thus improving the estimation accuracy of the corrected first position and posture. For example, in the case where the position of person 3 in frame 4_M moves upward, if the upward movement of person 3 in space is due to an action such as jumping, the second position and posture will not change due to the correction. On the other hand, if the position of camera 1 moves downward, the second position and posture will change due to the correction. Therefore, by using the first position and posture estimated with such high accuracy, the processing unit 2a can estimate the three-dimensional absolute position of each joint with high accuracy.
[0020] As described above, the information processing device 2 according to the first embodiment can accurately estimate the three-dimensional position of a person from a moving image captured by a moving camera. [Second Embodiment] Figure 2 shows an example of the configuration of an information processing system according to the second embodiment. The information processing system shown in Figure 2 includes an information processing device 100 and a camera 50 connected to the information processing device 100.
[0021] The information processing device 100 is an example of the information processing device 2 shown in Figure 1, and can be implemented as a computer such as a personal computer or a server device. The information processing device 100 estimates the three-dimensional coordinates of a person in a video based on the video captured by the camera 50. The camera 50 is positioned so that it can be moved by the photographer, and transmits the captured video data to the information processing device 100.
[0022] In this embodiment, the information processing device 100 receives the captured video data and estimates the three-dimensional coordinates of the person in real time. However, the captured video data may be temporarily stored in a storage device, and the information processing device 100 may estimate the three-dimensional coordinates of the person from the video data at a timing asynchronous to the capture.
[0023] Here, we will explain an example of the hardware configuration of the information processing device 100 using Figure 2. The information processing device 100 is implemented as a computer with a hardware configuration as shown in Figure 2, for example.
[0024] The information processing device 100 shown in Figure 2 includes a processor 101, RAM (Random Access Memory) 102, HDD (Hard Disk Drive) 103, GPU (Graphics Processing Unit) 104, input interface (I / F) 105, reading device 106, communication interface (I / F) 107, and network interface (I / F) 108.
[0025] The processor 101 comprehensively controls the entire information processing unit 100. The processor 101 is, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), or a PLD (Programmable Logic Device). Alternatively, the processor 101 may be a combination of two or more elements from among the CPU, MPU, DSP, ASIC, and PLD.
[0026] RAM 102 is used as the main memory of the information processing device 100. At least a portion of the OS (Operating System) program and application programs to be executed by the processor 101 are temporarily stored in RAM 102. In addition, various data necessary for processing by the processor 101 are stored in RAM 102.
[0027] HDD103 is used as an auxiliary storage device for the information processing device 100. The HDD103 stores the OS program, application programs, and various data. Other types of non-volatile storage devices, such as SSDs (Solid State Drives), can also be used as auxiliary storage devices.
[0028] A display device 104a is connected to the GPU 104. The GPU 104 displays images on the display device 104a according to instructions from the processor 101. The display device 104a can be an LCD display or an OLED (Electroluminescence) display, among others.
[0029] An input device 105a is connected to the input interface 105. The input interface 105 transmits signals output from the input device 105a to the processor 101. Examples of input devices 105a include keyboards and pointing devices. Examples of pointing devices include mice, touch panels, tablets, touchpads, and trackballs.
[0030] A portable recording medium 106a is attached to and detached from the reading device 106. The reading device 106 reads the data recorded on the portable recording medium 106a and transmits it to the processor 101. The portable recording medium 106a can be an optical disc, a semiconductor memory, or the like.
[0031] The communication interface 107 receives video data captured by the camera 50 and transmits it to the processor 101. The network interface 108 transmits and receives data with other devices via the network 108a. Alternatively, the video data from the camera 50 may be transmitted via the network 108a, and the network interface 108 may receive this data.
[0032] The processing functions of the information processing device 100 can be realized with the hardware configuration described above. Incidentally, the information processing device 100 estimates the position and orientation of the camera 50 (hereinafter sometimes referred to as "camera position and orientation") from the video image captured by the movable camera 50. Then, based on the estimated camera position and orientation, the information processing device 100 estimates the three-dimensional coordinates in the global coordinate system of the body parts (specifically, joint points) of the person shown in the video image.
[0033] This configuration makes it possible to estimate the three-dimensional position of a person from various types of captured images. For example, it becomes possible to estimate the three-dimensional position of a person's body parts from video footage of a person (athlete) in a sports broadcast, or from video footage of a person (performer) in a concert. Furthermore, it becomes possible to estimate the three-dimensional position of a person's body parts even from video footage captured by a smartphone at such events.
[0034] In recent years, metaverse-related activities have become increasingly active, such as viewing sports events and concerts on the metaverse. If the three-dimensional position of a person's body parts can be estimated using the method described above, it would be possible to use it, for example, to recreate real space on the metaverse and place an avatar corresponding to a photographed person in that recreated space.
[0035] However, when the camera 50 moves, it becomes difficult to accurately estimate the three-dimensional position of the person in space. In particular, as shown in Figure 3 below, there is a problem in that the three-dimensional position of the person cannot be accurately estimated when the person moves within the frame of the video.
[0036] Figure 3 illustrates the problems that arise when a person moves within a frame. In the example in Figure 3, the same person 60 is shown in consecutive frames 61 and 62. The position of person 60 is higher in frame 62 than in frame 61. That is, in frame 62, person 60 appears to have moved upwards relative to frame 61.
[0037] Such situations can occur, for example, in the following cases 1 and 2. In case 1, the person 60 is actually moving upward in space. For example, this could be the case when the person 60 jumps. On the other hand, in case 2, the person 60 is not moving upward, but the camera 50 is moving downward. When trying to estimate the three-dimensional position of the person 60 from the captured video, it is difficult to accurately estimate the three-dimensional position of the person 60 because it is impossible to determine whether the person 60 or the camera 50 has moved, as described above. Furthermore, in reality, both the person 60 and the camera 50 may be moving, making it even more difficult to accurately estimate the three-dimensional position of the person 60.
[0038] To address these problems, for example, it is possible to improve the accuracy of estimating the three-dimensional position of a person 60 by pre-providing three-dimensional positional information about the surrounding environment of the person 60 and detecting feature points of the surrounding environment from the captured image. However, this method requires pre-preparation of three-dimensional positional information about the surrounding environment, which limits its versatility. Furthermore, in sports broadcasts and concert footage, the focus is often solely on the person, making it impossible to detect feature points of the surrounding environment that are at a certain distance from the person. Another possible method is to detect the positional information of a moving camera 50, but this requires sensors for position detection, increasing the equipment cost.
[0039] Therefore, the information processing device 100 in this embodiment estimates the camera position and orientation based on the positional relationship of the joint points of a person in a moving image, and also estimates how forces are applied to the person based on that positional relationship. The information processing device 100 then optimizes the camera position and orientation and how forces are applied so that the predicted trajectory of the joint points in the global coordinate system, which is assumed when the estimated force is applied, matches the three-dimensional position of the joint points based on the camera position and orientation. This makes it possible to improve the accuracy of the camera position and orientation estimation, and by using the corrected camera position and orientation, it becomes possible to accurately estimate the three-dimensional coordinates of the joints in the global coordinate system.
[0040] Figure 4 shows an example of the configuration of processing functions provided by an information processing device. The information processing device 100 includes a storage unit 110 and a processing unit 120. The memory unit 110 is a memory area reserved in the memory device provided by the information processing device 100, such as the RAM 102 or HDD 103. Camera parameters 111, cinema information 112, and parameter set 113 are stored in the memory unit 110.
[0041] Camera parameter 111 includes information indicating the internal parameters of camera 50, such as the focal length and optical center of camera 50. The cinema information 112 is information that defines the joint model of a person. For example, the cinema information 112 includes information that shows the connection relationships and relative positional relationships of a predetermined number of joint points. The cinema information 112 may also include information that shows the length of the parts between adjacent joint points.
[0042] The parameter set 113 includes various parameters used in the processing of the motion simulator 124 and the optimization processing unit 125. For example, the parameter set 113 includes the coefficient of friction between the person and the contact surface, parameters for gradient calculation, and the convergence threshold in the optimization calculation.
[0043] The processing unit 120 includes a person detection unit 121, a contact surface detection unit 122, a camera position and pose estimation unit 123, a motion simulator 124, and an optimization processing unit 125. The processing unit 120 is, for example, a processor 101. In this case, the processing of the person detection unit 121, the contact surface detection unit 122, the camera position and pose estimation unit 123, the motion simulator 124, and the optimization processing unit 125 is achieved, for example, by the processor 101 executing a predetermined program.
[0044] The person detection unit 121 detects joint points included in the above joint model from the frames of the moving image. The person detection unit 121 also estimates the three-dimensional coordinates of each detected joint point based on the positions of multiple joint points on the frame image, the relative positional relationships between the joint points based on the cinema information 112, and the camera parameters 111. The estimated three-dimensional coordinates are the coordinates in the camera coordinate system.
[0045] The contact surface detection unit 122 detects contact surfaces that come into contact with a person from within the frame. For example, the contact surface detection unit 122 detects the ground or floor that comes into contact with the joint point corresponding to the ankle as a contact surface. The contact surface detection unit 122 tracks the detected contact surfaces and estimates the three-dimensional coordinates of the contact surfaces based on the positional relationship of the contact surfaces between frames. The estimated three-dimensional coordinates become the coordinates in the camera coordinate system.
[0046] The camera position and orientation estimation unit 123 estimates the camera position and orientation based on the camera position and orientation estimated in the previous frame, the positional relationship of corresponding feature points between the previous frame and the current frame, and the camera parameters 111. For example, joint points detected by the person detection unit 121 are used as feature points. Furthermore, by using the camera position and orientation estimated in the previous frame, a time-series prediction is performed from the initial position and orientation, thereby estimating the camera position and orientation in the global coordinate system.
[0047] Furthermore, the camera position and orientation estimation unit 123 uses the camera position and orientation estimation results to convert the three-dimensional coordinates of each joint point estimated by the person detection unit 121 and the three-dimensional coordinates of the contact surface estimated by the contact surface detection unit 122 into three-dimensional coordinates in the global coordinate system. This allows the global coordinates of each joint point to be estimated based on the camera position and orientation.
[0048] The motion simulator 124 estimates mechanical parameters that represent the forces acting on a person based on the positional relationships of joint points and contact surfaces between frames. The estimated mechanical parameters include joint torques acting on each joint point.
[0049] Furthermore, the motion simulator 124 estimates the three-dimensional coordinates of each joint point in the current frame in the global coordinate system based on the three-dimensional coordinates of each joint point and contact surface in the global coordinate system estimated in the previous frame, and the estimated dynamic parameters. This allows the global coordinates of each joint point to be estimated based on the simulation. The simulation is performed using a differentiable simulator (equations of motion) to enable gradient calculation in the optimization process.
[0050] The optimization processing unit 125 optimizes the camera position and orientation and joint torques so that the error between the global coordinates of each joint point estimated based on the camera position and orientation and the global coordinates of each joint point estimated based on the simulation is minimized. The optimization processing unit 125 outputs the global coordinates of each joint point in the current frame using the optimal values for the camera position and orientation.
[0051] Figure 5 shows an overview of the process for estimating the global coordinates of each joint point. In the example in Figure 5, it is assumed that frame 72 is the target of processing among the consecutive frames 71 and 72. Although not shown, the person detection unit 121 detects the two-dimensional coordinates of each joint point of the person from frame 72, and estimates the three-dimensional coordinates of each detected joint point in the camera coordinate system.
[0052] In this state, the camera position and pose estimation unit 123 estimates the camera position and pose corresponding to frame 72 based on the camera position and pose estimated from frame 71 by the optimization processing unit 125, the positional relationship of each joint point between frames 71 and 72, and the camera parameters 111 (step S11). Using the estimated camera position and pose, the camera position and pose estimation unit 123 converts the three-dimensional coordinates of each joint point estimated by the person detection unit 121 into global coordinates (step S12).
[0053] Meanwhile, the motion simulator 124 estimates the joint torque acting on each joint point based on the positional relationship of the joint points between frames 71 and 72 (step S13). Based on the global coordinates of each joint point estimated from frame 71 and the estimated joint torque, the motion simulator 124 estimates the global coordinates of each joint point in frame 72 through simulation (step S14). This allows the simulator to estimate where the position of each joint point in the global coordinate system in frame 71 will move to in frame 72 when the joint torque estimated in step S13 occurs.
[0054] The optimization processing unit 125 compares the global coordinates estimated in step S12 with the global coordinates estimated in step S14. The optimization processing unit 125 optimizes the camera position, orientation, and joint torques so as to minimize the error in the global coordinates for each joint point (step S15). This calculates the optimal values for the camera position, orientation, and joint torques that bring them closer to the estimated global coordinates in steps S12 and S14. This process estimates the camera position and orientation by reflecting not only geometric factors but also factors related to how forces are applied to the person. As a result, it becomes possible to accurately determine the movement of the camera 50 and the movement of the person and estimate the camera position and orientation.
[0055] The optimization processing unit 125 estimates the global coordinates of each joint point based on the camera position and orientation optimized in step S15 (step S16). This improves the accuracy of estimating the global coordinates of each joint point.
[0056] Here, in the estimation steps S13 and S14, the estimation accuracy can be improved by also using the position of the contact surface that comes into contact with the person. As a contact surface, for example, the surface of an object that comes into contact with the lower end of the person, such as the ground or floor, is detected.
[0057] Figure 6 is a diagram illustrating the contact surface detection process. The left side of Figure 6 shows the state of the previous frame, and the right side of Figure 6 shows the state of the current frame. When the contact surface detection unit 122 detects a contact surface in a given frame, it continues to detect the position of the contact surface in subsequent frames. For example, the contact surface detection unit 122 tracks multiple feature points included in the detected contact surface. This tracking is performed in parallel with the joint point detection process by the person detection unit 121.
[0058] In the example in Figure 6, in the previous frame, a contact surface 91 that is in contact with the ankle joints 81 and 82 of the person is detected. The detected contact surface 91 is also detected in the current frame. In the example in Figure 6, in the current frame, by comparing the three-dimensional coordinates of the ankle joints 81 and 82 in the camera coordinate system with the three-dimensional coordinates of the contact surface 91, it is detected that the ankle joints 81 and 82 have moved away from the contact surface 91.
[0059] The motion simulator 124 can calculate the distance between a joint point and a contact surface for each frame based on the estimation results of the three-dimensional coordinates of the joint points by the person detection unit 121 and the estimation results of the three-dimensional coordinates of the contact surface by the contact surface detection unit 122. Therefore, the motion simulator 124 can estimate not only joint torque but also reaction force and friction force from the contact surface to the joint point as motion parameters. Consequently, by using the three-dimensional coordinates of the contact surface, the motion simulator 124 can estimate the global coordinates of each joint point with high accuracy. For example, when a person moves upward in frame 62 as shown in Figure 3, it can accurately determine whether the person moved upward from the lower contact surface.
[0060] Furthermore, depending on the field of view of the camera 50, the contact surface detection unit 122 may not be able to continue detecting the same contact surface in all subsequent frames after the initial detection. For example, there may be cases where the contact surface temporarily moves out of the shooting range, causing the contact surface detection unit 122 to temporarily lose detection of the contact surface.
[0061] Therefore, if the contact surface detection unit 122 can no longer detect the contact surface, for example, it estimates the position of the contact surface in the current frame by interpolation calculation based on the detection results from past frames. Furthermore, if the contact surface detection unit 122 has not detected the contact surface up to the current frame, it may acquire positional information of external objects that could come into contact with the person from past frames and estimate the position of the contact surface in the current frame by interpolation calculation. Also, if the processing shown in Figure 5 is performed in offline processing, the contact surface detection unit 122 may detect the contact surface that will come into contact with the person from future frames and estimate the position of the contact surface in the current frame by interpolation calculation. In addition, if the contact surface is not visible in any frame, the contact surface detection unit 122 may estimate the distance between the contact surface and the person based on the person's height and posture.
[0062] Next, the processing of the information processing device 100 will be explained using a flowchart. Figures 7 and 8 are flowcharts showing examples of three-dimensional position estimation processing by an information processing device.
[0063] [Step S21] The person detection unit 121 acquires frame data from the input video. [Step S22] The person detection unit 121 detects multiple joint points of a person from the frame and obtains their two-dimensional coordinates.
[0064] [Step S23] The person detection unit 121 estimates the three-dimensional coordinates of each detected joint point based on the two-dimensional coordinates of each joint point, the relative positional relationship between the joint points based on the cinema information 112, and the camera parameters 111. The estimated three-dimensional coordinates are the coordinates in the camera coordinate system.
[0065] [Step S24] The contact surface detection unit 122 detects the contact surface that comes into contact with the person from the frame. The contact surface detection unit 122 estimates the three-dimensional coordinates of the contact surface based on the positional relationship of the contact surface between the previous frame and the current frame. The estimated three-dimensional coordinates are coordinates in the camera coordinate system. In practice, for example, the three-dimensional coordinates of each feature point included in the current frame are estimated based on the positional relationship of corresponding feature points between frames among the feature points included in the contact surface.
[0066] [Step S25] The camera position and orientation estimation unit 123 estimates the camera position and orientation corresponding to the current frame based on the camera position and orientation estimated from the previous frame by the optimization processing unit 125, the positional relationship of each joint point between the current frame and the previous frame, and the camera parameters 111. Since time-series prediction from the initial position and orientation is performed using the camera position and orientation estimated from the previous frame, the camera position and orientation in the global coordinate system is estimated. Furthermore, the estimated camera position and orientation becomes the initial parameter for the camera position and orientation in the optimization process.
[0067] [Step S26] The camera position and orientation estimation unit 123 uses the estimated camera position and orientation to convert the three-dimensional coordinates of each joint point estimated by the person detection unit 121 into global coordinates. The camera position and orientation estimation unit 123 also uses the estimated camera position and orientation to convert the three-dimensional coordinates of the contact surface estimated by the contact surface detection unit 122 into global coordinates.
[0068] [Step S27] The motion simulator 124 acquires the three-dimensional coordinates of each joint point for the current frame, which were estimated by the person detection unit 121 in step S23. The motion simulator 124 also acquires the three-dimensional coordinates of the contact surface for the current frame, which were estimated by the contact surface detection unit 122 in step S24.
[0069] Furthermore, the motion simulator 124 acquires the three-dimensional coordinates of each joint point for the previous frame, which were estimated by the person detection unit 121 in step S23 for the previous frame. Also, the motion simulator 124 acquires the three-dimensional coordinates of the contact surface, which were estimated by the contact surface detection unit 122 in step S24 for the previous frame.
[0070] The motion simulator 124 uses the acquired three-dimensional coordinates to estimate dynamic parameters representing the forces acting on the figure in the current frame, based on the positional relationship of corresponding joint points and contact surfaces between the current frame and the previous frame. These dynamic parameters include, for example, the joint torque acting on each joint point, the reaction force from the contact surface to the figure, and the frictional force between the contact surface and the figure. The estimated joint torque serves as the initial parameter for joint torque in the optimization process.
[0071] [Step S28] The motion simulator 124 obtains the three-dimensional coordinates of each joint point and contact surface in the global coordinate system estimated by the optimization processing unit 125 for the previous frame, and the dynamic parameters corresponding to the current frame estimated in step S27. Based on the acquired information, the motion simulator 124 estimates the global coordinates of each joint point of the current frame by simulation.
[0072] Next, in steps S29 to S32, the optimization processing unit 125 performs optimization processing. [Step S29] The optimization processing unit 125 generates an objective function f for optimization. The objective function f can be expressed, for example, by the following equation (1).
[0073]
number
[0074] In equation (1), T is the position and attitude matrix that represents the camera position and attitude estimated in step S25. n This shows the joint torque estimated in step S27 for the nth joint point.n represents the three-dimensional coordinates estimated for the n-th joint point in step S26. r n represents the three-dimensional coordinates estimated for the n-th joint point in step S27. r n ({F m} m ) means that when the joint torques estimated for m joint points (all joint points) act, the three-dimensional coordinates r of the n-th joint point are obtained by solving the equation of motion. Note that r n is obtained. Note that r n ({F m} m ) is calculated by the motion simulator 124. [[ID=If step S29 is executed after step S28, the values of the camera position and orientation and joint torque, which are the parameters to be optimized, are set as initial values estimated in steps S25 and S27, respectively. On the other hand, if step S29 is executed after step S31, the values of the camera position and orientation and joint torque are set as the values updated in the most recent step S30.
[0078] The optimization processing unit 125 performs gradient calculation of the objective function f. For example, the Jacobian matrix is used in this gradient calculation. [Step S30] The optimization processing unit 125 obtains the change amounts of the camera position and orientation and joint torque, respectively, based on the results of the gradient calculation in step S29. The optimization processing unit 125 adds the obtained change amounts to the current camera position and orientation and joint torque values and updates the camera position and orientation and joint torque values.
[0079] [Step S31] The optimization processing unit 125 inputs the camera position and orientation and joint torque values updated in step S30 into equation (1) and recalculates equation (1). At this time, the global coordinates of each joint point and contact surface based on the updated camera position and orientation are recalculated by the camera position and orientation estimation unit 123 in accordance with instructions from the optimization processing unit 125. n ({F m} m The calculation of ) is re-executed by the motion simulator 124 in accordance with instructions from the optimization processing unit 125.
[0080] The optimization processing unit 125 determines whether the calculation of equation (1) has converged. For example, if the calculated value of equation (1) is less than or equal to a predetermined threshold, it is determined that the calculation has converged. If the calculation has not converged, the process proceeds to step S29; if the calculation has converged, the process proceeds to step S32.
[0081] [Step S32] The optimization processing unit 125 finally outputs the global coordinates of each joint point calculated in step S30 (i.e., the global coordinates of each joint point estimated using the optimal values of the camera position and orientation) as the global coordinates of each joint point corresponding to the frame acquired in step S21.
[0082] [Step S33] The optimization processing unit 125 determines whether the frame acquired in step S21 is the final frame in the video. If it is not the final frame, the process proceeds to step S21 and the next frame is acquired. On the other hand, if it is the final frame, the process ends.
[0083] Furthermore, the optimization process described above also optimizes joint torque in addition to camera position and orientation. This reduces the impact of calculation errors in joint torque and improves the accuracy of estimating the global coordinates of each joint point. For example, depending on the orientation of a person in the frame, one joint point may be hidden by another, making it impossible to detect the former, which can reduce the accuracy of joint torque estimation. By optimizing joint torque as described above, the impact of such a decrease in estimation accuracy on the accuracy of estimating the global coordinates of each joint point can be reduced.
[0084] Furthermore, the processing of the information processing device 100, as described above, involves simulation during the processing, resulting in the following effects. For example, when displaying a virtual person using the estimated global coordinates of each joint point, it becomes possible to manipulate the person's movements and try executing different movements or scenarios that differ from reality midway through the process. It also becomes possible to train such a virtual person to learn the movements of an athlete's body parts. The results of such training can then be used, for example, to pass on or preserve an athlete's skills.
[0085] Furthermore, the processing functions of the devices shown in each of the above embodiments (for example, the information processing device 2,100) can be implemented by a computer. In this case, a program describing the processing content of the functions that each device should have is provided, and by executing that program on the computer, the above processing functions are implemented on the computer. The program describing the processing content can be recorded on a computer-readable recording medium. Computer-readable recording mediums include magnetic storage devices, optical discs, and semiconductor memory. Magnetic storage devices include hard disk drives (HDDs) and magnetic tapes. Optical discs include CDs (Compact Discs), DVDs (Digital Versatile Discs), and Blu-ray Discs (BD, registered trademark).
[0086] When distributing a program, portable recording media such as DVDs and CDs containing the program are sold. Alternatively, the program can be stored in the storage device of a server computer and transferred from the server computer to other computers via a network.
[0087] A computer executing a program stores programs, for example, those recorded on a portable storage medium or transferred from a server computer, in its own memory. The computer then reads the program from its memory and executes the processing according to the program. Alternatively, the computer can directly read the program from the portable storage medium and execute the processing according to that program. Furthermore, the computer can sequentially execute the processing according to the programs received from a server computer connected via a network, each time a program is transferred.
[0088] The above merely illustrates the principle of the present invention. Furthermore, numerous modifications and changes are possible for those skilled in the art, and the present invention is not limited to the exact configurations and applications shown and described above. All corresponding modifications and equivalents are considered to be within the scope of the present invention as defined by the appended claims and their equivalents. [Explanation of symbols]
[0089] 1 Camera 2. Information Processing Device 2a Processing Unit 3 people 4_1,4_M-1,4_M,4_N Frame S1~S3 Steps
Claims
1. Computers For each of the multiple frames included in a video image of a person captured by the camera, a coordinate estimation process is performed to estimate the first position and orientation of the camera, and to estimate the first three-dimensional absolute coordinates for each of the multiple joints included in the person based on the first position and orientation. The coordinate estimation process for the first frame among the plurality of frames is as follows: A first process for estimating a second position and orientation of the camera corresponding to the first frame, based on the positional relationship of the plurality of joints between the first frame and a second frame preceding the first frame, and the first position and orientation estimated from the second frame, A second process that estimates the joint torque acting on each of the plurality of joints in the first frame based on the positional relationship, and estimates the trajectory of each of the plurality of joints from the second frame based on the joint torque, A third process involves estimating the first position and orientation corresponding to the first frame by correcting the second position and orientation based on the trajectory estimation result, and estimating the first three-dimensional absolute coordinates for each of the plurality of joints corresponding to the first frame based on the estimated first position and orientation. including, Location estimation method.
2. In the first processing for the first frame, based on the positional relationship and the second positional orientation, a second three-dimensional absolute coordinate is estimated for each of the plurality of joints corresponding to the first frame. In the second process for the first frame, a third three-dimensional absolute coordinate is estimated for each of the multiple joints corresponding to the first frame, based on the first three-dimensional absolute coordinate and the joint torque estimated from the second frame. In the third processing for the first frame, the second position and orientation are optimized so that the error between the second three-dimensional absolute coordinates and the third three-dimensional absolute coordinates is reduced, and the optimized second position and orientation are estimated as the first position and orientation. The position estimation method according to claim 1.
3. The coordinate estimation process for the first frame further includes a fourth process of detecting contact objects that are in contact with the person or that were in contact in previous frames from the first frame, and estimating the three-dimensional position of the contact objects. In the second process for the first frame, the third three-dimensional absolute coordinates are estimated based on the first three-dimensional absolute coordinates, the joint torque, and the three-dimensional position of the contact object estimated from the second frame. The position estimation method according to claim 2.
4. In the third processing for the first frame, the second position and orientation and the joint torque are optimized so that the error between the second three-dimensional absolute coordinates and the third three-dimensional absolute coordinates is reduced. The position estimation method according to claim 2.
5. On the computer, For each of the multiple frames contained in a video image of a person captured by the camera, a coordinate estimation process is performed to estimate the first position and orientation of the camera, and to estimate the first three-dimensional absolute coordinates for each of the multiple joints contained in the person based on the first position and orientation. The coordinate estimation process for the first frame among the plurality of frames is as follows: A first process for estimating a second position and orientation of the camera corresponding to the first frame, based on the positional relationship of the plurality of joints between the first frame and a second frame preceding the first frame, and the first position and orientation estimated from the second frame, A second process that estimates the joint torque acting on each of the plurality of joints in the first frame based on the positional relationship, and estimates the trajectory of each of the plurality of joints from the second frame based on the joint torque, A third process involves estimating the first position and orientation corresponding to the first frame by correcting the second position and orientation based on the trajectory estimation result, and estimating the first three-dimensional absolute coordinates for each of the plurality of joints corresponding to the first frame based on the estimated first position and orientation. including, A location estimation program.
6. The device has a processing unit that performs coordinate estimation processing to estimate a first position and orientation of the camera for each of a plurality of frames included in a video image of a person captured by the camera, and to estimate a first three-dimensional absolute coordinate for each of a plurality of joints included in the person based on the first position and orientation. The coordinate estimation process for the first frame among the plurality of frames is as follows: A first process for estimating a second position and orientation of the camera corresponding to the first frame, based on the positional relationship of the plurality of joints between the first frame and a second frame preceding the first frame, and the first position and orientation estimated from the second frame, A second process that estimates the joint torque acting on each of the plurality of joints in the first frame based on the positional relationship, and estimates the trajectory of each of the plurality of joints from the second frame based on the joint torque, A third process involves estimating the first position and orientation corresponding to the first frame by correcting the second position and orientation based on the trajectory estimation result, and estimating the first three-dimensional absolute coordinates for each of the plurality of joints corresponding to the first frame based on the estimated first position and orientation. including, Information processing device.