3D position calculation method, device, and program
By interpolating and clustering two-dimensional positions from multiple images, the method enhances the accuracy of three-dimensional position calculation, addressing errors from undetected or falsely detected objects in multi-viewpoint images.
Patent Information
- Application Number
- JP2024530249
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Existing methods for calculating the three-dimensional position of an object using multi-viewpoint images are inaccurate due to undetected or falsely detected objects, leading to errors in triangulation.
The method involves acquiring images before and after a target time from multiple cameras, interpolating the two-dimensional positions, and using camera parameters to calculate the three-dimensional position, with outlier removal through clustering and center of gravity calculation.
This approach improves the accuracy of three-dimensional position calculation by supplementing missing or erroneous two-dimensional data, ensuring precise positioning even with outliers.
Smart Images

Figure 0007754315000007 
Figure 0007754315000008 
Figure 0007754315000009
Abstract
Description
[Technical Field]
[0001] The disclosed technology relates to a three-dimensional position calculation method, a three-dimensional position calculation device, and a three-dimensional position calculation program. [Background technology]
[0002] Conventionally, triangulation has been used to calculate the three-dimensional position of an object in a world coordinate system from the two-dimensional position of the object in multi-view images of the object captured from multiple different viewpoints. For example, a system has been proposed that uses two or more cameras to track the path and orientation of a portion of portable sports equipment swung by an athlete. In this system, at least two sets of video images of the portable sports equipment being swung are acquired using at least two different cameras with different positions. Then, motion regions within the video images are identified, and candidate positions in two-dimensional space of an identifiable portion of the portable sports equipment (e.g., the head) are identified within the motion regions. Based on this, possible positions in three-dimensional space of the identifiable portion are identified for each of multiple moments when the portable sports equipment is swung. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] U.S. Patent Application Publication No. 2017 / 0270354 Summary of the Invention [Problem to be solved by the invention]
[0004] However, if the multi-view images include images in which the object is not detected or images in which the object is falsely detected, there is a problem that the three-dimensional position of the object may not be calculated accurately.
[0005] According to one aspect, the disclosed technology aims to improve the accuracy of calculating the three-dimensional position of an object using multi-viewpoint images. [Means for solving the problem]
[0006] In one aspect, the disclosed technology acquires a first image captured at one or more times before a target time and a second image captured at the target time using each of multiple cameras that capture images of an object from different viewpoints. The disclosed technology also acquires a third image captured at one or more times after the target time. The disclosed technology then interpolates the two-dimensional position of the object in the second image based on the two-dimensional positions of the object detected from each of the first and third images. The disclosed technology also calculates the three-dimensional position of the object at the target time based on the two-dimensional position of the object detected from the second image, the interpolated two-dimensional position of the object in the second image, and camera parameters of each of the multiple cameras. [Effects of the Invention]
[0007] As one aspect, it has an effect of improving the accuracy of calculating the three-dimensional position of an object using multi-viewpoint images. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a schematic diagram showing a connection between a three-dimensional position calculation device and a camera according to the present embodiment. [Figure 2] FIG. 1 is a diagram for explaining a general method for calculating a three-dimensional position from a multi-viewpoint image. [Figure 3] FIG. 1 is a diagram for explaining a general method for calculating a three-dimensional position from a multi-viewpoint image. [Figure 4] FIG. 10 is a diagram illustrating the removal of outliers. [Figure 5] 1A and 1B are diagrams for explaining problems with a general method for calculating a three-dimensional position from multi-viewpoint images. [Figure 6]1 is a functional block diagram of a three-dimensional position calculation device according to the present embodiment. [Figure 7] FIG. 2 is a diagram for explaining an example of a two-dimensional position of an object. [Figure 8] FIG. 2 is a diagram for explaining an example of a two-dimensional position of an object. [Figure 9] FIG. 10 is a diagram for explaining interpolation of two-dimensional positions using time-space information. [Figure 10] FIG. 10 is a diagram for explaining calculation of a three-dimensional position based on clustering of three-dimensional position candidates. [Figure 11] FIG. 1 is a block diagram showing a schematic configuration of a computer that functions as a three-dimensional position calculation device. [Figure 12] 10 is a flowchart showing an example of a three-dimensional position calculation process according to the present embodiment. [Figure 13] FIG. 10 is an image diagram showing an example of a calculation result of a three-dimensional position according to this embodiment. [Figure 14] FIG. 10 is an image diagram showing an example of a calculation result of a three-dimensional position according to this embodiment. [Figure 15] FIG. 1 is a diagram for explaining application of a three-dimensional position calculation device according to the present embodiment to a scoring system for gymnastics competitions. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an example of an embodiment of the disclosed technology will be described with reference to the drawings.
[0010] As shown in Fig. 1, a three-dimensional position calculation device 10 according to this embodiment is connected to a plurality of cameras 30n that capture images of an object 90 (a person in the example of Fig. 1) at viewpoint n from each of different directions. In the example of Fig. 1, n = 0, 1, 2, and a camera 300 that captures images from viewpoint 0, a camera 301 that captures images from viewpoint 1, and a camera 302 that captures images from viewpoint 2 are connected to the three-dimensional position calculation device 10. Note that the number of cameras 30n connected to the three-dimensional position calculation device 10 is not limited to the example of Fig. 1, and may be two, four or more.
[0011] The camera 30n is installed at an angle and position such that the target object 90 falls within the shooting range. The images captured by the camera 30n are sequentially input to the three-dimensional position calculation device 10. A synchronization signal is sent to each camera 30n, and the images captured by each camera 30n are synchronized.
[0012] Here, a general method and problems involved in calculating the three-dimensional position of an object 90 from a plurality of images taken from a plurality of different viewpoints (hereinafter referred to as "multi-viewpoint images") using triangulation will be described.
[0013] As shown in FIG. 2, the two-dimensional position of the object detected in the image 40n taken from the viewpoint n is denoted by p 2d,obs cn (white circle in Figure 2), the true 2D position is p 2d,gt cn In the example of Figure 2, n = 0, 1, 2. The calculated three-dimensional position is p^ 3d (black circle in Figure 2), the true 3D position is p 3d,gt (The black star in Figure 2). In Figure 2, "p^" is written as "^ (hat)" above "p". This is the same in the following figures. 2d =[x,y]∈real number R 2 , p 3d = [X,Y,Z]∈real number R 3 The three-dimensional position p^ of the object 3d is the camera parameters (internal and external parameters) of the camera 30n that captured each image 40n, and the two-dimensional position p 2d,obs cn It is calculated by triangulation using
[0014] Here, for example, the two-dimensional position p of the object 90 detected from the image 400 2d,obs c0 If the detection error of is large, that is, p 2d,obs c0 and p 2d,gt c0 If the difference is large, the calculated three-dimensional position p^ 3d The error of p^ also becomes larger.3d noise However, the detection error is large p 2d,obs c0 In this case, the 2D position p 2d,obs c0 is excluded as an outlier, and p 2d,obs c1 and p 2d,obs c2 Using the three-dimensional position p^ 3d refine By calculating the true 3D position p 3d,gt It is desirable to calculate the three-dimensional position so that it approaches
[0015] A more specific explanation will be given using another example shown in Fig. 3. In the example of Fig. 3, n = 0, 1, 2, 3. The superscript t of each symbol indicates the time when the image 40n was captured, that is, the time information associated with the image (frame) 40n. In the example of Fig. 3, the two-dimensional position p of the object can be determined from the image 402. 2d,obs t,c2 has gone undetected.
[0016] 3D position of the object p^ 3d t is calculated as shown below using the function cv::sfm::triangulatePoints implemented in OpenCV (Reference: https: / / docs.opencv.org / 3.4 / d0 / dbd / group__triangulation.html). cn is a perspective projection matrix representing the camera parameters of the camera 30n.
[0017]
number
[0018] The two-dimensional position p of the detected object 2d,obs t,cn If there is an outlier in the 3d t and the true 3D position p 3d,gtt The error between p and 2d,obs t,cn In order to remove outliers from the 3D position p̂, for example, RANSAC (Random Sample Consensus) is applied. First, as shown in Figure 4, the 3D position p̂ calculated without removing outliers is 3d t Then, the projected two-dimensional position and p 2d,obs t,cn The error (indicated by the double arrow in Figure 4) is calculated, and if the error is greater than a predetermined threshold, the p 2d,obs t,cn In the example in Figure 4, p 2d,obs t,c0 is an outlier. This outlier p 2d,obs t,c0 is excluded and the three-dimensional position is calculated again.
[0019] The problem here is that many p 2d,obs t,cn If is excluded as an outlier, p 2d,obs t,cn In such cases, as shown in Figure 5, there may be insufficient information on the two-dimensional position for calculating the three-dimensional position, making it impossible to calculate the three-dimensional position with high accuracy.
[0020] Therefore, the three-dimensional position calculation device 10 according to this embodiment calculates the three-dimensional position of the object by using spatiotemporal information, specifically, information on images taken at times before and after the image taken at the target time. The three-dimensional position calculation device 10 according to this embodiment will be described in detail below.
[0021] 6, the three-dimensional position calculation device 10 functionally includes an acquisition unit 12, an interpolation unit 14, and a calculation unit 16. A camera parameter DB (Database) 20 is stored in a predetermined storage area of the three-dimensional position calculation device 10. The camera parameter DB 20 stores internal parameters and external parameters of each camera 30n.
[0022] The acquisition unit 12 acquires time-series multi-viewpoint images captured by a plurality of cameras 30n. Here, in the time-series multi-viewpoint images, an image captured at a time t (t=0, 1, . . . , T, where T is time information of the final frame) to be processed to calculate the three-dimensional position of the object 90 is defined as image 40n(t). Also, an image captured at time t-1, one time before time t, is defined as image 40n(t-1), and an image captured at time t+1, one time after time t, is defined as image 40n(t+1). Image 40n(t-1) is an example of a "first image" in the disclosed technology, image 40n(t) is an example of a "second image" in the disclosed technology, and image 40n(t+1) is an example of a "third image" in the disclosed technology.
[0023] Furthermore, information on the two-dimensional position of the object 90 is assigned to each image 40n included in the multi-view image. The information on the two-dimensional position of the object 90 may be the coordinate values of a predetermined point within a region surrounding the object 90 detected from each image 40n included in the multi-view image using a detection model previously generated by machine learning to detect the region of the object 90 from the image 40n. For example, as shown in FIG. 7, if the region of the object 90 is detected by a two-dimensional bounding box (hereinafter referred to as "2D-BBOX") 42n, the coordinates of a predetermined position in the 2D-BBOX 42n may be used as information on the two-dimensional position of the object 90. The predetermined position may be, for example, the center, the midpoint of the base, or one of the corners (e.g., the upper left corner) of the 2D-BBOX 42n. In the example of FIG. 7, the midpoint of the base of the 2D-BBOX 42n (black circle in FIG. 7) is used as the two-dimensional position of the object 90. This is treated as information indicating the position of the feet of a person, who is the object 90.
[0024] Furthermore, the information on the two-dimensional position of the object 90 may be the coordinate values of each part of the object 90 recognized from each image 40n included in the multi-view image using a recognition model that has been generated in advance by machine learning to recognize one or more parts of the person who is the object 90 from the image 40n. For example, as shown in FIG. 8, when the positions of each joint, etc. of the person who is the object 90 (black circles in FIG. 8) are recognized by the recognition model, the coordinate values of the positions of each joint, etc. may be used as the information on the two-dimensional position of the object 90.
[0025] In addition, when acquiring a multi-view image that does not have information on the two-dimensional position of the object 90, the acquisition unit 12 may acquire information on the two-dimensional position of the object 90 using the above-mentioned detection model or recognition model.
[0026] The interpolation unit 14 interpolates the two-dimensional position of the object 90 in the image 40n(t) based on the two-dimensional positions of the object 90 detected from each of the images 40n(t-1) and 40n(t+1). Specifically, the interpolation unit 14 predicts and interpolates the two-dimensional position of the object 90 in the image 40n(t) by linearly interpolating the two-dimensional positions of the object 90 detected from each of the images 40n(t-1) and 40n(t+1).
[0027] A specific description will be given using the example of Fig. 9. In the example of Fig. 9, n = 0, 1, and the two-dimensional position of the object 90 is detected from each of the image 40n(t-1), image 40n(t), and image 40n(t+1) as shown below.
[0028]
number
[0029] In this case, the interpolator 14 calculates the interpolated two-dimensional position p of the object 90 in each image 40n(t) as follows: 2d,pred t,cn (The dotted circle in Figure 9) is calculated.
[0030]
number
[0031] The time points of the images used for interpolation are not limited to t-1 and t+1. For example, the interpolation unit 14 may interpolate the two-dimensional position of the image 40n(t) using the images 40n at times t-5, t-4, t-3, t-2, t-1, t+1, t+2, t+3, t+4, and t+5. In this case, the images 40n(t-5), 40n(t-4), 40n(t-3), 40n(t-2), and 40n(t-1) are examples of the "first image" of the disclosed technology. The images 40n(t+1), 40n(t+2), 40n(t+3), 40n(t+4), and 40n(t+5) are examples of the "third image" of the disclosed technology.
[0032] The calculation unit 16 calculates the three-dimensional position of the object 90 at time t based on the two-dimensional position of the object 90 detected from the image 40n(t), the two-dimensional position interpolated by the interpolation unit 14, and the camera parameters of each camera 30n. Specifically, the calculation unit 16 calculates candidates for the three-dimensional position of the object 90 for each combination of the detected and interpolated two-dimensional positions of the object 90 between the images 40n(t).
[0033] More specifically, the calculation unit 16 selects one of the two-dimensional positions detected and interpolated from the image 40i captured by the camera 30i and calculates p 2d i Then, one of the two-dimensional positions detected and interpolated from the image 40j taken by the camera 30j is selected and p 2d j The calculation unit 16 calculates p 2d i =(x i ,y i ,1) and p 2d j =(x j ,y j , 1), a perspective projection matrix P representing the camera parameters of each of the cameras 30i and 30j ci and P cj By solving the following equation using 3d,cand t Calculate P n is the nth row of P.
[0034]
number
[0035] In the example of FIG. 9, the calculation unit 16 calculates p 2d,obs t,c0 and p 2d,pred t,c0 And, image 401 p 2d,obs t,c1 and p 2d,pred t,c1 From the combination of , four 3D position candidates p 3d,cand t (The shaded circle in Figure 9) is calculated.
[0036] The calculation unit 16 calculates the plurality of three-dimensional position candidates p 3d,cand t Based on this, the three-dimensional position p^ of the object 90 3d t For example, the calculation unit 16 calculates a plurality of candidates p 3d,cand t The position of the center of gravity is the three-dimensional position p^ 3d t In this case, the two-dimensional position p detected from the image 40n(t) is calculated as follows. 2d,obs t,cn , and the interpolated two-dimensional position p 2d,pred t,cn Therefore, the calculation unit 16 calculates the candidate p 3d,cand t Based on the distance between the candidates p 3d,cand t Clustering is performed and candidates p 3d,cand t The center of gravity of the cluster with the largest number of 3d t This assumes that the position of the center of gravity of the largest cluster has the highest probability of approximating the true three-dimensional position. Note that a hierarchical clustering method such as the complete linkage method may be applied as the clustering method.
[0037] A specific description will be given using the example of Fig. 10. In the example of Fig. 10, n = 0, 1, 2, 3, and the two-dimensional position of the object 90 is detected from each of the images 40n(t-1), 40n(t), and 40n(t+1) as shown below.
[0038]
number
[0039] In the example of FIG. 10, the interpolation unit 14 calculates the interpolated two-dimensional position p of the object 90 in each image 40n(t) as follows: 2d,pred t,cn (The dotted circle in Figure 10) has been calculated.
[0040]
number
[0041] The calculation unit 16 calculates two candidates p 3d,cand t If the distance between the two candidates p is smaller than a threshold K, 3d,cand t The calculation unit 16 also calculates the center of gravity of the cluster and the other candidates p 3d,cand t If the distance between p and p is less than the threshold K, 3d,cand t The calculation unit 16 performs this process by assigning candidates p whose distance is smaller than the threshold K to the corresponding cluster. 3d,cand Between clusters or candidates p 3d,cand This is repeated until there is no gap between the clusters. In the example of Fig. 10, six clusters (solid ellipses and dashed ellipses in Fig. 10) are generated by clustering.
[0042] The calculation unit 16 calculates the candidate p included in the generated cluster. 3d,cand tThe cluster with the largest number of clusters is selected, and the position of the center of gravity of the selected cluster is calculated as the three-dimensional position p^ of the object 90 at time t. 3d t In the example of FIG. 10, the cluster indicated by the solid ellipse is selected. As a result, the candidates p belonging to other clusters are calculated as follows. 3d,cand t is excluded as an outlier, and the true 3D position p 3d,gt t The 3D position p^ close to 3d t is calculated.
[0043] The calculation unit 16 calculates the three-dimensional position p^ of the object 90 at each time t. 3d t Calculate the three-dimensional position p^ at t=0,1,···,T 3d t of the series (time series 3D position p^ 3d ) is output.
[0044] The three-dimensional position calculation device 10 may be realized by, for example, a computer 50 shown in Fig. 11. The computer 50 includes a CPU (Central Processing Unit) 51, a memory 52 as a temporary storage area, and a non-volatile storage device 53. The computer 50 also includes an input / output device 54 such as an input device and a display device, and an R / W (Read / Write) device 55 that controls reading and writing of data from and to a storage medium 59. The computer 50 also includes a communication I / F (Interface) 56 that is connected to a network such as the Internet. The CPU 51, memory 52, storage device 53, input / output device 54, R / W device 55, and communication I / F 56 are connected to one another via a bus 57.
[0045] The storage device 53 is, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage device 53 serving as a storage medium stores a three-dimensional position calculation program 60 for causing the computer 50 to function as the three-dimensional position calculation device 10. The three-dimensional position calculation program 60 includes an acquisition process control command 62, an interpolation process control command 64, and a calculation process control command 66. The storage device 53 also includes an information storage area 70 in which information constituting the camera parameter DB 20 is stored.
[0046] The CPU 51 reads out the three-dimensional position calculation program 60 from the storage device 53, loads it into the memory 52, and sequentially executes the control instructions contained in the three-dimensional position calculation program 60. The CPU 51 operates as the acquisition unit 12 shown in FIG. 6 by executing an acquisition process control instruction 62. The CPU 51 also operates as the interpolation unit 14 shown in FIG. 6 by executing an interpolation process control instruction 64. The CPU 51 also operates as the calculation unit 16 shown in FIG. 6 by executing a calculation process control instruction 66. The CPU 51 also reads out information from the information storage area 70 and loads the camera parameter DB 20 into the memory 52. As a result, the computer 50 that has executed the three-dimensional position calculation program 60 functions as the three-dimensional position calculation device 10. The CPU 51 that executes the program is hardware.
[0047] The functions realized by the three-dimensional position calculation program 60 may be realized by, for example, a semiconductor integrated circuit, more specifically, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or the like.
[0048] Next, the operation of the three-dimensional position calculation device 10 according to the present embodiment will be described. When a time-series multi-viewpoint image is input to the three-dimensional position calculation device 10 and the calculation of the three-dimensional position of the object 90 is instructed, the three-dimensional position calculation process shown in FIG. 12 is executed in the three-dimensional position calculation device 10. Note that the three-dimensional position calculation process is an example of the three-dimensional position calculation method of the disclosed technology.
[0049] In step S10, the acquisition unit 12 acquires a time-series multi-viewpoint image to which information on the two-dimensional position of the object 90 is attached. Next, in step S12, the acquisition unit 12 sets 1 to a variable t representing the time to be processed. Next, in step S14, the interpolation unit 14 interpolates the two-dimensional position of the object 90 in the image 40n(t) based on the two-dimensional positions of the object 90 detected from each of the images 40n(t - 1) and 40n(t + 1) included in the multi-viewpoint image.
[0050] Next, in step S16, the calculation unit 16 calculates candidates for the three-dimensional position of the object 90 using the camera parameters of the camera 30n that captured the corresponding image 40n(t) for each combination of the detected and interpolated two-dimensional positions of the object 90 between the images 40n(t). Next, in step S18, the calculation unit 16 clusters the candidates based on the distances between the calculated candidates for the three-dimensional position, and calculates the center of gravity of the cluster with the largest number of candidates included in the cluster as the three-dimensional position of the object 90 at time t.
[0051] Next, the acquisition unit 12 determines whether t is smaller than the time information T of the last frame of the multi-viewpoint image. If t < T, the process proceeds to step S22, the acquisition unit 12 increments t by 1, and returns to step S14. On the other hand, if t ≥ T, the process proceeds to step S24, the calculation unit 16 outputs the calculated three-dimensional position, and the three-dimensional position calculation process ends.
[0052] As described above, the three-dimensional position calculation device according to this embodiment acquires a first image taken at a time before a target time, a second image taken at a target time, and a third image taken at a time after the target time, using each of a plurality of cameras that capture images of an object from multiple viewpoints. The three-dimensional position calculation device also interpolates the two-dimensional position of the object in the second image based on the two-dimensional positions of the object detected from each of the first and third images. The three-dimensional position calculation device then calculates the three-dimensional position of the object at the target time based on the two-dimensional position of the object detected in the second image, the interpolated two-dimensional position, and the camera parameters of each of the plurality of cameras. This supplements the two-dimensional position information used to calculate the three-dimensional position, even when there are many outliers or undetected objects in the two-dimensional positions on the images, thereby improving the accuracy of calculating the three-dimensional position of the object using multi-view images.
[0053] Furthermore, the 3D position calculation device according to this embodiment calculates candidates for the 3D position of the object for each combination of detected and interpolated 2D positions between images. The 3D position calculation device then clusters the candidates based on the distance between the candidates, and calculates the center of gravity of the cluster containing the largest number of candidates as the 3D position of the object at the target time. This allows outliers to be appropriately excluded from the candidates, improving the accuracy of calculating the 3D position of the object using multi-view images.
[0054] In the above embodiment, outliers are removed by clustering 3D position candidates and selecting the largest cluster. However, this is not limiting. For example, a method such as the above-mentioned RANSAC may be applied. Specifically, the 3D position calculation device calculates the positions of the centers of gravity of multiple candidates as 3D positions and projects the calculated 3D positions onto each of the second images based on the camera parameters of each of the multiple cameras. The 3D position calculation device may then recalculate the 3D position of the object using 2D positions detected and interpolated in the second image whose distance from the projected positions is within a predetermined threshold. However, while the method using clustering, as in the above embodiment, functions properly even when the proportion of candidates other than outliers is below 50%, other outlier removal methods, such as RANSAC, may have difficulty removing outliers appropriately.
[0055] Here, an example of the calculation result of the three-dimensional position by the three-dimensional position calculation device according to the embodiment will be described. Fig. 13 is an image diagram showing an example in which the position of a person in the world coordinate system is calculated based on skeletal information of the person recognized from the image. In both the upper (shelf) and lower (campus) examples, it can be seen that the three-dimensional position of each person can be calculated with high accuracy, despite the fact that people are crowded together and overlap on the image.
[0056] Fig. 14 is an image diagram showing an example in which the three-dimensional position of the midpoint of the bottom of a 2D-BBOX detected from an image, i.e., the position of a person's feet, is calculated, and the foot positions represented by the three-dimensional position are mapped on a map showing the area to be photographed. In the example of Fig. 14, even though people overlap on the image and are obscured by obstacles such as shelves, the three-dimensional position of each person's feet is calculated with high accuracy. As shown in the example of Fig. 14, the three-dimensional position calculation device according to this embodiment can be applied to a system that acquires the movement trajectories of customers in a store.
[0057] Furthermore, the three-dimensional position calculation device according to the above embodiment can be applied to, for example, a scoring system for gymnastics competitions. Here, an outline of the processing of the scoring system for gymnastics competitions will be described with reference to FIG.
[0058] When multi-perspective images are input, the scoring system detects a person's area from each image included in the multi-perspective images. Next, the scoring system determines whether the person indicated by the detected area is an athlete or a non-athlete, based on whether the person's location is within the competition area or not, and identifies the area representing the athlete. The scoring system tracks the athlete by matching areas representing the same athlete in the time-series multi-perspective images. The scoring system recognizes the athlete's two-dimensional skeletal information from each of the tracked images using a recognition model or the like. The scoring system estimates three-dimensional skeletal information from the two-dimensional skeletal information using camera parameters. The scoring system then performs post-processing such as smoothing on the time-series three-dimensional skeletal information, estimates the phases (breaks) of the performance, and recognizes the technique.
[0059] In the processing of the scoring system, the three-dimensional position calculation device according to the above embodiment can be applied to the processing of estimating three-dimensional skeletal information from two-dimensional skeletal information.
[0060] The disclosed technology is not limited to gymnasts as the target object, but can be applied to various people as the target object, such as athletes of other sports, ordinary pedestrians, etc. Furthermore, it can also be applied to animals, vehicles, etc. as the target object other than people.
[0061] In the above embodiment, the three-dimensional position calculation program is stored (installed) in advance in a storage device, but this is not limiting. The program according to the disclosed technology may be provided in a form stored in a storage medium such as a CD-ROM, a DVD-ROM, or a USB memory. [Explanation of symbols]
[0062] 10 3D position calculation device 12 Acquisition Department 14 Interpolation section 16 Calculation section 20 Camera parameter DB 30n camera 40n images 50 Computers 51 CPU 52 memory 53 Storage device 54 Input / Output Devices 55 R / W device 56 Communication I / F 57 Bus 59 Storage medium 60-dimensional position calculation program 62 Acquisition Process Control Instructions 64 Interpolation Process Control Instructions 66 Calculation Process Control Instructions 70 Information storage area 90 Objects
Claims
1. a first image taken at one or more times before a target time, a second image taken at the target time, and a third image taken at one or more times after the target time, each of the first and second images being taken with a plurality of cameras that photograph an object from different viewpoints; interpolating a two-dimensional position of the object in the second image based on the two-dimensional positions of the object detected from each of the first image and the third image; Calculating a three-dimensional position of the object at the target time based on the two-dimensional position of the object detected from the second image, the interpolated two-dimensional position of the object in the second image, and camera parameters of each of the plurality of cameras. A three-dimensional position calculation method in which a computer executes a process including the steps of:
2. 2. The three-dimensional position calculation method according to claim 1, wherein the process of calculating the three-dimensional position of the object includes calculating the three-dimensional position of the object at the target time based on candidate three-dimensional positions of the object calculated for each combination of the two-dimensional position of the object detected and interpolated from the second image taken by a first camera included in the plurality of cameras and the two-dimensional position of the object detected and interpolated from the second image taken by a second camera included in the plurality of cameras.
3. 3. The three-dimensional position calculation method according to claim 2, wherein the process of calculating the three-dimensional position of the object includes clustering the candidates based on the distances between the candidates, and calculating the center of gravity of the cluster containing the largest number of candidates as the three-dimensional position of the object at the target time.
4. The three-dimensional position calculation method according to any one of claims 1 to 3, wherein the three-dimensional position of the object at the target time is recalculated using the two-dimensional position of the object detected from the second image, and the interpolated two-dimensional position of the object in the second image, whose distance from the calculated three-dimensional position of the object projected onto each of the second images based on the camera parameters of each of the multiple cameras is within a predetermined threshold.
5. The three-dimensional position calculation method according to any one of claims 1 to 3, wherein the process of interpolating the two-dimensional position of the object in the second image includes predicting the two-dimensional position of the object in the second image by linear interpolation of the two-dimensional positions of the object detected from each of the first image and the third image.
6. The three-dimensional position calculation method according to any one of claims 1 to 3, wherein the two-dimensional position of the object is the coordinate value of a predetermined point within an area surrounding the object detected from each of the first image, the second image, and the third image using a detection model generated in advance by machine learning to detect the area of the object from the image, or the coordinate value of a part of the object recognized from each of the first image, the second image, and the third image using a recognition model generated in advance by machine learning to recognize one or more parts of the object from the image.
7. an acquisition unit that acquires, by each of a plurality of cameras that capture images of an object from a plurality of different viewpoints, a first image captured at one or more times before a target time, a second image captured at the target time, and a third image captured at one or more times after the target time; an interpolation unit that interpolates a two-dimensional position of the object in the second image based on the two-dimensional positions of the object detected from each of the first image and the third image; a calculation unit that calculates a three-dimensional position of the object at the target time based on the two-dimensional position of the object detected from the second image, the interpolated two-dimensional position of the object in the second image, and camera parameters of each of the plurality of cameras; A three-dimensional position calculation device including:
8. a first image taken at one or more times before a target time, a second image taken at the target time, and a third image taken at one or more times after the target time, each of the first and second images being taken with a plurality of cameras that photograph an object from different viewpoints; interpolating a two-dimensional position of the object in the second image based on the two-dimensional positions of the object detected from each of the first image and the third image; Calculating a three-dimensional position of the object at the target time based on the two-dimensional position of the object detected from the second image, the interpolated two-dimensional position of the object in the second image, and camera parameters of each of the plurality of cameras. A three-dimensional position calculation program for causing a computer to execute processing including the above.
Citation Information
Patent Citations
Three-dimensional information detecting device and three-dimensional information detecting method
JP2002008040A
Motion capture method
JP2014211404A
Tracking of handheld sporting implements using computer vision
US20170270354A1
Stereo camera
WO2011096251A1
Image capturing device and image capturing method
WO2016132950A1