Method and system for generating a training dataset for keypoint detection, and method and system for predicting the 3D location of virtual markers on markerless objects

The method generates a high-quality training dataset using an optical marker-based system and neural networks to predict 3D marker locations, addressing accuracy and cost issues in existing keypoint detection models, particularly in markerless motion capture systems.

JP7712706B2Active Publication Date: 2025-07-24NANYANG TECH UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023577120
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-14
Filing Date
2022-06-10
Publication Date
2025-07-24
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

Existing keypoint detection models rely on manually annotated datasets with inconsistent quality, limiting accuracy, especially in applications requiring high precision like sports science and rehabilitation, and existing markerless systems face challenges with occlusion and cost.

Method used

A method and system using an optical marker-based motion capture system to generate a high-quality training dataset by projecting 3D marker trajectories onto 2D images, and a neural network to predict 3D locations of virtual markers on markerless subjects, employing global and rolling shutter cameras with calibration techniques to minimize errors.

Benefits of technology

Improves keypoint detection accuracy by reducing manual annotation errors and system costs, while enabling markerless motion capture with fewer cameras and less occlusion, achieving sub-centimeter precision in 3D marker localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712706000007
    Figure 0007712706000007
  • Figure 0007712706000008
    Figure 0007712706000008
  • Figure 0007712706000009
    Figure 0007712706000009
Patent Text Reader

Abstract

According to an embodiment of the present invention, a method and system for generating a training dataset for keypoint detection is provided. The system includes an optical marker-based motion capture system for capturing markers as 3D trajectories and a video camera for simultaneously capturing a sequence of 2D images. Each marker is positioned on a bony landmark or keypoint of a subject. The method, performed by a computer in the system, includes projecting each trajectory onto each image to determine a 2D location for each marker, interpolating a 3D position therefrom, generating a bounding box around the subject, and generating a training dataset including at least one image and the determined 2D location and bounding box of each marker therein. According to a further embodiment, a method and system for predicting the 3D location of a virtual marker on a marker-free subject using a neural network trained by the generated training dataset is also provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of priority of Singapore Patent Application No. 10202106342T, filed on June 14, 2021, the contents of which are hereby incorporated by reference in their entirety for all purposes.

[0002] Various embodiments relate to methods and systems for generating a training dataset for keypoint detection, and methods and systems for predicting the 3D location of virtual markers on markerless objects (e.g., humans, animals, or objects) using a neural network trained by the generated training dataset.

Background Art

[0003] The ability to sense and digitize the kinematics of human movement has unlocked research and applications in many areas such as motion analysis in sports science, anomaly diagnosis in rehabilitation, and character animation in the film industry, or such ability can serve the purpose of human-computer interaction in video games, interactive technologies, or different types of computer applications. Technologies that provide such ability can be in various forms. One early off-the-shelf technology that is still widely used today is the multi-camera marker-based motion capture form. In this technology, the bone landmarks of the subject are attached using retro-reflective markers that are viewed by infrared cameras with active infrared light sources. When one marker is viewed by two or more infrared cameras, assuming that those infrared cameras are calibrated and synchronized, the three-dimensional (3D) position of the marker is calculated from triangulation. Subsequently, a sequence of these 3D positions is used in subsequent applications.

[0004] Since the introduction of the Deep Convolutional Neural Network (AlexNet) in 2012 for image classification, an even larger number of more complex computer vision problems have since been approached in a similar way under the subsequent data-driven paradigm. One of the challenges encountered has been human pose estimation or human keypoint detection. Neural network models in this field have rapidly evolved over the past decade due to the popularity of model-centric approaches. Scientists have typically downloaded available public datasets and proposed new neural network architectures or new training methods that can improve test accuracy over existing models. This trend has led to many significant contributions to the models, but not as many contributions have been made to the datasets and data quality. In the field of human keypoint detection, the two largest data collection efforts are from the COCO dataset and the MPII dataset, which have 118K images and 40K images respectively. These datasets contain the 2D joint positions that have been manually annotated for every human in the dataset. For example, in the COCO dataset, all keypoints are manually annotated through crowdsourcing. As far as the inventors are aware, all prior art data-driven human keypoint detection models rely on these manually annotated datasets for training, regardless of the annotation accuracy.

[0005] The task of selecting pixels on a high-resolution image representing the joint center is difficult for humans to perform accurately due to various possible problems. The main reason is that there is no clear consensus among the annotating workers on exactly where each joint center is in relation to the actual human bone landmarks. Even when a definition is given, the 2D images may not provide enough clues about those bone landmarks, so the definition remains extremely difficult to find. Thus, the annotated positions appear to be more like blurred 2D areas rather than pinpoint locations at the pixel level. Using these datasets surely limits how accurate the trained model can be. This level of quality may be sufficient for entertainment purposes such as video game control or interactive technologies. However, in more contentious applications such as sports science, biomechanical analysis, or rehabilitation analysis, such a level of quality is often not considered suitable or sufficient.

[0006] In another existing approach, the Kinect skeleton tracking system operates using depth images. A human model in a random pose can be created in 3D. Random forest regression is used to predict each pixel in order to determine every part of the human. However, the accuracy of such synthetic models is not great because the constraints used may not be realistic.

[0007] Therefore, there is a need for a method and / or system for at least addressing the above problems, and more particularly, a method and / or system in which the system involves the use of multiple RGB cameras, the system and / or method generates 3D marker position outputs, and no markers or sensors are placed on the subject body. One obvious benefit of being markerless is the reduction of time and manpower in subject creation, which makes the motion capture (mocap) workflow more practical for unpredictable applications such as medical diagnosis. The present method and / or system also does not involve overly restrictive constraints.

[0008] Furthermore, instead of continuing with the model-centric trend, the present method and / or system may involve a data-centric approach that generates a dataset with the best prior art keypoint detection model recognized by players in the field and annotations of the highest possible quality. Such annotations must not be derived from human decisions but from accurate sensors such as marker positions from marker-based motion capture systems. If the markers are correctly placed on the bone landmarks and the marker-based motion capture system can accurately extract the 3D trajectories of the markers, then the markers can in turn be projected onto the video frames to obtain pixel-accurate 2D ground truths for the training of keypoint detection. To ensure that all calibrations and synchronizations work under a relatively low budget, the data collection infrastructure is also designed from the ground up. Additionally, this can very well avoid or at least reduce the inconsistent quality of camera calibration parameters and time synchronization used when obtaining existing datasets of the same type that may cause significantly large errors after projection. For example, based on cropped images from the MoVi dataset using marker projection in 2D, it can be seen that due to insufficient camera calibration and synchronization, the projection does not align with the markers.

Prior Art Documents

Non-Patent Documents

[0009] [Non-Patent Document 1] P. Liang et al., "An asian-centric human movement database capturing activities of daily living", Scientific Data, vol. 7, no. 1, pp. 1-13, 2020 [Summary of the Invention]

[0010] According to one embodiment, a method for generating a training dataset for keypoint detection is provided. The method can be based on a plurality of markers respectively captured as 3D trajectories by an optical marker-based motion capture system, each marker being placed on a bone landmark of a human or animal subject, or a keypoint of an object, and the human or animal subject or object being substantially simultaneously captured by a plurality of color video cameras over a time period as a sequence of 2D images. The method includes, for each marker, projecting the 3D trajectory onto each of the 2D images to determine the 2D location in each 2D image, and for each marker, interpolating the 3D position for each of the 2D images based on each of the 2D locations in the sequence of 2D images and the exposure relationship time of the plurality of color video cameras, and for each 2D image, generating a 2D bounding box around the human or animal subject or object based on the interpolated 3D positions of the respective markers and an extended volume derived from two or more of the markers having an anatomical or functional relationship with each other, and generating a training dataset including at least one 2D image selected from the sequence of 2D images, the determined 2D locations of each marker in the at least one selected 2D image, and the generated 2D bounding box for the at least one selected 2D image.

[0011] According to one embodiment, a method is provided for predicting the 3D location of a virtual marker on a markerless human or animal subject or a markerless object. The method is based on a markerless human or animal subject or a markerless object captured by a plurality of color video cameras as a sequence of 2D images. For each 2D image captured by each color video camera, the method includes predicting a 2D bounding box using a trained neural network, generating a plurality of heatmaps with reliability scores by the trained neural network for each 2D image, selecting for each heatmap the pixel with the highest reliability score, associating the selected pixel with the virtual marker, thereby determining the 2D location of the virtual marker, and triangulating each determined 2D location to predict a sequence of 3D locations of the virtual marker based on the sequence of 2D images captured by the plurality of color video cameras. Each heatmap is for the 2D localization of the virtual marker on the markerless human or animal subject or the markerless object, and for each heatmap, the reliability score indicates the probability of having the associated virtual marker at different 2D locations within the predicted 2D bounding box. The trained neural network is trained using a training dataset generated by a method for generating at least a training dataset for keypoint detection according to one of the above embodiments.

[0012] According to one embodiment, a computer program is provided that is adapted to execute a method for generating a training dataset for keypoint detection and / or a method for predicting the 3D location of a virtual marker on a markerless human or animal subject or a markerless object according to the various embodiments described above.

[0013] According to one embodiment, a non-transitory computer-readable medium is provided that, when executed on a computer, causes the computer to execute a method for generating a training dataset for keypoint detection according to the various embodiments described above, and / or a method for predicting the 3D location of virtual markers on markerless human or animal subjects or markerless objects.

[0014] According to one embodiment, a data processing apparatus is provided that includes means for executing a method for generating a training dataset for keypoint detection according to the various embodiments described above, and / or a method for predicting the 3D location of virtual markers on markerless human or animal subjects or markerless objects.

[0015] According to one embodiment, a system for generating a training dataset for keypoint detection is provided. The system may include an optical marker-based motion capture system configured to capture a plurality of markers over a time period, where each marker is placed on a bone landmark of a human or animal subject or a keypoint of an object and is captured as a 3D trajectory, a plurality of color video cameras configured to capture a human or animal subject or an object over a time period as a sequence of 2D images, and a computer. The computer is configured to receive a sequence of 2D images captured by the plurality of color video cameras and respective 3D trajectories captured by the optical marker-based motion capture system, project the 3D trajectories onto each of the 2D images to determine 2D locations in each 2D image for each marker, interpolate 3D positions for each of the 2D images based on respective 2D locations in the sequence of 2D images for each marker and exposure relationship times of the plurality of color video cameras, generate a 2D bounding box around the human or animal subject or object based on respective interpolated 3D positions of the plurality of markers and an extended volume derived from two or more of the markers having an anatomical or functional relationship with each other, and generate a training dataset including at least one 2D image selected from the sequence of 2D images, determined 2D locations of each marker in at least one selected 2D image, and the generated 2D bounding box for at least one selected 2D image.

[0016] According to one embodiment, a system is provided for predicting the 3D location of a virtual marker on a markerless human or animal subject or a markerless object. The system may include a plurality of color video cameras configured to capture a markerless human or animal subject or a markerless object as a sequence of 2D images, and a computer. The computer is configured to receive a sequence of 2D images captured by the plurality of color video cameras, predict a 2D bounding box for each 2D image captured by each color video camera using a trained neural network, generate a plurality of heatmaps with reliability scores for each 2D image using a trained neural network, select, for each heatmap, the pixel with the highest reliability score, associate the selected pixels with the virtual marker to determine the 2D location of the virtual marker, and triangulate each determined 2D location to predict a sequence of 3D locations of the virtual marker based on the sequence of 2D images captured by the plurality of color video cameras. Each heatmap is for the 2D localization of the virtual marker on the markerless human or animal subject or the markerless object, and for each heatmap, the reliability score indicates the probability of having the virtual marker associated with different 2D locations within the predicted 2D bounding box. The trained neural network is trained using a training dataset generated by a system and / or method for generating at least a training dataset for keypoint detection according to the various embodiments described above.

[0017] In the drawings, like reference numerals generally refer to like parts throughout different figures. The drawings are not necessarily to scale; instead, emphasis is generally placed on illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings.

Brief Description of the Drawings

[0018]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

[0019] The following detailed description refers to the accompanying drawings which illustrate specific details and examples by which the present invention may be practiced. These examples are described in sufficient detail to enable those skilled in the art to practice the present invention. Other examples may be utilized and structural, logical, and electrical changes may be made without departing from the scope of the present invention. Since some examples can be combined with one or more other examples to form new examples, the various examples are not necessarily mutually exclusive.

[0020] Examples described in the context of one of the methods or devices are similarly valid for other methods or devices. Similarly, examples described in the context of a method are similarly valid for a device, and vice versa.

[0021] Features described in the context of one example may correspondingly be applicable to the same or similar features in other examples. Features described in the context of one example may be applicable to these other examples correspondingly, even if not explicitly described in these other examples. Further, the additional and / or combinations and / or alternatives described for a feature in the context of one example may be applicable to the same or similar features in other examples correspondingly.

[0022] In the context of the various examples, the articles “a,” “an,” and “the” used with respect to a feature or element include reference to one or more of the feature or elements.

[0023] In the context of the various examples, the phrase “at least substantially” can include “strictly” and reasonable differences.

[0024] In the context of various embodiments, the terms "about" or "substantially" as applied to a numerical value encompass both the exact value and reasonable differences.

[0025] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0026] As used herein, a phrase in the form of "at least one of A or B" can include A or B, or both A and B. Correspondingly, a phrase in the form of "at least one of A or B or C", or including further listed items, can include any and all combinations of one or more of the associated listed items.

[0027] As used herein, the expression "configured to" can mean "constructed to" or "arranged to".

[0028] Various embodiments can provide a data-driven markerless multi-camera human motion capture system. In order for such a system to be data-driven, it is important to use a suitable and accurate training dataset.

[0029] Figure 1A shows a flowchart of a method 100 for generating a training dataset for keypoint detection according to various embodiments. In a preceding step 102, a plurality of markers are each captured as a 3D trajectory by an optical marker-based motion capture system. Each marker can be placed on a bone landmark of a human or animal subject, or on a keypoint of an object. The human or animal subject or object is substantially simultaneously captured by a plurality of color video cameras over a time period as a sequence of 2D images. The time period can vary depending on how much time is required to capture the movement of the subject. The object can be a moving object, such as a sports equipment that can be tracked when in use, for example, a tennis racket. Method 100 includes the following active steps. In step 104, for each marker, the 3D trajectory is projected onto each of the 2D images to determine the 2D location in each 2D image. In step 106, for each marker, based on each of the 2D locations in the sequence of 2D images and the exposure relationship time of the plurality of color video cameras, the 3D position for each of the 2D images is interpolated. In step 108, for each 2D image, based on the interpolated 3D positions of each of the plurality of markers and an extended volume derived from two or more of the markers having an anatomical or functional relationship with each other, a 2D bounding box is generated around the human or animal subject or object. In step 110, a training dataset is generated, and the training dataset includes at least one 2D image selected from the sequence of 2D images, the determined 2D locations of each marker in the at least one selected 2D image, and the generated 2D bounding box for the at least one selected 2D image. For example, in the case of a human or animal subject, two or more markers for deriving the extended volume can have at least one of an anatomical relationship or a functional relationship with each other.In another exemplary case of an object, two or more markers for deriving an extended volume may have a functional (and / or structural) relationship to each other.

[0030] In other words, method 100 focuses on learning from marker data instead of manually annotated data. Using training data collected from a marker-based motion capture system significantly improves the accuracy and efficiency of data collection. With respect to position accuracy, manual annotation can often miss the joint center by several centimeters, while marker position accuracy is within the range of a few millimeters. With respect to data generation efficiency, for example, manual annotation performed in existing techniques can take at least 20 seconds per image. On the other hand, method 100 can generate and annotate data at an average rate of 80 images per second (including manual data cleanup time). This advantageously enables efficient scaling of data collection up to millions of images. With a larger data set, more accurate training data improves the accuracy of any sequential machine learning model for this task.

[0031] In various embodiments, in preceding step 102, a plurality of markers captured as 3D trajectories and a human or animal subject or object captured substantially simultaneously as a sequence of 2D images over a time period may be coordinated using a synchronization signal communicated to a plurality of color video cameras by an optical marker-based motion capture system. In the context of various embodiments, the phrase "preceding step" refers to this step preceding or being performed beforehand. The preceding step may be an inactive step of the method.

[0032] Method 100 may further include identifying the captured 3D trajectory using a label representing a bone landmark or key point on which a marker is placed, prior to the step of projecting the 3D trajectory in step 104. For each marker, the label may be configured to be propagated along with each determined 2D location such that in the generated training dataset, each determined 2D location of each marker includes the corresponding label.

[0033] Method 100 may further include, after the step of projecting the 3D trajectory onto each 2D image to determine the 2D location in each 2D image in step 104, for each 2D image and for each marker, drawing a 2D radius on the determined 2D location according to a distance with a predefined margin between the color video camera (that captured the specific (each) 2D image) and the marker to form an enclosed area, and applying a learning-based context-aware image inpainting technique to the enclosed area to remove the marker blob from the 2D location. For example, the learning-based context-aware image inpainting technique may include a Generative Adversarial Network (GAN)-based context-aware image inpainting technique.

[0034] In various embodiments, the plurality of color video cameras (or RGB cameras) may include a plurality of global shutter cameras. The exposure relationship time may be the mid-exposure time, which is in the middle of the exposure period, for capturing each 2D image using each global shutter camera. Each global shutter camera may include at least one visible light emitting diode (LED) operable to facilitate the inverse reflection marker coupled to the wand being perceived as a bright spot detectable. For example, the visible LED may include a white LED.

[0035] In the context of various embodiments, the term "wand" refers to an elongated object to which a retroreflective marker can be coupled that facilitates the wavy movement of the retroreflective marker.

[0036] Multiple global shutter cameras can be pre-calibrated as follows. Based on the retroreflective markers captured by an optical marker-based motion capture system as a 3D trajectory in which the wand is continuous and covers a target capture volume (or target motion capture volume), and the retroreflective markers captured substantially simultaneously by each global shutter camera as a sequence of 2D calibration images during a time period, for each 2D calibration image, the 2D calibration position of the retroreflective marker can be extracted by searching for bright pixels and scanning across the 2D calibration image to identify the 2D location of the bright pixels. The time period for capture in the pre-calibration can be shorter than two minutes or an amount sufficient for the trajectory to cover the capture volume. An iterative algorithm can be applied at the 2D location of the searched bright pixels to converge the 2D location at the centroid of the bright pixel cluster. Further, based on the mid-exposure time, also referred to interchangeably as the mid-exposure period or mid-exposure timing, within the exposure period in each 2D calibration image, and the 3D trajectory covering the target capture volume, a 3D calibration position can be linearly interpolated at the mid-exposure time from each of the 2D calibration images. For at least some of the multiple 2D calibration images, multiple 2D-3D correspondence pairs can be formed. Each 2D-3D correspondence pair can include the converged 2D location and the interpolated 3D calibration position for each of at least some of the multiple 2D calibration images. A camera calibration function can be applied to the multiple 2D-3D correspondence pairs to determine the extrinsic camera parameters and fine-tune the intrinsic camera parameters of the multiple global shutter cameras.

[0037] In existing motion capture systems in the market, in order to reduce troublesome problems in calculations, almost always, the camera system needs to use a global shutter sensor (or a global shutter camera).

[0038] In other embodiments, the plurality of color video cameras can be a plurality of rolling shutter cameras. Since additional errors related to rolling shutter artifacts can occur, replacing a global shutter camera with a rolling shutter camera is not a plug-and-play process. To adapt a rolling shutter camera to the type of motion capture system used in method 100, careful modeling of the camera timing, synchronization, and calibration is required to minimize errors due to the rolling shutter effect. However, since rolling shutter cameras are significantly less expensive than global shutter cameras, this benefit of compatibility is a reduction in system cost.

[0039] In these other embodiments, the step of projecting a 3D trajectory onto each 2D image in step 104 may further include determining an intersection time from the intersection between a first line connecting the 3D trajectories projected over a time period to capture each pixel row of the 2D image and a second line representing the moving midpoint of the exposure time for each 2D image captured by each rolling shutter camera; interpolating 3D intermediate positions based on the intersection time to obtain a 3D interpolated trajectory from the sequence of 2D images for each 2D image captured by each rolling shutter camera; and projecting the 3D interpolated trajectory onto each 2D image to determine the 2D location in each 2D image for each marker. The exposure relationship time when using a plurality of rolling shutter cameras is the intersection time.

[0040] Similar to the embodiments with global shutter cameras, each rolling shutter camera here may include at least one visible light emitting diode operable to facilitate the retroreflective marker coupled to the wand being perceived as a detectable bright spot.

[0041] Multiple rolling shutter cameras can be pre-calibrated as follows. The wand is a continuous wave, based on the retroreflective markers captured by an optical marker-based motion capture system as a 3D trajectory covering the target capture volume, and the retroreflective markers captured substantially simultaneously by each rolling shutter camera as a sequence of 2D calibration images over a time period. For each 2D calibration image, the 2D calibration position of the retroreflective marker can be extracted by searching for bright pixels and scanning across the entire 2D calibration image to identify the 2D location of the bright pixels. An iterative algorithm can be applied at the 2D location of the searched bright pixels to converge the 2D location at the 2D centroid of the bright pixel cluster. Further, based on the observation times of the 2D centroids from multiple rolling shutter cameras, the 3D calibration positions can be interpolated from the 3D trajectory covering the target capture volume. The observation time of each 2D centroid of each bright pixel cluster from each 2D calibration image is T i +b - e / 2 + dv, Equation 1 calculated by where T i is the trigger time of the 2D calibration image, b is the trigger readout delay of the rolling shutter camera, e is the exposure time set for the rolling shutter camera, d is the line delay of the rolling shutter camera, and v is the pixel row of the 2D centroid of the bright pixel cluster.

[0042] For at least some of the plurality of 2D calibration images, a plurality of 2D-3D correspondence pairs can be formed. Each 2D-3D correspondence pair can include a converged location for each of at least some of the plurality of 2D calibration images and an interpolated 3D calibration position. A camera calibration function can be applied to the plurality of 2D-3D correspondence pairs to determine extrinsic camera parameters and to fine-tune the intrinsic camera parameters of the plurality of rolling shutter cameras.

[0043] When pre-calibrating a plurality of color video cameras, the iterative algorithm can be an average shift algorithm. Retroreflective markers captured as 3D trajectories covering a target capture volume by an optical marker-based motion capture system and retroreflective markers captured substantially simultaneously as a sequence of 2D calibration images can be coordinated using a synchronization signal communicated to the plurality of color video cameras by the optical marker-based motion capture system.

[0044] Generally, in the hardware layer of a motion capture system, in order to avoid the influence of the detection delay between the upper pixel row and the lower pixel row received by a rolling shutter camera, the camera needs to use a global shutter sensor. However, the implementation of a global shutter camera requires a more complex electronic circuit to execute simultaneous start and stop of the exposure of all pixels. As a result, a global shutter camera becomes significantly more expensive than a rolling shutter camera with the same resolution. Since human motion is not fast enough to be excessively distorted by the rolling shutter effect, it may be possible to reduce the system cost by using a rolling shutter camera with a detailed modeling of the rolling shutter effect to compensate for errors. This rolling shutter model can be integrated into the entire workflow starting from camera calibration, data collection, and culminating in the triangulation of 3D keypoints, which will be further described below. Therefore, advantageously, additional flexibility in the selection of the camera is obtained.

[0045] In various embodiments, the markers referred to with respect to method 100 include retroreflective markers.

[0046] Figure 1B shows a flow chart illustrating a method 120 for predicting the 3D location of virtual markers on markerless human or animal subjects or markerless objects according to various embodiments. In a preceding step 122, the markerless human or animal subject or markerless object is captured by a plurality of color video cameras as a sequence of 2D images. Method 120 includes the following active steps. In step 125, for each 2D image captured by each color video camera, a 2D bounding box is predicted using a trained neural network. In step 124, for each 2D image, a plurality of heatmaps with confidence scores are generated by the trained neural network. Each heatmap is for the 2D localization of a virtual marker of a markerless human or animal subject or markerless object. In the context of various embodiments, 2D localization refers to the process of identifying the 2D location or position of a virtual marker, and thus each heatmap is associated with one virtual marker. The trained neural network can be trained using at least a training dataset generated by method 100. In step 126, for each heatmap, the pixel with the highest confidence score is selected or chosen, and the selected pixel is associated with the virtual marker, thereby determining the 2D location of the virtual marker. For each heatmap, the confidence score indicates the probability of having an associated virtual marker at different 2D locations within the predicted 2D bounding box. In step 128, based on the sequence of 2D images captured by the plurality of color video cameras, each determined 2D location is triangulated to predict a sequence of 3D locations of the virtual marker. Optionally, the triangulation step in step 128 can include weighted triangulation of each 2D location of the virtual marker based on each confidence score as a weight for triangulation. For example, the weighted triangulation is according to the formula (Σi w i Q i ) -1 (Σ i w i Q i C i ) may include deriving each predicted 3D location of the virtual marker to be used, where i is 1, 2, …, N (N is the total number of color video cameras), and w i is the weight for triangulation or the reliability score of the i-th ray from the i-th color video camera, and C i is the 3D location of the i-th color video camera associated with the i-th ray, and U i is a 3D unit vector representing the back-projected direction associated with the i-th ray, given that I3 is a 3×3 identity matrix

Number

[0047] In other words, method 120 outputs virtual marker positions instead of joint centers. Since a typical biomechanical analysis workflow begins calculations from 3D marker positions, it is important to maintain the marker positions (more specifically, the virtual marker positions) in the output of method 120 to ensure that method 120 is compatible with existing workflows. Different from existing systems that learn from manually annotated joint centers, learning to predict marker positions generates not only computable joint positions but also the orientations of body segments. These orientations of body segments cannot be recovered from a set of joint centers for each posture. For example, when the shoulder, elbow, and wrist are approximately in a straight line, the specificity of this arm posture makes it impossible to recover the orientations of the upper arm and forearm joints. However, these orientations can be calculated from the shoulder marker, elbow marker, and wrist marker. Therefore, enabling the machine learning model to predict marker positions (more specifically, virtual marker positions) instead of joint center positions is essential for more controversial application examples. Direct Linear Transformation (DLT) is an established method for performing triangulation to obtain 3D positions from multiple 2D positions observed by two or more cameras. For this application, a new triangulation formula has been derived to improve triangulation accuracy by utilizing a reliability score (or equivalently called a confidence score), which is additional information provided by a neural network model for each predicted 2D location. For each ray (e.g., representing a 2D location on one image), the reliability score can be included as a weight in this new triangulation formula. Method 120 can significantly improve triangulation accuracy with respect to the DLT method.

[0048] In various embodiments, the plurality of color video cameras can include a plurality of global shutter cameras.

[0049] In other embodiments, the plurality of color video cameras can be a plurality of rolling shutter cameras. In these other embodiments, method 120 can further include determining, for each rolling shutter camera, an observation time based on the determined 2D locations in two consecutive 2D images, before the step of triangulating the respective 2D locations for predicting the sequence of 3D locations of the virtual marker in step 128. The observation time can be calculated using Equation 1, where T i refers to the trigger time of each of the two consecutive 2D images, and v is the pixel row of the 2D location in each of the two consecutive 2D images. Based on the observation time, the 2D location of the virtual marker is interpolated at the trigger time. The step of triangulating the respective 2D locations in step 128 can include triangulating each interpolated 2D location derived from the plurality of rolling shutter cameras.

[0050] In one embodiment, the plurality of color video cameras can be externally calibrated as follows. Based on one or more checkerboards simultaneously captured by the plurality of color video cameras, for each pair of the plurality of color video cameras, calculate the relative transformation between those two color video cameras. When the plurality of color video cameras have their respective calculated relative transformations and all existing cameras are linked by the relative transformations, apply an optimization algorithm to fine-tune the extrinsic camera parameters of the plurality of color video cameras. More specifically, the optimization algorithm is the Levenberg-Marquardt algorithm and its cv2 function applied to the 2D checkerboard observations and the initial relative transformation. The one or more check car erboards can include unique markings.

[0051] In another embodiment, the plurality of color video cameras can alternatively be externally calibrated as follows. Each color video camera can include at least one visible light emitting diode (LED) operable to facilitate the perception of a plurality of retroreflective markers coupled to a wand as bright spots detectable thereby. Based on the wand being continuous and wavy and the retroreflective markers captured by the plurality of color video cameras as a sequence of 2D calibration images, an optimization function is applied to the captured 2D calibration images to fine-tune the extrinsic camera parameters of the plurality of color video cameras. The optimization algorithm can be the Levenberg-Marquardt algorithm and its cv2 function, as described above.

[0052] Although the methods described above have been illustrated and described as a series of steps or events, it will be understood that no ordering of such steps or events should be construed in a limiting sense. For example, some steps may be performed in a different order and / or concurrently with other steps or events than those illustrated and / or described herein. Further, not all illustrated steps may be required to implement one or more aspects or embodiments described herein. Also, one or more of the steps shown herein may be performed in one or more separate acts and / or phases.

[0053] Various embodiments may also provide a computer program adapted to execute method 100 and / or method 120 according to various embodiments.

[0054] Various embodiments may further provide a non-transitory computer-readable medium including instructions that, when executed on a computer, cause the computer to execute method 100 and / or method 120 according to various embodiments.

[0055] Various embodiments may also provide a data processing apparatus comprising means for performing method 100 and / or method 120 according to various embodiments.

[0056] Figure 1C shows a schematic diagram of a system 140 for generating a training dataset for keypoint detection according to various embodiments. The system 140 can include an optical marker-based motion capture system 142 configured to capture a plurality of markers over a time period, and a plurality of color video cameras 144 configured to capture a human or animal subject or object as a sequence of 2D images over the time period. Each marker can be placed on a bone landmark of a human or animal subject or on a keypoint of an object and can be captured as a 3D trajectory. The system 140 can also include a computer 146 configured to receive, as indicated by dashed lines 152, 150, a sequence of 2D images captured by the plurality of color video cameras 144 and respective 3D trajectories captured by the optical marker-based motion capture system 142. The time period can vary depending on how much time is required to capture the movement of the subject or object. The computer 146 is further configured to, for each marker, project the 3D trajectory onto each of the 2D images to determine a 2D location in each 2D image, and for each marker, interpolate a 3D position for each of the 2D images based on the respective 2D locations in the sequence of 2D images and the exposure relationship times of the plurality of color video cameras 144, and for each 2D image, generate a 2D bounding box around the human or animal subject or object based on the respective interpolated 3D positions of the plurality of markers and an extended volume derived from two or more of the markers having an anatomical or functional relationship to each other, and generate a training dataset including at least one 2D image selected from the sequence of 2D images, the determined 2D locations of each marker in the at least one selected 2D image, and the generated 2D bounding box for the at least one selected 2D image.In one embodiment, computer 146 can be the same computer that communicates with a plurality of color video cameras 144 and an optical marker-based motion capture system 142 to record respective data. In different embodiments, computer 146 can be a processing computer separate from the computer that communicates with a plurality of color video cameras 144 and an optical marker-based motion capture system 142 to record respective data.

[0057] System 140 can further include a synchronization pulse generator that communicates with optical marker-based motion capture system 142 and a plurality of color video cameras 144, and the synchronization pulse generator can be configured to receive a synchronization signal from the optical marker-based motion capture system 142 to coordinate a human or animal subject or object to be captured substantially simultaneously by the plurality of color video cameras 144, as indicated by line 148. For example, the plurality of color video cameras 144 can include at least two color video cameras, preferably eight color video cameras.

[0058] In various embodiments, the optical motion capture system 142 can include a plurality of infrared cameras. For example, there can be at least two infrared cameras configured to be spaced apart from each other to capture a subject from different views.

[0059] The plurality of color video cameras 144 and the plurality of infrared cameras are configured to be spaced apart from each other and along at least a path taken by a human or animal subject or object, or at least substantially surround a capture volume of the human or animal subject or object.

[0060] The 3D trajectory can be identifiable using a label that represents a bone landmark or key point on which the marker is placed. For each marker, the label can be configured to be propagated with each determined 2D location such that in the generated training dataset, each determined 2D location of each marker includes the corresponding label.

[0061] In some examples, the computer 146 is further configured to draw a 2D radius on the 2D location determined for each marker according to a distance with a predefined margin between the color video camera 144 (that captured the specific (each) 2D image) and the marker to form an enclosed area in each 2D image, and to apply a learning-based context-aware image inpainting technique to the enclosed area to remove the marker blob from the 2D location. For example, the learning-based context-aware image inpainting technique can include an adversarial generative network (GAN)-based context-aware image inpainting technique.

[0062] In various embodiments, the plurality of color video cameras 144 can be a plurality of global shutter cameras.

[0063] In other embodiments, the plurality of color video cameras 144 may be a plurality of rolling shutter cameras. In these other embodiments, the computer 146 is further configured to determine, for each 2D image captured by each rolling shutter camera, the intersection time from the intersection between a first line connecting 3D trajectories projected over a time period to capture each pixel row of the 2D image and a second line representing the moving midpoint of the exposure time, and to interpolate 3D intermediate positions based on the intersection times to obtain 3D interpolated trajectories from the sequence of 2D images for each 2D image captured by each rolling shutter camera, and to project the 3D interpolated trajectories onto each of the 2D images to determine the 2D locations in each 2D image for each marker.

[0064] System 140 may be used to facilitate the execution of method 100. Accordingly, system 140 may include elements or components that are the same as or similar to the elements or components of method 100 of FIG. 1A, and thus, the similar elements may be as described in the context of method 100 of FIG. 1A, and thus, the corresponding description may be omitted here.

[0065] An exemplary setup 200 of system 140 is schematically shown in FIG. 2. As seen in FIG. 2, a plurality of color (RGB: red - green - blue) video cameras 144 and infrared (IR) cameras 203 are configured around an object 205 with retro - reflective markers placed on bone landmarks or key points. Different configurations may be possible (not shown in FIG. 2). As the object 205 moves, the retro - reflective markers also move within the capture volume. The synchronization pulse generator 201 may communicate with the optical motion capture system 142 using the synchronization signal 211, with the color video cameras 144 via the synchronization channel 207, and also with the computer 146. The computer 146 and the color video cameras 144 may communicate using the data channel 209.

[0066] Figure 1D shows a schematic diagram of a system 160 for predicting the 3D location of virtual markers on a markerless human or animal subject or a markerless object according to various embodiments. The system 160 may include a plurality of color video cameras 164 configured to capture a markerless human or animal subject or a markerless object as a sequence of 2D images, as indicated by the dotted line 168, and a computer 166 configured to receive a sequence of 2D images captured by the plurality of color video cameras 164. The computer 166 may further predict a 2D bounding box for each 2D image captured by each color video camera 164 using a trained neural network, generate a plurality of heatmaps with reliability scores for each 2D image using a trained neural network, select the pixel with the highest reliability score for each heatmap, associate the selected pixel with the virtual marker to determine the 2D location of the virtual marker, and triangulate each determined 2D location to predict a sequence of 3D locations of the virtual marker based on the sequence of 2D images captured by the plurality of color video cameras 164. Each heatmap may be for the 2D localization of the virtual marker on the markerless human or animal subject or the markerless object. For each heatmap, the reliability score indicates the probability of having the associated virtual marker at different 2D locations within the predicted 2D bounding box. The trained neural network may be trained using at least a training dataset generated by method 100. In one embodiment, the computer 166 may be the same computer that communicates with the plurality of color video cameras 164 for recording data. In different embodiments, the computer 166 may be a separate processing computer from the computer that communicates with the plurality of color video cameras 164 for recording data.

[0067] Optionally, the respective 2D locations of the virtual markers can be triangulated based on their respective reliability scores as weights for triangulation. For example, triangulation can be performed where i = 1, 2, …, N (where N is the total number of color video cameras), w i is the weight for triangulation, or the reliability score of the ith ray from the ith color video camera, and C i is the 3D location of the ith color video camera associated with the ith ray, and U i is a 3D unit vector representing the back-projected direction associated with the ith ray, given that I3 is the 3×3 identity matrix

Equation

[0068] In various embodiments, the plurality of color video cameras 164 can be a plurality of global shutter cameras.

[0069] In other embodiments, the plurality of color video cameras 164 can be a plurality of rolling shutter cameras. In these other embodiments, the computer 166 can further determine the observation time for each rolling shutter camera based on the determined 2D locations in two consecutive 2D images, and based on the observation time, can be configured to interpolate the 2D location of the virtual marker at the trigger time. The observation time can be calculated using Equation 1, where T iis the trigger time for each of two consecutive 2D images, and v is the pixel row of the 2D location in each of the two consecutive 2D images. Each interpolated 2D location derived from multiple rolling - shutter cameras can be triangulated to predict a sequence of 3D locations of virtual markers.

[0070] As seen from an exemplary setup 300 of the system 160 schematically shown in FIG. 3, when a marker - less human or animal subject (e.g., patient 305) walks into the doctor's room along a passage or capture volume 313 as indicated by arrow 315, a sequence of 2D images captured by multiple color video cameras 164 can be processed by the system 160 to predict the 3D locations of virtual markers on the marker - less human or animal subject. The multiple color video cameras 164 can be configured to be operable spaced apart from each other along at least a portion of the passage to the doctor's room or capture volume 313 (which can be part of a corridor in a clinic / hospital). In other words, after the patient 305 walks into the doctor's room through the passage or capture volume 313 to meet the doctor, the system 160 has predicted the 3D locations of the virtual markers on the patient 305, and these 3D locations can be used to facilitate information such as an animation (in digitized form) showing the movement of the patient 305. The computer 166 can be located near the multiple color video cameras 164 in the doctor's room or other locations. In the latter case, the predicted / processed information can be remotely transmitted to a computing device or display device located in the doctor's room, or to a mobile device for processing / display. Points D, E simply represent the electrical connection of some of the color video cameras 164 (seen on the right side of FIG. 3) to the computer 166 (seen on the left side of FIG. 3). Other configurations of the color video cameras 164 may be possible. For example, all of the multiple color video cameras 164 can be configured along one side of the passage 313.

[0071] System 160 can be used to facilitate the execution of method 120. Accordingly, system 160 may include elements or components that are the same as or similar to the elements or components of method 120 of FIG. 1B, and accordingly, the similar elements may be those described in the context of method 120 of FIG. 1B, and accordingly, the corresponding description may be omitted here. System 160 may also include some of the elements or components that are the same as or similar to the elements or components of system 140 of FIG. 1C, and accordingly, the same ending numbers are assigned, and the similar elements may be those described in the context of system 140 of FIG. 1C, and accordingly, the corresponding description may be omitted here. For example, in the context of various embodiments, a plurality of color video cameras 164 are the same as the plurality of color video cameras 144 of FIG. 1C.

[0072] Examples of methods 100, 120 and systems 140, 160 will be described in more detail below.

[0073] i. Advantages and improvements Some advantages and improvements of methods 100, 120 and systems 140, 160 according to various embodiments are evaluated to be superior to existing methods / systems.

[0074] Advantages over non-optical motion capture systems Non-optical motion capture systems can be in various forms. One of the most prevalent types in the market uses an inertial measurement unit (IMU) suitable for measuring acceleration, angular velocity, and ambient magnetic fields to approximate the orientation, position, and trajectory of sensors. Ultra-wideband technology can also be integrated for better localization. Another existing tracking technology can use electromagnetic transmitters to track sensors within a spherical capture volume with a small radius of 66 cm. One common drawback among such systems is that the sensors on the subject's body are obtrusive. Attaching sensors to the subject not only takes time for subject preparation but can also cause unnatural movements and / or interfere with movement. Using the markerless motion capture system (e.g., system 160) described in this application eliminates the need for additional items on the subject's body, thereby reducing human involvement / intervention during the process and making the motion capture workflow smoother.

[0075] Advantages over Commercial Marker-based Systems Careful marker placement for full-body motion capture by an expert typically takes at least 30 minutes. If markers are removed from the workflow, one human (expert) can be removed from the workflow and at least 30 minutes can be saved per new subject. After recording, an existing marker-based motion capture system provides only the trajectories of unlabeled markers and is not available for any analysis until the data is post-processed using marker labeling and gap filling. This process is typically done in a semi-automatic manner that takes about one person-hour to process just one minute of recording time. Using a markerless motion capture system (e.g., system 160), since system 160 essentially outputs virtual marker positions with labels, manual post-processing steps are no longer applicable. Since all virtual marker processing is fully automated, one person-hour per minute of recording is saved and replaced by about 20 machine-minutes of recording time, and even faster recording times can be achieved with higher computing power. From a cost perspective, commercial marker-based systems range from $100,000 to $500,000 Singapore dollars, however, all materials in markerless system 160 can cost only about $10,000 Singapore dollars, which is about 10% of a low-end marker-based system. One technical advantage that a data-driven markerless system (e.g., system 160) has over marker-based systems is the way the markless system avoids occlusion. As far as the inventors know, the only way to avoid occlusion in a marker-based system is to add more cameras to ensure that at least two cameras always see one marker at the same time. However, a markerless system (e.g., system 160) can infer virtual markers in occluded regions and thus it requires fewer cameras and generates far fewer gaps in the marker trajectories. Moreover, the use of markers can cause unnatural movements, marker drop during recording, or sometimes skin inflammation. Removing the use of markers simply eliminates at least these problems mentioned above.

[0076] Advantages for a single depth camera system A depth camera is a camera that gives the depth value in each pixel instead of the color value. Thus, only one camera views the 3D surface of the object from one side. This information can be used to estimate the human pose for motion capture purposes. However, the resolution of off-the-shelf depth cameras is relatively low compared to color cameras, and the depth values are usually noisy. As a result, the motion capture results from a single depth camera become relatively inaccurate with special problems due to occlusion. For example, the error in the wrist position from the Kinect SDK and Kinect 2.0 is usually in the range of 3 - 7 cm even without occlusion. Markerless systems (e.g., system 160) produce more accurate results with an average error of less than 2 cm.

[0077] Advantages for freely available open-source human tracking software In particular, there are many open-source projects that share 2D human keypoint detection software for free, such as MediaPipe from Google, OpenVINO from Intel, and Detectron2 from Facebook. These projects also work in a data-driven manner, but they rely on datasets annotated by humans as training data. To demonstrate the advantages of using marker-based annotation M-BA (such as used in method 100 and / or system 140) over manual annotation, the following table 1 and in FIG. 4 a set of preliminary results has been created for comparison.

[0078] [Table 1]

[0079] Figure 4 shows a plot showing the overall accuracy profiles from 12 joints (e.g., shoulders, elbows, wrists, hips, knees, and ankles) from different tools, namely M-BA 402, Thia Markerless 404, Detectron2 406 from Facebook, OpenVINO 408, and MediaPipe 410. As described in Figure 4, M-BA 402, which serves as the basis for methods 100, 120, produces the highest accuracy across the distance thresholds.

[0080] To obtain the results in Table 1 and Figure 4, a system for predicting the 3D locations of virtual markers on a markerless human or animal subject (160) (and described in a similar context as a plurality of color video cameras 164), which involves taking more than 50,000 frames from one male test subject and one female test subject (each frame containing 8 viewpoints) while performing a list of random actions. At the same time, a marker-based motion capture system (Qualisys) (described in a similar context as the optical marker-based motion capture system 142) is used to record the ground truth positions for accuracy comparison. For this system (e.g., 160), the data creation, training, inference, and triangulation methods will be described in the following ii. Technical Description section. The training data used in this experiment contains approximately 2.16 million images from 27 subjects, but the two test subjects are not included in the training data. For MediaPipe, OpenVINO, and Detectron2, the 2D joint positions output from these tools are triangulated and compared with the absolute reference measurements from the marker-based motion capture system in the same manner as that performed for this system (e.g., 160). In the case of MediaPipe, since MediaPipe does not work well when the subject size is relatively small compared to the image size, the images are cropped using the ground truth subject bounding box before the inference of the 2D joint positions. The experimental results show that this method (e.g., 120) produces lower average errors than those of their open-source tools at all six joints. It is important to know that Detectron2 and this method (e.g., 120) use exactly the same neural network architecture. This means that the focus here is on designing better training data that directly reduces the average error by about 28%.

[0081] Advantages over Commercial Markerless Motion Capture Systems One of the existing commercial markerless motion capture systems compared is Theia Markerless. Theia Markerless is a software system that strictly supports videos only from two camera systems, Qualisys' Miqus Video and Sony's RX0M2. The hardware layers for these two camera systems already cost approximately S$63,000 or S$28,000 each (for 8 cameras + computers), along with an additional S$28,000 for software costs. In contrast, the material cost for the entire hardware layer of System 160 is only approximately S$10,000. To evaluate the accuracy, similar tests were also performed on Theia Markerless. Note that the videos used for the evaluation of Theia Markerless were recorded by an expensive Miqus Video global shutter camera system (where all 8 cameras are positioned side by side with multiple color video cameras 164). All tracking and triangulation algorithms are performed in executable software and are not revealed. Despite the more expensive hardware used for Theia Markerless, System 160 is superior in performance at every joint in the evaluation (see Table 1 and Figure 4). One minus aspect of Theia Markerless is the data gap from joint extraction. When the software is unsure about a particular joint in a particular frame, the software decides not to give an answer for that joint. This relatively high percentage of gaps (0.6 - 2.4%) can easily cause further problems in subsequent analysis. On the other hand, System 160 always predicts the results according to various embodiments.

[0082] ii. Technical Description In this section, we will describe the important components, techniques, and ideas for operating a system (described in a context similar to systems 140 and 160). Since ablation studies have not been conducted, it remains unclear how much each idea contributes to the final accuracy. However, reasons are provided for every part of the design.

[0083] Detection Hardware and Camera Configuration (Described in a context similar to method 100 for generating a training dataset for keypoint detection) To collect training data, one marker-based motion capture system (e.g., 142) and multiple color video cameras (e.g., 144) are required. The motion capture system 142 is capable of generating a synchronization signal, and the video camera 144 is capable of taking a shot when a synchronization pulse is received. Since normal video cameras usually operate at a much lower frame rate than motion capture systems, hardware clock multipliers and dividers can be used to enable synchronization at two different frame rates.

[0084] All video cameras are set approximately 170 cm above the ground and face the central capture area. To minimize the variation of data that can be controlled during training and system deployment, it is important to have training images taken from substantially the same height. Additionally, 170 cm can be the height that a common tripod can reach without the need to construct a framework for mounting the camera.

[0085] To support accurate calibration or pre-calibration, each video camera (e.g., 144) can be provided with three visible (white) LEDs and is equipped with at least one such LED 500 as in the example shown in FIG. 5. With these LEDs 500, a normal video camera that senses only light in the visible spectrum can see a circular retroreflective marker as a detectable bright spot on the captured (captured) image. When the marker-based motion capture system 142 sees this marker in 3D space and the video camera 144 simultaneously sees this marker in 2D on the image, they form a 2D-3D correspondence pair. Sufficient collection of these correspondence pairs across the capture volume can be used to calculate the accurate camera pose (extrinsic parameters) and fine-tune the intrinsic camera parameters. One important camera setting is the exposure time. The exposure needs to be short enough to minimize motion blur. During video recording, the target object is a human. Therefore, the exposure time is selected to be 2 -8 seconds or about 3.9 ms. At this timing, the edges of the human silhouette during extremely fast movements are still sharp. During calibration, the target object is a retroreflective marker that can move faster than a human body. Therefore, the exposure time is selected to be 2 -10 seconds or about 1 ms. At this exposure time the capture environment is significantly dark, but the reflection from the marker is still bright enough to be detected. The video camera 144 can use a global shutter sensor or a rolling shutter sensor -wo Since global shutter types are generally used for this type of application example and rolling shutter cameras require additional modeling and calculations, the following description focuses more on the integration of rolling shutter cameras in this application.

[0086] Rolling Shutter Camera Model In this section, we will describe the rolling shutter model developed for the FSCAM_CU135 camera from e-con System. However, since most rolling shutter cameras operate in a similar manner, this model may be applicable to most rolling shutter cameras. In the hardware trigger mode of FSCAM, a rising edge pulse is used to trigger image capture. When the trigger pulse is received, the camera sensor undergoes a delay of b seconds before starting the readout. The camera sensor then reads the pixels row by row starting from the top, with a line delay of d seconds per row until the last row is reached. The exposure for the next frame is automatically started based on a predetermined timing relative to the previous trigger. The readout for the next image starts in the same manner from the next rising edge pulse. The trigger readout delay (b) and line delay (d) depend on the camera model and configuration. For FSCAM operating at a resolution of 1920×1440, b and d are approximately 5.76×10 -4 seconds and 1.07×10 -5 seconds respectively. This rolling shutter model 600 developed for the FSCAM_CU135 camera is shown in Figure 6. In this model 600, it is assumed that all pixels in the same row always operate simultaneously. According to Figure 6, the center line of the exposure zone (middle exposure line) represents a linear relationship between the pixel row and time. This means that when an object is observed in a specific pixel row of a specific video frame, the exact time (t) to capture that object can be calculated. This relationship can be formulated as Equation 1 t=T i +b - e / 2 + dv, Equation 1 and can be expressed as follows, where T i is the trigger time of video frame (i), e is the exposure time, and v is the pixel row.

[0087] As can be seen in FIG. 6, the gray area is the time during which the pixel row was exposed to light. Note that the first row of the image starts from the topmost row. This model 600 is used as follows.

[0088] Interpolation of 2D marker trajectories at trigger time: When multiple rolling shutter cameras observe the same object (such as a marker), the object does not project onto the same pixel row across all cameras, so their observation times usually do not match. This time mismatch causes a large error when the process requires observations from multiple cameras. For example, triangulation of 2D observations from multiple cameras assumes that those observations are from the same instant, but if not, especially when the object is moving fast, the triangulation can give a large error. To improve the process during the resulting triangulation, the rolling shutter model can be used to estimate the 2D position of the observed marker or object at the trigger time so that observations can be obtained at exactly the same timing across all cameras. The calculations from the rolling shutter model 600 are shown in the graphical representation 700 of FIG. 7. In FIG. 7, each black dot represents an observation point on one video frame. These dots always remain on the intermediate exposure line by the rolling shutter model 600. For each observation in one particular frame, the known pixel row (v) can be used to solve for the observation time (t) from Equation 1. When the observation times are known in two consecutive video frames (t1 and t2), linear interpolation of the 2D position at the intermediate trigger time can be easily performed. In other words, to estimate the position of the observed 2D trajectory at the trigger time (T m ), first the observation times (t1 and t2) are calculated from the observation rows (v1 and v2). Using t1 and t2, interpolation of the 2D position at T m can be performed. The interpolated value is used in triangulation as if it were from a global shutter camera.

[0089] Projection of 3D marker trajectories onto 2D images: In training data generation, one crucial step is to generate the 2D locations of 40 body markers for each video frame. When the camera uses a global shutter sensor, the observation time is precisely known for the entire image. This time can be used to interpolate the 3D positions from the marker trajectories and project it directly onto the video camera. In contrast, the observation time from a rolling shutter depends on the result (row) of the projection, which is not known until the projection is performed. Therefore, the new projection method 800 shown in FIG. 8 is developed.

[0090] First, the target 3D trajectories from a marker-based mocap system (e.g., the optical motion capture system 142) are directly projected onto the target camera (e.g., each of the plurality of color video cameras 144) for each sample. In other words, when the 3D marker trajectories from the marker-based mocap system are projected onto the camera, they can be plotted as shown in FIG. 8 where each point represents one sample. For each sample, the projection gives a pixel row (v), and the time of that sample is also known. Since the sample frequency of the marker-based mocap system is relatively higher, if the dots of the projection are connected in a plot like FIG. 8, there are several lines or adjacent pairs that intersect the intermediate exposure lines (from Equation 1). Since any two consecutive samples can form a linear equation (the line connecting the two dots), if this equation intersects any of the intermediate exposure lines from Equation 1 within its own time section, the solution of these two linear equations tells the exact time of interpolation. This intersection time is used to interpolate the 3D positions from the trajectory. The interpolated 3D positions can then be projected onto the camera (or image) to obtain an accurate projection that matches the observation and can be used in training.

[0091] Video camera calibration Initialization of Camera Intrinsic Parameters: For each video camera, a standard process with the OpenCV library is used to estimate the camera intrinsic parameters. A 10×7 checkerboard with 35mm blocks is kept stationary in front of the camera in 30 different poses to capture 30 different images. Then, cv2.findChessboardCorners is used to find the corners of the 2D checkerboard on each image. Next, cv2.calibrateCamera is used to obtain an estimate of the intrinsic matrix and distortion coefficients. These values are fine-tuned in the next stage of calibration.

[0092] Camera Calibration for Training Data Collection: Since the extrinsic parameter solution from the subsequent calibration is in the marker-based mocap reference frame, in this calibration, more specifically, in the pre-calibration process, it is assumed that the marker-based mocap system (e.g., the optical motion capture system 142) is already calibrated. This calibration is performed by waving a wand with one retroreflective marker at the tip for about 2 - 3 minutes throughout the capture volume. This marker is captured by both the marker-based motion capture system and a video camera with white LEDs (e.g., multiple color video cameras 144). From the perspective of the marker-based motion capture system, it records the 3D trajectory of the marker. From the perspective of the video camera, it looks at a series of dark images with bright spots that can be extracted as 2D positions on each image.

[0093] To extract this 2D position, the algorithm searches for bright pixels across the image to converge the location at the centroid of the bright pixel cluster and applies the mean shift algorithm to that location.

[0094] When the camera uses a global shutter sensor, 2D-3D correspondence pairs are collected simply by linearly interpolating the 3D position from the 3D marker trajectory using the time midway through the exposure interval of the video camera frame. Then, by applying the cv2.calibrateCamera function to that set of correspondence pairs, the extrinsic camera parameters are given and the intrinsic camera parameters are fine-tuned.

[0095] However, since not all pixel rows are captured simultaneously, this cannot be done directly on a rolling shutter camera. The time of the observed 2D marker on the video frame varies according to the row of pixels where it is seen. Equation 1 is used to calculate the time of the 2D marker observation, which is used to linearly interpolate the 3D position from the 3D marker trajectory to form one 2D-3D corresponding (or corresponding) pair. Then, by applying the cv2.calibrateCamera function to that set of correspondence pairs, the extrinsic camera parameters are given and the intrinsic camera parameters are fine-tuned.

[0096] The method described works well when there are no other bright or reflective items in the camera's field of view. However, that assumption is not very realistic since a motion capture environment usually includes many light sources, a computer screen, and LEDs from a video camera on the opposite side. Therefore, additional procedures are required to handle these noises.

[0097] For example, 5 seconds of video recording is performed immediately before the wand undulation step to find bright pixels in the image and mask those bright pixels in every frame before searching for markers in the wand undulation recording. This removes static bright areas in the camera's field of view, but dynamic noise from moving shiny objects such as watches or glasses is included in the 2D-3D correspondence pool. To remove those dynamic noises from the 2D-3D correspondence pool, a method based on the idea of Random Sample Consensus (RANSAC) for removing outliers from model fitting has been developed. In this method, it is assumed that the noise occurs less than 5% from the samples of all 2D-3D correspondence pairs so that most can correctly form a consensus.

[0098] This method is described as follows. (a) Randomly sample 100 2D-3D correspondence pairs from the pool. (b) Use those 100 correspondence pairs to calculate the camera parameters using cv2.calibrateCamera. (c) Use the calculated camera parameters to project all 3D points from all pairs in the pool to observe the Euclidean error between the projection and the 2D observation. Pairs with an error of less than 10 pixels are classified as good pairs. (d) All good pairs from the latest round of classification are also used to calculate the camera parameters using cv2.calibrateCamera. (e) Repeat steps (c) and (d) until the set of good points remains the same in subsequent iterations, i.e., until the model converges.

[0099] If the first 100 samples contain a large number of noisy pairs, the calculated camera parameters will be inaccurate and will not match a large number of corresponding pairs in the pool. In this case, the model converges using a small number of good pairs.

[0100] On the other hand, if the first 100 samples contain only valid pairs, the calculated camera parameters will be quite accurate and will match a large number of valid pairs in the pool. In this case, the number of good pairs is expanded to cover all valid points, but the noisy pairs remain excluded as they do not match the valid consensus.

[0101] To achieve the latter case, processes (a)-(e) are repeated 200 times to select the final model with the maximum number of good pairs. By evaluation, this method of noise removal can reduce the average projection error to the sub-pixel level, which is ideal for data collection.

[0102] Extrinsic Camera Calibration for System Deployment: In the actual deployment of a system (e.g., system 160), there is no marker-based motion capture system to provide 3D information of marker trajectories to collect 2D-3D correspondences for camera calibration. Therefore, alternative extrinsic calibration methods can be used. If the cameras are not equipped with LEDs, a checkerboard can be simultaneously captured by two cameras to calculate the relative transformation between the two cameras using the cv2.StereoCalibrate method. When the relative transformations between all cameras in the system are known, their extrinsic parameters are also fine-tuned using the Levenberg-Marquardt optimization in this case to obtain the final result. (Explained in a similar context as externally calibrating a color video camera in a method for predicting the 3D location of virtual markers on a markerless human or animal subject 120) To facilitate this calibration process, multiple checkerboards can be used in the same environment by adding unique Aruco markers into the checkerboard 900 seen as a Charuco board in FIG. 9. These Charuco boards can be detected using the cv2.aruco.estimatePoseCharucoBoard function with their board identification information.

[0103] If the cameras are equipped with LEDs, it is possible to extend the calibration to be more accurate in a larger volume using a wand with reflective markers and bundle adjustment optimization techniques.

[0104] Training Data Collection and Preprocessing This section can be described in a context similar to method 100 for generating a training dataset for keypoint detection, and describes how the dataset is collected and preprocessed before training. The training data (or training dataset) includes three important elements: an image from a video camera, the positions of 2D keypoints on each image, and the bounding box of the target object.

[0105] Marker set: A set of 40 markers is selected from the marker set in the RRIS's Ability Data protocol (see P. Liang et al., "An A Asian-centric human movement database capturing activities of daily living", Scientific Data, vol. 7, no. 1, pp. 1-13, 2020). Since the placement of all clusters is not consistent across multiple subjects and their large sizes cause difficulties in the impainting step later, all clusters are removed. There are 4 markers on the head (RTEMP, RHEAD, LHEAD, LTEMP), 4 markers on the torso (STER, XPRO, C7, T10), 4 markers on the pelvis (RASIS, LASIS, LPSIS, RPSIS), 7 markers on each upper limb (ACR, HLE, HME, RSP, USP, CAP, HMC2), and 7 markers on each lower limb (FLE, FME, TAM, FAL, FCC, FMT1, FMT5). The marker placement task is standardized according to the bone landmarks and is most preferably performed by trained people.

[0106] Marker Projection for a Rolling Shutter Camera: All 3D marker trajectories are projected onto each video camera using the projection method described in the above section that describes the projection of 3D marker trajectories onto 2D images under the rolling shutter camera model. The results from the 2D projection are 2D keypoints for training. For example, refer to step 104 of method 100.

[0107] Marker Removal: Images captured from a video camera always contain visible marker blobs that can cause problems for the model learned during inference. When the model sees a pattern where the predicted position of the keypoints always lands on the gray blob from the visible marker, the model memorizes this pattern and always looks for the gray blob as an important feature to identify the position of the marker itself. This overfitting can degrade performance in markerless use when there are no longer markers on the body. Therefore, the video data is created as if there were no markers on the subject. This can be done by using an image inpainting technique that uses an adversarial generative network (GAN) to replace the pixel color in the target area by noticing the surrounding context. In this case, DeepFillv2 is used to remove the markers. To remove the markers, the pixels occupied by the markers are removed from the list. This can be done automatically by taking the 2D projection (e.g., step 104 of method 100) and drawing a 2D radius according to the distance between the camera and the marker with some additional margin to cover the base and shadow of the marker.

[0108] Removing Non-targets: With multiple video cameras facing all directions, it is difficult to avoid non-target humans in the field of view. Since those non-target humans are not wearing markers, they are not labeled and are interpreted as background during the training process, which can cause confusion in the model. Therefore, those non-target humans are automatically detected by the default human detection from Detectron2 and blurred by the smooth edge.

[0109] Bounding Box Formulation: One important piece of information required by the training process is the 2D bounding box around each human target. This 2D bounding box in the form of a simple rectangle covers not only all the projected marker positions but also the complete silhouette of all body parts. Therefore, the formulation is developed by expanding that coverage by different amounts up to the point where the coverage of each marker covers adjacent body parts. For example, there are no markers on the fingers, so elbow markers, wrist markers, and hand markers are used to approximate the volume that the fingers can reach. Then, those 3D points on the surface of that volume are projected onto each camera to approximate the bounding box. See, for example, step 108 of method 100.

[0110] Neural Network Architecture and Training Framework The keypoint detection version of Mask-RCNN with a Feature Pyramid Network (FPN) as the feature extraction backbone is used as the neural network architecture. Since the network already has a PyTorch implementation on the Detectron2 projection repository, (as described above in the Training Data Collection and Preprocessing section, and also refer to steps 125 and 124 of method 120) the set of keypoints from the joint centers can be changed to a set of 40 markers, and modifications can be made to enable the training images to be loaded from video files. The data loader module is also modified to use shared memory across all worker processes to reduce redundancy in memory usage and make it possible to greatly increase the size of the training data.

[0111] Strategic triangulation After training, the model is capable of predicting the 2D locations of all 40 markers from an image of the object without markers. For example, refer to step 126 of method 120. In some specific situations, such as when the object is half-cropped by the camera's field of view, some of the markers may not give a location output because their confidence levels are too low. In the case of a rolling shutter camera, as described in the above section that explains the interpolation of the 2D marker trajectories at the trigger time under the rolling shutter camera model, the 2D locations used for triangulation are the interpolated results between two consecutive frames to obtain the location at the trigger time. If a marker from one of the adjacent frames is not available for interpolation, the camera should be treated as being unavailable for that marker in that frame.

[0112] In an ideal situation where the predictions output from all available cameras are fairly accurate, triangulation of the results from all cameras can be performed using direct linear transformation. One 2D location on an image from one camera can be represented by a 3D ray pointing out from the camera origin to the point. Direct linear transformation directly calculates the 3D point, which is the virtual intersection point of all those rays. In this ideal case, the distance between the 3D point and any ray is likely to be small (i.e., less than 10 cm), and the solution can be easily accepted.

[0113] However, in reality, the predictions in a few cameras can be incorrect. Sometimes, for example, because the torso is obstructing, some cameras may not see the exact position of the wrist. Sometimes, some cameras can be confused between the left and right sides of the body. To make the triangulation more robust, this method rejects the contributions from cameras that do not agree with the consensus.

[0114] A method for triangulating one marker in one specific frame can be performed as follows.

[0115] (a) List all available cameras (cameras that can provide the 2D location of the target marker).

[0116] (b) Triangulate all available cameras to obtain the 3D location. Triangulation can be performed using the commonly used DLT method. Optionally, if a confidence score for each 2D marker prediction is given, the triangulation method can be significantly improved using the weighted triangulation formula (see Equation 2) described below in the section on new weighted triangulation.

[0117] (c) Identify the camera that gives the maximum distance between the triangulated 3D point and the ray from that camera among the cameras in the available list. If the maximum distance is less than 10 cm, the triangulated one is accepted. Otherwise, that camera is removed from the list of available cameras.

[0118] (d) Repeat steps (b) and (c) until a solution is accepted. If the number of cameras in the list is less than two, there is no solution for that marker in this frame.

[0119] Using this method, the maximum number of triangulations performed per marker per frame is exactly n - 1, where n is the number of cameras. These n - 1 calculations are much faster than trying all possible combinations of triangulations that would require n - 2n - 1 calculations.

[0120] New weighted triangulation In a neural network that performs 2D keypoint localization, it may be common to also generate a confidence score associated with each 2D location output. For example, the keypoint detection version of Mark-RCNN generates a heatmap of confidence inside the bounding box for each keypoint. Then, the 2D location with the highest confidence in the heatmap is selected as the answer. In this case, the confidence score at the peak is the associated score for that 2D keypoint prediction. In normal triangulation, that confidence score is usually ignored. However, the weighted triangulation formula allows for the use of the score as a weight for triangulation to improve the accuracy of triangulation, as described below.

[0121] Weighted triangulation formula: The triangulated 3D position (P) is P = (Σ i w i Q i ) -1 (Σ i w i Q i C i ) (Equation 2 can be derived as, where w i is the weight, or the confidence score of the i-th ray from the i-th camera, and C iis the 3D camera location associated with the i-th ray, U i is a 3D unit vector representing the back-projected direction associated with the i-th ray, I3 is a 3×3 identity matrix being given

Number

[0122] The direction vector (Ui) of each back-projected ray is 1) Undistorting the 2D observations using cv2.undistortPointsIter on the normalized coordinates, and 2) Forming a 3D direction vector in the camera reference frame [x_undistorted, y_undistorted, 1] T and, 3) Rotating the direction into the global reference frame using the current estimate of the camera orientation, and 4) Normalizing the vector to obtain the unit vector (Ui) is calculated by.

[0123] This formula is derived by minimizing the weighted sum of squares of the distances between the triangulated points and all the rays. So, when the prediction reliability is low, the influence in triangulation becomes smaller, and it becomes possible to bring the triangulated points closer to the rays with higher prediction reliability, resulting in improved overall accuracy.

[0124] iii. Commercial application examples The potential customers of the present invention are those who seek a non-real-time markerless human motion capture system. The potential customers can be scientists who want to study human movements, animators who want to create animations from human movements, or hospitals / clinics that want to generate objective diagnoses from the movements of patients.

[0125] The advantage in reducing the time and manpower used to run a motion capture system is that the patient can perform a short motion capture and have the opportunity to meet with a physician who has the analysis results within the same time, thus opening up the opportunity for clinicians to adopt this technology for objective diagnosis / analysis from the patient's movements.

[0126] The present invention has been illustrated and described in detail with reference to specific embodiments, but it should be understood by those skilled in the art that various changes in form and detail can be made therein without departing from the spirit and scope of the present invention as defined by the appended claims. The scope of the present invention is, therefore, indicated by the appended claims, and all changes that come within the meaning and range of equivalents of the claims are, therefore, included.

Claims

1. A method for generating a training dataset for keypoint detection, the method comprising: Based on a plurality of markers, each captured as a 3D trajectory by an optical marker-based motion capture system, each marker being placed on a bone landmark of a human or animal subject or a keypoint of an object, and the human or animal subject or the object being substantially simultaneously captured by a plurality of color video cameras over a time period as a sequence of 2D images, For each marker, projecting the 3D trajectory onto each of the 2D images to determine a 2D location in each 2D image; For each marker, interpolating a 3D position for each of the 2D images based on each of the 2D locations in the sequence of 2D images and the exposure relationship time of the plurality of color video cameras; For each 2D image, generating a 2D bounding box around the human or animal subject or the object based on the interpolated 3D positions of each of the plurality of markers and an extended volume derived from two or more of the markers having an anatomical or functional relationship with each other; Generating the training dataset including at least one 2D image selected from the sequence of 2D images, the determined 2D locations of each marker in at least one selected 2D image, and the generated 2D bounding box for at least one selected 2D image. A method as described above.

2. The method according to claim 1, wherein the plurality of markers captured as the 3D trajectories respectively and the human or animal subject or the object captured substantially simultaneously as the sequence of 2D images over the time period are coordinated using a synchronization signal communicated to the plurality of color video cameras by the optical marker-based motion capture system.

3. Before projecting the 3D trajectory, identifying the captured 3D trajectory having a label representing the bone landmark or key point on which the marker is placed, wherein for each marker, the label is propagated with each determined 2D location such that in the generated training dataset, each determined 2D location of each marker corresponds to a contained label, further comprising identifying the captured 3D trajectory, the method according to claim 1 or 2.

4. After projecting the 3D trajectory onto each of the 2D images to determine the 2D location in each 2D image, In each 2D image, for each marker, drawing a 2D radius on the determined 2D location according to a distance having a predefined margin between the color video camera and the marker to form an enclosed area, and applying a learning-based context-aware image inpainting technique to the enclosed area to remove the marker blob from the 2D location, further comprising the method according to claim 1 or 2.

5. The method according to claim 4, wherein the learning-based context-aware image inpainting technique comprises an adversarial generation network-based context-aware image inpainting technique.

6. The method according to claim 1 or 2, wherein the plurality of color video cameras are a plurality of global shutter cameras.

7. The method according to claim 6, wherein the exposure relationship time is intermediate the exposure times for capturing each 2D image using each global shutter camera.

8. Each global shutter camera comprises at least one visible light emitting diode operable to facilitate the perception of a retroreflective marker coupled to a wand as a bright spot detectable, the plurality of global shutter cameras, The wand is in a continuous wave-like form, and based on the retroreflective markers captured by the optical marker-based motion capture system as a 3D trajectory covering the target capture volume, and the retroreflective markers substantially simultaneously captured by each global shutter camera as a sequence of 2D calibration images during a time period, For each 2D calibration image, searching for bright pixels and extracting the 2D calibration position of the retroreflective marker by scanning across the entire 2D calibration image to identify the 2D location of the bright pixels, and applying an iterative algorithm at the 2D location of the searched bright pixels to converge the 2D location at the centroid of the bright pixel cluster, Based on the middle of the exposure time and the 3D trajectory in each 2D calibration image, linearly interpolating the 3D calibration position for each of the 2D calibration images, Forming a plurality of 2D-3D correspondence pairs for at least a portion of the plurality of 2D calibration images, each 2D-3D correspondence pair including the converged 2D location and the interpolated 3D calibration position for each of the at least a portion of the plurality of 2D calibration images, forming a plurality of 2D-3D correspondence pairs, Determining extrinsic camera parameters and applying a camera calibration function to the plurality of 2D-3D correspondence pairs to fine-tune the intrinsic camera parameters of the plurality of global shutter cameras The method according to claim 7, which is pre-calibrated by.

9. The plurality of color video cameras are a plurality of rolling shutter cameras, and projecting the 3D trajectory onto each of the 2D images, For each 2D image captured by each rolling shutter camera, determining an intersection time from the intersection between a first line connecting the projected 3D trajectory over the time period and a second line representing the moving middle of the exposure time to capture each pixel row of the 2D image, For each 2D image captured by each rolling shutter camera, interpolating a 3D intermediate position based on the intersection time to obtain a 3D interpolated trajectory from the sequence of 2D images, For each marker, projecting the 3D interpolated trajectory onto each of the 2D images to determine the 2D location in each 2D image The method according to claim 1 or 2, further comprising.

10. The method according to claim 9, wherein the exposure relationship time is the intersection time.

11. Each rolling shutter camera comprises at least one visible light emitting diode operable to facilitate the retroreflective marker coupled to the wand being perceived as a bright spot detectable, and the plurality of rolling shutter cameras, Based on the retroreflective marker captured by the optical marker-based motion capture system as a 3D trajectory in which the wand is continuous and wavy and covers a target capture volume, and the retroreflective marker captured substantially simultaneously by each rolling shutter camera as a sequence of 2D calibration images during a time period, For each 2D calibration image, searching for bright pixels and extracting the 2D calibration position of the retroreflective marker by scanning across the 2D calibration image to identify the 2D location of the bright pixels; and applying an iterative algorithm at the 2D location of the searched bright pixels to converge the 2D location at the 2D centroid of the bright pixel cluster. Interpolating a 3D calibration position from the 3D trajectory covering the target capture volume based on the observation times of the 2D centroids from the plurality of rolling shutter cameras, wherein the observation time of each 2D centroid of each bright pixel cluster from each 2D calibration image i is T i +b - e / 2 + dv Calculated by Here, T i is the trigger time of the i-th 2D calibration image, b is the trigger readout delay received by the rolling shutter camera, e is the exposure time set for the rolling shutter camera, d is the line delay received by the rolling shutter camera, v is the pixel row of the 2D centroid of the bright pixel cluster, Interpolating the 3D calibration position Forming a plurality of 2D-3D correspondence pairs for at least a portion of the plurality of 2D calibration images, each 2D-3D correspondence pair including the converged 2D location for each of the at least a portion of the plurality of 2D calibration images and the interpolated 3D calibration position, forming the plurality of 2D-3D correspondence pairs Determining extrinsic camera parameters and applying a camera calibration function to the plurality of 2D-3D correspondence pairs to fine-tune the intrinsic camera parameters of the plurality of rolling shutter cameras The method according to claim 9, which is pre-calibrated thereby.

12. The method according to claim 8, wherein the iterative algorithm is a mean shift algorithm.

13. The retroreflective marker captured as the 3D trajectory covering the target capture volume and the retroreflective marker captured substantially simultaneously as the sequence of 2D calibration images are coordinated by the optical marker-based motion capture system using a synchronization signal communicated to the plurality of color video cameras.

14. The method according to claim 1 or 2, wherein the marker includes a retroreflective marker.

15. A method for predicting the 3D location of a markerless human or animal subject or a virtual marker on a markerless object, the method comprising Based on the markerless human or animal subject or the markerless object captured by a plurality of color video cameras as a sequence of 2D images For each 2D image captured by each color video camera, predicting a 2D bounding box using a trained neural network For each 2D image, generating a plurality of heatmaps with reliability scores by the trained neural network Each heatmap being for the 2D localization of a virtual marker of the markerless human or animal subject or the markerless object The trained neural network being trained using at least the training dataset generated by the method according to claim 1 Generating a plurality of heatmaps with reliability scores For each heatmap, selecting the pixel with the highest reliability score and associating the selected pixel with the virtual marker, thereby determining the 2D location of the virtual marker, wherein for each heatmap, the reliability score indicates the probability of having a virtual marker associated with different 2D locations in the predicted 2D bounding box, and determining the 2D location of the virtual marker; triangulating each determined 2D location to predict a sequence of 3D locations of the virtual marker based on the sequence of 2D images captured by the plurality of color video cameras; A method comprising. [

16. ] The method according to claim 15, wherein triangulating comprises weighted triangulation of the respective 2D locations of the virtual marker based on each of the reliability scores as weights for triangulation. [

17. ] The weighted triangulation includes deriving each predicted 3D location of the virtual marker using the formula (Σ i w i Q i ) -1 (Σ i w i Q i C i ) wherein where assuming N is the total number of color video cameras, i is 1, 2,..., N, w i is the weight for the triangulation or the reliability score of the i-th light ray from the i-th color video camera, C i is the 3D location of the i-th color video camera associated with the i-th light ray, U i is a 3D unit vector representing the back-projected direction associated with the i-th ray, I 3 is a 3×3 identity matrix given 【Number 1】 The method according to claim 16. [

18. ] The method according to any one of claims 15 to 17, wherein the plurality of color video cameras are a plurality of global shutter cameras. [

19. ] The plurality of color video cameras are a plurality of rolling shutter cameras, and the method comprises, prior to triangulating the respective 2D locations to predict the sequence of 3D locations of the virtual marker, determining an observation time for each rolling shutter camera based on the determined 2D locations in two consecutive 2D images, wherein the observation time is T i +b - e / 2 + dv calculated by Here, T i is the trigger time of each of the two consecutive 2D images, where b is the trigger readout delay of the rolling shutter camera, e is the exposure time set for the rolling shutter camera, d is the line delay of the rolling shutter camera, v is the pixel row of the 2D location in each of the two consecutive 2D images, determining an observation time for each rolling shutter camera; Interpolating the 2D location of the virtual marker at the trigger time based on the observation time, wherein triangulating the respective 2D locations comprises triangulating the respective interpolated 2D locations derived from the plurality of rolling shutter cameras, and interpolating the 2D location of the virtual marker The method according to any one of claims 15 to 17, further comprising.

20. Based on one or more checkerboards simultaneously captured by the plurality of color video cameras Calculating a relative transformation between every two of the plurality of color video cameras for every two of the plurality of color video cameras Applying an optimization function to fine-tune the extrinsic camera parameters of the plurality of color video cameras when the plurality of color video cameras each have the calculated relative transformation The method according to any one of claims 15 to 17, further comprising externally calibrating the plurality of color video cameras by.

21. The method according to claim 20, wherein the one or more checkerboards include unique markings.

22. Each color video camera includes at least one visible light emitting diode operable to facilitate the perception of a retroreflective marker coupled to a wand as a bright spot detectable, the method comprising Based on the retroreflective marker captured by the plurality of color video cameras as a sequence of 2D calibration images, wherein the wand is continuous and wavy Applying an optimization function to the captured 2D calibration images to fine-tune the extrinsic camera parameters of the plurality of color video cameras The method according to any one of claims 15 to 17, further comprising externally calibrating the plurality of color video cameras by.

23. A computer program adapted to perform the method according to any one of claims 1, 2, 15 to 17.

24. A non-transitory computer-readable medium comprising instructions that, when executed on a computer, cause the computer to perform the method according to any one of claims 1, 2, 15 to 17.

25. A data processing apparatus including means for executing the method according to any one of claims 1, 2, 15 to 17.

26. A system for generating a training data set for keypoint detection, the system comprising: An optical marker-based motion capture system configured to capture a plurality of markers over a time period, each marker being placed on a bone landmark of a human or animal subject or a keypoint of an object and captured as a 3D trajectory; an optical marker-based motion capture system; A plurality of color video cameras configured to capture the human or animal subject or the object over the time period as a sequence of 2D images; A computer And the computer is Receiving the sequence of 2D images captured by the plurality of color video cameras and each of the 3D trajectories captured by the optical marker-based motion capture system; For each marker, projecting the 3D trajectory onto each of the 2D images to determine a 2D location in each 2D image; For each marker, interpolating a 3D position for each of the 2D images based on each of the 2D locations in the sequence of 2D images and the exposure relationship time of the plurality of color video cameras; For each 2D image, based on the interpolated 3D positions of each of the plurality of markers and an extended volume derived from two or more of the markers having an anatomical or functional relationship to each other, generating a 2D bounding box around the human or animal subject or the object; Generating the training data set including at least one 2D image selected from the sequence of 2D images, the determined 2D locations of each marker in at least one selected 2D image, and the generated 2D bounding box for at least one selected 2D image A system configured to perform.

27. The system according to claim 26, further comprising a synchronization pulse generator communicating with the optical marker-based motion capture system and the plurality of color video cameras, the synchronization pulse generator being configured to receive a synchronization signal from the optical marker-based motion capture system to coordinate the human or animal subject or the object to be captured substantially simultaneously by the plurality of color video cameras.

28. The system according to claim 26 or 27, wherein the optical marker-based motion capture system comprises a plurality of infrared cameras.

29. The system according to claim 28, wherein the plurality of color video cameras and the plurality of infrared cameras are configured to be spaced apart from each other and along at least a path to be taken by the human or animal subject or the object, or at least substantially surround a capture volume of the human or animal subject or the object.

30. The system according to claim 26 or 27, wherein the 3D trajectory is identifiable using a label representing the bone landmark or key point on which the marker is placed, and for each marker, the label is propagated with each determined 2D location such that in the generated training dataset, each determined 2D location of each marker includes a corresponding label.

31. The system according to claim 26 or 27, wherein the computer is further configured to draw a 2D radius on the determined 2D location for each marker according to a distance having a predefined margin between the color video camera and the marker to form an enclosed area in each 2D image, and to apply a learning-based context-aware image inpainting technique to the enclosed area to remove the marker blob from the 2D location.

32. The system according to claim 31, wherein the learning-based context-aware image inpainting technique includes an adversarial generation network-based context-aware image inpainting technique.

33. The system according to claim 26 or 27, wherein the plurality of color video cameras are a plurality of global shutter cameras.

34. The plurality of color video cameras are a plurality of rolling shutter cameras, and the computer further For each 2D image captured by each rolling shutter camera, to connect the projected 3D trajectory over the time period to obtain a first line for capturing each pixel row of the 2D image, and a second line representing the moving midpoint of the exposure time, determining an intersection time from the intersection of the two lines; For each 2D image captured by each rolling shutter camera, interpolating a 3D intermediate position based on the intersection time to obtain a 3D interpolated trajectory from the sequence of 2D images; For each marker, projecting the 3D interpolated trajectory onto each of the 2D images to determine a 2D location in each of the 2D images The system according to claim 26 or 27, configured to perform the above.

35. A system for predicting the 3D location of a markerless human or animal subject or a virtual marker on a markerless object, the system comprising: A plurality of color video cameras configured to capture the markerless human or animal subject or the markerless object as a sequence of 2D images; A computer The computer is configured to: Receive the sequence of 2D images captured by the plurality of color video cameras; For each 2D image captured by each color video camera, use a trained neural network to predict a 2D bounding box; For each 2D image, use the trained neural network to generate a plurality of heatmaps with confidence scores, Each heatmap being for the 2D localization of the virtual marker on the markerless human or animal subject or the markerless object, The trained neural network being trained using at least the training dataset generated by the method according to claim 1 to generate a plurality of heatmaps. For each heatmap, selecting the pixel with the highest said reliability score and associating the selected pixel with the virtual marker to determine the 2D location of the virtual marker, wherein for each heatmap, the reliability score indicates the probability of having the virtual marker associated with different 2D locations in the predicted 2D bounding box, and associating the selected pixel with the virtual marker; triangulating each determined 2D location based on the sequence of 2D images captured by the plurality of color video cameras to predict a sequence of 3D locations of the virtual marker; A system configured to perform the above. **Claim 36** The system according to claim 35, wherein each of the 2D locations of the virtual marker should be triangulated based on each of the reliability scores as weights for triangulation. **Claim 37** The triangulation includes deriving each predicted 3D location of the virtual marker using the formula (Σ i w i Q i ) -1 (Σ i w i Q i C i ) where assuming N is the total number of color video cameras, i is 1, 2,..., N, w i is the weight for the triangulation or the reliability score of the ith light ray from the ith color video camera, C i is the 3D location of the i-th color video camera associated with the i-th light ray, U i is a 3D unit vector representing the back-projected direction associated with the i-th ray, I 3 is a 3×3 identity matrix given 【Number 2】 The system according to claim 36. **Claim 38** The system according to any one of claims 35 to 37, wherein the plurality of color video cameras are a plurality of global shutter cameras. **Claim 39** The plurality of color video cameras are a plurality of rolling shutter cameras, and the computer is further configured to determine the observation time for each rolling shutter camera based on the determined 2D locations in two consecutive 2D images, wherein the observation time is T i +b - e / 2 + dv, calculated by Here, T i is the trigger time for each of the two consecutive 2D images, b is the trigger readout delay of the rolling shutter camera, e is the exposure time set for the rolling shutter camera, d is the line delay of the rolling shutter camera, v is the pixel row of the 2D location in each of the two consecutive 2D images, determining the observation time for each rolling shutter camera; Interpolating the 2D location of the virtual marker at the trigger time based on the observation time, wherein each interpolated 2D location derived from the plurality of rolling shutter cameras should be triangulated to predict a sequence of 3D locations of the virtual marker, and interpolating the 2D location of the virtual marker The system according to any one of claims 35 to 37, configured to perform the above.

40. The plurality of color video cameras are spaced apart from each other and operably configured along at least a portion of the passage to the doctor's room so that when a human or animal subject without a marker walks into the doctor's room along the passage, the system processes a sequence of 2D images captured by the plurality of color video cameras to predict the 3D location of the virtual marker on the human or animal subject without a marker. The system according to claim 37.

41. The method according to claim 11, wherein the iterative algorithm is a mean shift algorithm.

42. The retroreflective marker captured as the 3D trajectory covering the target capture volume and the retroreflective marker captured substantially simultaneously as the sequence of 2D calibration images are coordinated using a synchronization signal communicated to the plurality of color video cameras by the optical marker-based motion capture system. The method according to claim 11.

Citation Information

Patent Citations

  • Device and method for body posture evaluation

    CN112102947A

  • Method and apparatus for acquiring joint position, and method and apparatus for acquiring motion

    JP2020042476A

  • Method, System and Device for Direct Prediction of 3D Body Poses from Motion Compensated Sequence

    US20170316578A1

  • Motion database structure, motion data normalization method for the motion database structure, and searching device and method using the motion database structure

    WO2009145071A1