Methods and systems for generating a training dataset to train a machine learning model and for inferring virtual keypoint locations

By integrating projected-marker-based and in-the-wild datasets, the method enhances the accuracy and robustness of markerless motion capture systems in diverse environments and clothing, addressing the limitations of both marker-based and markerless systems.

WO2025244579A1PCT designated stage Publication Date: 2025-11-27NANYANG TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/SG2025/050320
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-20
Filing Date
2025-05-13
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing marker-based motion capture systems are labor-intensive, prone to errors due to marker placement variability and soft tissue artifacts, and require extensive post-processing, while markerless systems face challenges with inconsistent annotation and limited generalization to diverse environments and clothing, leading to inaccurate keypoint detection.

Method used

A method and system that combines projected-marker-based annotated images with in-the-wild datasets to train a machine learning model, using a neural network to infer virtual keypoint locations, and triangulate 3D positions from multiple camera views, minimizing errors and enhancing robustness across various environments and clothing.

Benefits of technology

The approach provides accurate and reliable virtual keypoint localization, reducing the need for manual annotation and marker placement, and improving the generalization of markerless motion capture systems to unseen environments and clothing conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2025050320_27112025_PF_FP_ABST
    Figure SG2025050320_27112025_PF_FP_ABST
Patent Text Reader

Abstract

According to embodiments of the present invention, a method for generating a training dataset to train a machine learning model for inferring virtual keypoint locations is provided. The method includes obtaining at least one augmented dataset including projected-marker-based annotated images of subjects; and sampling from the augmented dataset and at least one in-the-wild dataset based on a sampling ratio to generate the training dataset. The in-the-wild dataset includes manually annotated images of random subjects under unrestricted conditions. The projected-marker-based annotated images include 2D marker-based keypoint locations, augmented keypoints and projected-marker-based bounding boxes. The manually annotated images include 2D manually-annotated keypoint locations and manually-annotated bounding boxes. One or more of the augmented keypoints respectively coincide with one or more of the 2D manually-annotated keypoint locations. According to further embodiments, a system for generating a training dataset, and a method and system for inferring virtual keypoint locations are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR GENERATING A TRAINING DATASET TO TRAIN A MACHINE LEARNING MODEL AND FOR INFERRING VIRTUAL KEYPOINT LOCATIONSCross-Reference To Related Application

[0001] This application claims the benefit of priority of Singapore patent application No. 10202401429Q, filed 20 May 2024, the content of it being hereby incorporated by reference in its entirety for all purposes.Technical Field

[0002] Various embodiments relate to a method and a system for generating a training dataset to train a machine learning model for inferring virtual keypoint locations, as well as a method and a system for inferring virtual keypoint locations. Various embodiments also relate to a method for training a machine learning model for inferring virtual keypoint locations.Background

[0003] Human motion capture (mocap) technology is a valuable tool for biomechanics research as it provides objective data and insights into human movement, thereby facilitating precise analyses thereof. For instance, accurate and reliable mocap data is important for sports analysis to improve athletic performance, as well as clinical assessment to evaluate movement disorders.

[0004] Traditionally, if special keypoints are needed for a wide range of clothing and environments, a lot of manpower would be required to look at a large number of human- in-the wild images to click the 2D locations of those landmarks. This process is highly labor intensive, and it is prone to spatial error. Among the publicly released image databases with such annotation, none of the surface anatomical markers are selected as a keypoint in those datasets. Therefore, it is impossible for a data-driven method to learnthose surface anatomical keypoints from any human-in-the-wild dataset. Alternatively, those surface anatomical keypoint locations may be collected using a marker-based motion capture system for a high-accuracy annotation.

[0005] Marker-based (optical) mocap systems are generally considered the state-of-the-art equipment or the “gold standard” for assessing human movement. Marker-based mocap systems rely on anatomical landmarks such as bony landmarks to estimate the position and orientation of the segment. These systems require retro-reflective markers strategically placed on the subject’s body to be tracked by infrared cameras equipped with active infrared light sources. The precise triangulation and tracking of the 3D positions of these markers are facilitated by multiple synchronized and calibrated cameras. For example, a basic marker set uses markers positioned on the skin near joint centers to define limb segment position and orientation.

[0006] For a more comprehensive marker set, the Vicon’s Plug-in Gait uses around 38 markers for full-body modeling. The Rehabilitation Research Institute of Singapore (RRIS)’s Asian-centric movement database uses 84 markers based on a modified Calibrated Anatomical System Technique (CAST). Marker sets based on anatomical landmarks are advantageous due to their ease of placement and interpretation. However, their accuracy and reliability may be affected by marker placement variability and soft tissue artifacts. Between different skilled examiners, the marker placement root mean square (RMS) errors can be up to 24.8 mm. Soft-tissue artifacts occur because the skin and attached markers move relative to the bone during motion due to the deformation of tissues such as muscle or fat beneath the skin. This leads to unavoidable measurement errors. Although intracortical bone pins or biplanar videoradiography may provide more accurate motion data, the former is an invasive technique, while the latter involves exposure to radiation, making it impractical to collect large-scale data. The use of anatomical landmarks from marker-based mocap may be preferred as the errors from marker placement variability and soft tissue artifacts are relatively minor and more manageable when compared to the potential human errors involved in visually annotating joint centers or anatomical landmarks from 2D images.

[0007] Marker-based mocap systems have several limitations due to the need for markers. For example, the marker placement process is time-consuming and requires trainedpersonnel to palpate specific anatomical landmarks on the subject’s body. Further, the markers can affect the natural movement of the subject during data collection. Extensive postprocessing is also required such as marker labelling, gap filling due to marker occlusion, and correcting mislabeled markers due to marker swap. The time-consuming marker placement process and extensive postprocessing limit wider adoptions of the marker-based mocap systems.

[0008] A promising alternative to marker-based mocap is markerless mocap, driven by recent advancements in deep learning models capable of directly identifying virtual markers from a RGB image. In other words, one common step that runs in a data-driven markerless human motion capture system is the extraction of 2D keypoint locations from an image, and this step usually relies on a trained machine learning model such as artificial neural networks. This eliminates the need for real markers and overcomes various limitations associated with marker-based mocap by significantly reducing the time taken for subject preparation, data collection, and post-processing. Additionally, deep (or machine) learning models may learn a geometry-aware representation of the body and approximately infer the location of virtual markers in occluded regions based on other visible features. This helps to reduce occlusion issues and minimize gaps in marker trajectories.

[0009] Markerless mocap from multi-view videos generally involves detecting key features from the images and fitting an articulated human model to the features in 3D space. Early works on extracting image features relied on silhouette and edge-based techniques to create a 3D visual hull for tracking a single subject, but these techniques require good contrast between the captured subject and the background. Furthermore, it is difficult to track segments that are nearly rotationally symmetric around their local axes, as the silhouette of these segments barely changes when they rotate around these axes. Significant progress has been made in computer vision and deep learning. Convolutional neural networks (CNNs) have played a key role in frameworks like OpenPose and DensePose, enabling the detection of body joints and estimation of dense correspondence maps. Many methods rely on open-source frameworks like OpenPose, which can only detect around 20 body keypoints, but these joint centers are insufficient to determine the orientation of body segments. Hence, OpenCap uses two Long Short-Term Memory (LSTM) networks trainedon a large mocap dataset to augment the 3D positions of the sparse 20 keypoints with a more detailed 43 surface anatomical markers. Because those augmented anatomical positions are not extracted directly from the images, this process may introduce a bias towards average movement patterns and may not generalize to individuals with unique movement patterns such as rehabilitation patients.

[0010] While markerless mocap has rapidly advanced with a model-centric focus on developing new architecture and training methods, a data-centric approach is also important as the performance of deep learning models depends on the quality and quantity of the training dataset. The quality of the inference of such a data-driven model depends largely on the training process and the selection of training data. Some popular datasets for 2D human keypoint estimation include the MPII Human Pose with 25K images, and the COCO or COCO-WholeBody with 200K images. Although these datasets contain diverse human-in-the-wild images, the keypoints are manually annotated through crowdsourcing. This results in inconsistent annotation as different people may have different interpretations of a joint center and potential mislabeling can affect estimation accuracy, especially for the hip and knee joint centers. Furthermore, these datasets are annotated with around 16 to 23 body keypoints, providing only two joint centers for each lower and upper limb segment. Thus, it is difficult to evaluate rotation about the longitudinal axis of a bone, as a minimum of three non-aligned markers are required to fully reconstruct a rigid segment’s six-degree- of-freedom transformation.

[0011] To address the issues of inconsistent annotation and insufficient 2D keypoints, markerless mocap systems such as Theia3D use highly trained annotators with anatomical knowledge to manually annotate 51 salient features, including joint locations and other identifiable surface features, on images of over 500K humans in the wild. Although quality control is done by additional expert annotators, it may be challenging for humans to accurately annotate occluded points. PitchAI™ is another example of a single-camera markerless mocap system for baseball pitching that estimates 53 3D markers (19 joint centers, 34 bony landmarks) from 19 2D joint centers.

[0012] As mentioned above, datasets are important for training deep learning models and evaluating the performance of various methods. Most datasets for markerless human mocap contain diverse images of humans in the wild engaged in different activities and scenes.While these datasets help in the training of robust models that may be applied to real-world scenarios, they may not be suitable for biomechanics applications. This is because manually annotated keypoints may be inaccurate and the sparse body keypoints are insufficient to determine the orientation of body segments. rooi3i For example, the MoVi dataset automates the annotation process by collecting synchronized and calibrated marker-based mocap data with stationary video cameras to allow the overlay of the 3D skeletal pose in camera coordinates. However, for some of the images, the inconsistent quality of the given camera calibration parameters and time synchronization results in significant errors during projection.

[0014] Marker trajectory data from the marker-based mocap may also be combined with simultaneous video record to provide “human-in-mocap-lab” dataset for the purpose of training a keypoint detection model. A downside of using this kind of human-in-mocap- lab dataset is the limitation of the capture environment and clothing. This is because data collection is usually done in just one location, and many types of clothes are restricted as they block the markers or prevent marker placement at the right bone landmark. For example, shoes, long pants, and baggy clothes are not allowed. The low variety of clothes and environment in the dataset simply leads to the low transferability of the model to perform inferences in an unseen location or clothing. rooisi Thus, there is a need for an improved method and / or system to address at least the abovementioned problems by outputting the position of surface anatomical markers commonly used by biomechanical analysis workflow, where marker trajectories allow established biomechanical workflow / formulations to be applied directly with minimal to zero adjustments on the virtual marker data. Further, the markerless mocap user can enjoy the benefits of a markerless motion capture system but use the output from the method / system as if it is a marker-based motion capture system.Summary

[0016] According to an embodiment, a method for generating a training dataset to train a machine learning model for inferring virtual keypoint locations is provided. The method includes obtaining at least one augmented dataset comprising projected-marker-basedannotated images of subjects; and sampling from the at least one augmented dataset and at least one in-the-wild dataset based on a sampling ratio to generate the training dataset. The at least one in-the-wild dataset includes manually annotated images of random subjects under unrestricted conditions. The projected-marker-based annotated images of the subjects include 2D marker-based keypoint locations, augmented keypoints and projected- marker-based bounding boxes. The manually annotated images of the random subjects include 2D manually-annotated keypoint locations and manually-annotated bounding boxes. One or more of the augmented keypoints respectively coincide with one or more of the 2D manually-annotated keypoint locations.

[0017] According to an embodiment, a method for training a machine learning model for inferring virtual keypoint locations is provided. The method includes (i) feeding a training dataset generated by a method for generating a training dataset, according to an embodiment, into a neural network; (ii) generating, by the neural network, outputs for one or more target subjects; (iii) comparing the outputs and annotations in the training dataset to obtain differences thereof; (iv) adjusting weights and / or biases in the neural network based on the obtained differences; and (v) repeating (i) to (iv) until the outputs substantially match the annotations or for a predetermined number of iterations to obtain the trained machine learning model. The substantially matched outputs or the outputs after completing the predetermined number of iterations include a plurality of anatomical keypoint 2D location predictions, a plurality of wild keypoint 2D location predictions, and a bounding box prediction surrounding each target subject.

[0018] According to an embodiment, a method for inferring virtual keypoint locations is provided. The method includes based on a marker-less human or animal subject captured by a plurality of colour video cameras as sequences of 2D images, for each 2D image captured by each colour video camera, generating, using a trained machine learning model, a 2D bounding box; for each 2D image, generating, by the trained machine learning model, a plurality of heatmaps with scores of confidence, wherein each heatmap is for 2D localization of a keypoint of the marker-less human or animal subject; for each heatmap, selecting a pixel with the highest score of confidence, and associating the selected pixel to the keypoint, thereby determining the 2D location of the keypoint; and based on the sequences of 2D images captured by the plurality of colour video cameras, triangulatingthe respective determined 2D locations to predict a sequence of 3D locations of the keypoint, thereby inferring the virtual keypoint locations. The trained machine learning model is trained using at least a training dataset generated by a method for generating a training dataset, according to an embodiment. For each heatmap, the scores of confidence are indicative of probability of having the associated keypoint in different 2D locations in the generated 2D bounding box.

[0019] According to an embodiment, a system for generating a training dataset to train a machine learning model for inferring virtual keypoint locations is provided. The system includes a computer configured to obtain at least one augmented dataset including projected-marker-based annotated images of subjects; and sample from the at least one augmented dataset and at least one in-the-wild dataset based on a sampling ratio to generate the training dataset. The at least one in-the-wild dataset includes manually annotated images of random subjects under unrestricted conditions. The projected-marker-based annotated images of the subjects include 2D marker-based keypoint locations, augmented keypoints and projected-marker-based bounding boxes. The manually annotated images of the random subjects include 2D manually-annotated keypoint locations and manually- annotated bounding boxes. One or more of the augmented keypoints respectively coincide with one or more of the 2D manually-annotated keypoint locations.

[0020] According to an embodiment, a system for inferring virtual keypoint locations is provided. The system includes a plurality of colour video cameras configured to capture a marker-less human or animal subject as sequences of 2D images; and a computer configured to receive the sequences of 2D images captured by the plurality of colour video cameras; for each 2D image captured by each colour video camera, generate, using a trained machine learning model, a 2D bounding box; for each 2D image, generate, using the trained machine learning model, a plurality of heatmaps with scores of confidence, wherein each heatmap is for 2D localization of a keypoint of the marker-less human or animal subject; for each heatmap, select a pixel with the highest score of confidence, and associate the selected pixel to the keypoint to determine the 2D location of the keypoint; and based on the sequences of 2D images captured by the plurality of colour video cameras, triangulate the respective determined 2D locations to predict a sequence of 3D locations of the keypoint, thereby inferring the virtual keypoint locations. For each heatmap, the scoresof confidence are indicative of probability of having the associated keypoint in different 2D locations in the generated 2D bounding box. The trained machine learning model is trained using at least a training dataset generated by a method for generating a training dataset, according to an embodiment.Brief Description of the Drawings

[0021] In the drawings, like reference characters generally refer to like parts throughout the different views. The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, in which:

[0022] FIG. 1 shows a flow chart illustrating a method for generating a training dataset to train a machine learning model for inferring virtual keypoint locations, according to various embodiments.

[0023] FIG. 2 shows a flow chart illustrating a method for training a machine learning model for inferring virtual keypoint locations, according to various embodiments.

[0024] FIG. 3 shows a schematic cross-sectional view of a system for generating a training dataset to train a machine learning model for inferring virtual keypoint locations, according to various embodiments.

[0025] FIG. 4 shows a flow chart illustrating a method for inferring virtual keypoint locations, according to various embodiments.

[0026] FIG. 5 shows a schematic cross-sectional view of a system for inferring virtual keypoint locations, according to various embodiments.

[0027] FIG. 6 relates to prior art depicting a cropped image from the Human3.6M dataset with 3D model-based keypoints projected onto a 2D image.

[0028] FIG. 7 relates to prior art depicting cropped images from the MoVi dataset with 3D markers projected onto a 2D image plane.

[0029] FIG. 8 shows a flow chart depicting a simplified overview of the training data selection and training method used to obtain a final model that infers surface anatomicalbone landmarks from images of humans in random environments and clothes, according to various examples.

[0030] FIG. 9 shows a schematic view representing the data collection and training pipeline, according to one example.

[0031] FIG. 10 shows a top-view map of all the camera placements and the capture volume used in the RRIS40 test set, according to one example.

[0032] FIG. 11 shows a schematic view of marker placements, according to various examples.

[0033] FIG. 12 shows a cropped image with all the crosses display projected 2D marker positions, according to one example.

[0034] FIG. 13 shows cropped images illustrating an example of marker removal before (left) and after (right) using GAN-based context-aware image inpainting.

[0035] FIG. 14 shows a schematic view representing the 3D virtual marker tracking pipeline, according to one example

[0036] FIG. 15 shows the area under curve (AUC) and the error distribution for specific markers for benchmarking results of the proposed method of virtual marker trajectories with three different neural network models.

[0037] FIG. 16 shows a cumulative distribution plot illustrating overall accuracy profile of the proposed method with three different neural network models of FIG. 15 from all the 40 virtual markers.

[0038] FIG. 17 shows the AUC and the error distribution for specific markers for benchmarking results on the RRIS40 test set in evaluation against an anatomical-marker- based motion capture system.

[0039] FIG. 18 shows a cumulative distribution plot depicting overall accuracy profile of the proposed method and OpenCap from all the 23 overlapping virtual anatomical markers with a total of 2,337,253 Euclidean error samples from 10 test subjects from the RRIS40 test set.

[0040] FIG. 19 shows the AUC and the error distribution for specific markers for benchmarking results on the RRIS40 test set in evaluation against joint-center-base pose estimation tools.

[0041] FIG. 20 shows a cumulative distribution plot depicting overall accuracy profile of the proposed method, Detectron2, OpenCap, and MediaPipe from all the 12 key joint centers with a total of 1,169,671 Euclidean error samples from 10 test subjects from the RRIS40 test set.

[0042] FIG. 21 shows the AUC and the error distribution for specific markers for benchmarking results on the GPJATK dataset on 15 overlapping markers.

[0043] FIG. 22 shows the AUC and the error distribution for specific markers for benchmarking results of two models on the RRIS40 test set.

[0044] FIG. 23 shows snapshots of video records with 2D projections of 3D virtual markers while the subject is walking with different walking aids, according to various examples.

[0045] FIG. 24 shows images from the COCO dataset, with overlying 2D keypoints inferred by the trained machine learning model, according to various examples.

[0046] FIG. 25 shows a plot depicting the relationship between mean Euclidean errors and the number of cameras, thereby illustrating the average error from each camera combination, according to one example.

[0047] FIG. 26 shows a plot illustrating the relationship between Euclidean errors and the image resolution for the comparison of overall Euclidean errors with varying image resolutions.Detailed Description

[0048] The following detailed description refers to the accompanying drawings that show, by way of illustration, specific details and embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention. Other embodiments may be utilized and structural, logical, and electrical changes may be made without departing from the scope of the invention. The various embodiments are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments.

[0049] Embodiments described in the context of one of the methods or devices are analogously valid for the other methods or devices. Similarly, embodiments described in the context of a method are analogously valid for a device, and vice versa.

[0050] Features that are described in the context of an embodiment may correspondingly be applicable to the same or similar features in the other embodiments. Features that are described in the context of an embodiment may correspondingly be applicable to the other embodiments, even if not explicitly described in these other embodiments. Furthermore, additions and / or combinations and / or alternatives as described for a feature in the context of an embodiment may correspondingly be applicable to the same or similar feature in the other embodiments.

[0051] In the context of various embodiments, the articles “a”, “an” and “the” as used with regard to a feature or element include a reference to one or more of the features or elements.

[0052] In the context of various embodiments, the phrase “substantially” or “at least substantially” may include “exactly” and a reasonable variance.

[0053] In the context of various embodiments, the term “about” or “approximately” as applied to a numeric value encompasses the exact value and a reasonable variance.

[0054] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.

[0055] As used herein, the phrase of the form of “at least one of A or B” may include A or B or both A and B. Correspondingly, the phrase of the form of “at least one of A or B or C”, or including further listed items, may include any and all combinations of one or more of the associated listed items.

[0056] As used herein, the expression “configured to” may mean “constructed to” or “arranged to”.

[0057] Various embodiments relate to anatomical-marker-driven 3D markerless human motion capture. More specifically, a model training method to enhance robustness of virtual surface marker localization in unseen environment for a markerless motion capture system may be provided.

[0058] FIG. 1 shows a flow chart illustrating a method 100 for generating a training dataset to train a machine learning model for inferring virtual keypoint locations, according to various embodiments. As seen in FIG. 1, at Step 102, at least one augmented dataset including projected-marker-based annotated images of subjects is obtained. At Step 104, sampling from the at least one augmented dataset and at least one in-the-wild dataset based on a sampling ratio to generate the training dataset. The at least one in-the-wild datasetincludes manually annotated images of random subjects under unrestricted conditions. The projected-marker-based annotated images of the subjects include 2D marker-based keypoint locations, augmented keypoints and projected-marker-based bounding boxes. The manually annotated images of the random subjects include 2D manually-annotated keypoint locations and manually-annotated bounding boxes. One or more of the augmented keypoints respectively coincide with one or more of the 2D manually-annotated keypoint locations.

[0059] In the context of various embodiments, “virtual keypoint location” refers to a predicted 3D location of an important feature of a human or animal subject.

[0060] “Projected-marker-based annotated images of subjects” refer to images of subjects that are annotated using marker-based mocap.

[0061] Each “projected-marker-based bounding box” refer to a polygon (e.g. a rectangle) denoting a region of interest (e.g. each subject) in the projected-marker-based annotated images. The source of ground truth bounding box for the augmented dataset is calculated from 3D marker positions and project to the projected-marker-based annotated image as a 2D bounding box (i.e. the “projected-marker-based bounding box”).

[0062] Each “manually-annotated bounding box” refers to a polygon (e.g. a rectangle) drawn around a region of interest (e.g. each random subject) in the manually annotated images by someone.

[0063] The in-the-wild dataset may be obtained by capturing images of the random subjects, manually identifying a plurality of wild keypoints in each image with each manually-annotated bounding box surrounding each random subject, and labeling the plurality of wild keypoints in each image to generate the 2D manually- annotated keypoint locations in each manually annotated image. For example, the in-the-wild dataset may include COCO dataset, COCO-Wholebody dataset, MPTT dataset, AT challenger dataset, Leeds Sports Pose dataset, or Relative Human dataset.

[0064] The unrestricted conditions may include random environments and the random subjects with random clothes.

[0065] In various embodiments, obtaining the at least one augmented dataset at Step 102 may include based on a plurality of physical markers captured by an optical marker-based motion capture system, each as a 3D trajectory, wherein each physical marker is placed ona bone landmark or a keypoint of each subject, and the subjects substantially simultaneously captured by a plurality of colour video cameras over a period of time as sequences of 2D images, for each physical marker, identifying the captured 3D trajectory with a marker label representative of the bone landmark or keypoint on which the physical marker is placed; obtaining augmented markers; determining an augmented 3D trajectory with an augmented label representative of each augmented marker; for each physical marker, projecting the 3D trajectory and for each augmented marker, projecting the augmented 3D trajectory to each of the 2D images to determine a 2D location in each 2D image; for each physical marker and for each augmented marker, based on the respective 2D locations in the sequences of 2D images and an exposure-related time of the plurality of colour video cameras, interpolating a 3D position for each of the 2D images; for each 2D image, based on the respective interpolated 3D positions of the plurality of physical markers and the augmented markers, and an extended volume, generating each projected- marker-based bounding box around each subject; and generating the augmented dataset including at least one 2D image selected from the sequences of 2D images, the determined 2D location of each physical marker in the selected at least one 2D image, the determined 2D location of each augmented marker in the selected at least one 2D image, and the projected-marker-based bounding boxes for the selected at least one 2D image. r0066] The exposure-related time may involve a middle of exposure time to capture each 2D image or a part thereof using each colour video camera.

[0067] Each augmented marker may be calculated from positions of two or more of the physical markers placed on each subject. The extended volume may be derived from two or more of the physical markers and / or the augmented markers having an anatomical, functional and / or structural relationship with one another.

[0068] For each physical marker, the marker label may be arranged to be propagated with each determined 2D location such that in the augmented dataset, each determined 2D location of each physical marker contains the corresponding marker label. For each augmented marker, the augmented label may be arranged to be propagated with each determined 2D location such that in the augmented dataset, each determined 2D location of each augmented marker contains the corresponding augmented label.

[0069] The 2D marker-based keypoint locations may include the determined 2D locations of the plurality of physical markers.

[0070] The augmented keypoints may include the determined 2D locations of the augmented markers.

[0071] In various embodiments, the augmented markers may include each joint center of the positions of the two or more of the physical markers placed on each subject. Additionally or alternatively, the augmented markers may include each 2D wild keypoint location calculated from the positions of the two or more of the physical markers placed on each subject, wherein the 2D wild keypoint location corresponds to one of the 2D manually-annotated keypoint locations in the at least one in-the-wild dataset. For example, a calculated left shoulder joint center (i.e. an augmented marker derived by calculations based on the positions of the two or more of the physical markers e.g. that may be associated to the left shoulder of the subject) would exist in the at least one in-the-wild dataset. With such overlapping, the trained machine learning model gains benefit of merging the datasets with non-overlapping keypoint set.

[0072] The plurality of physical markers each being captured as the 3D trajectory and the subjects being substantially simultaneously captured as the sequences of 2D images over the period of time may be coordinated using a synchronized signal communicated by the optical marker-based motion capture system to the plurality of colour video cameras.

[0073] In various embodiments, the method 100 may further include after projecting the 3D trajectory to each of the 2D images to determine the 2D location in each 2D image, in each 2D image and for each physical marker, drawing a 2D radius on the determined 2D location according to a distance with a predefined margin between the colour video camera(s) and the physical marker to form an encircled area, and applying a learning-based context-aware image inpainting technique to the encircled area to remove a marker blob from the 2D location.

[0074] In other embodiments, the method 100 may further include after projecting the augmented 3D trajectory to each of the 2D images to determine the 2D location in each 2D image, in each 2D image and for each augmented marker, drawing a 2D radius on the determined 2D location according to a distance with a predefined margin between the colour video camera(s) and the augmented marker to form an encircled area, and applyinga learning-based context-aware image inpainting technique to the encircled area to remove a marker blob from the 2D location.

[0075] For example, the learning-based context-aware image inpainting technique may include a Generative Adversarial Network-based context-aware image inpainting technique.

[0076] In different embodiments, the method 100 may further include detecting, using a human detection algorithm, non-subject humans captured in the 2D images; and blurring the non-subject humans in the 2D images to remove the non-subject humans.

[0077] At Step 104, sampling from the augmented dataset and the in-the-wild dataset may be based on the sampling ratio of the augmented dataset and the in-the-wild dataset at A:W, where each of A and W is greater than 0. For example, sampling from the augmented dataset and the in-the-wild dataset may be based on the sampling ratio of the augmented dataset and the in-the-wild dataset at 80:20.

[0078] FIG. 2 shows a flow chart illustrating a method 220 for training a machine learning model for inferring virtual keypoint locations, according to various embodiments. As seen in FIG. 2, at Step 222, a training dataset generated by a method 100 (FIG. 1) is fed into a neural network. At Step 224, outputs for one or more target subjects are generated by the neural network. At Step 226, the outputs and annotations in the training dataset are compared to obtain differences thereof. At Step 228, weights and / or biases are adjusted in the neural network based on the obtained differences. At Step 230, Step 222 to Step 228 are repeated until the outputs substantially match the annotations or for a predetermined number of iterations to obtain the trained machine learning model (Step 232). The substantially matched outputs or the outputs after completing the predetermined number of iterations, from the trained machine learning model, include a plurality of anatomical keypoint 2D location predictions, a plurality of wild keypoint 2D location predictions, and a bounding box prediction surrounding each target subject are generated by the neural network.

[0079] In the context of various embodiments, “anatomical keypoint 2D location predictions” refer to 2D locations that are predicted during training based on the at least one augmented dataset. “Wild keypoint 2D location predictions” refer to 2D locations that are predicted during training based on the at least one in-the-wild dataset.

[0080] For Step 222, the annotations may include the projected marker-based bounding boxes and / or the manually-annotated bounding boxes that are the ground truth of the training data used during the training step for the neural network to try its best to find a calculation from the input image to predict an output that minimizes the differences towards the ground truth. In other words, when the outputs substantially match the annotations, the differences are minimized towards the ground truth. In an alternative approach, the differences may be minimized towards the ground truth after performing the training for the predetermined number of iterations. The predetermined number of iterations may be determined or estimated from prior repetitive experiments. For example, in prior repetitive experiments, it may be observed that upon the weight / bias adjustment to some number of iterations, the output may be sufficiently close the ground truth after a certain number of iterations. This certain number of iterations may then be the predetermined number of iterations used as a criterion in the training process.

[0081] In one embodiment, the method 220 may further include discarding the plurality of wild keypoint 2D location predictions from the trained machine learning model. This may be carried out prior to using the trained machine learning model to infer virtual keypoint locations.

[0082] The method 220 of FIG. 2 may encompass the same or similar elements of the method 100 of FIG. 1, and as such, the descriptions of these elements in relation to the method 220 of FIG. 2 are omitted here.

[0083] With reference to the method 100 of FIG. 1 and the method 220 of FIG. 2, a way of machine learning model training method that utilizes both kinds of datasets (i.e. human- in-the-wild dataset without surface anatomical marker annotation (e.g. the in-the-wild dataset), and human-in-mocap-lap dataset with surface anatomical marker annotation (e.g. the augmented dataset)) may be provided, and a model that localizes surface anatomical markers (i.e. markers that are attached on the bone landmarks in the traditional markerbased motion capture workflow) for unseen clothing and environments may be produced without having to annotate those surface keypoint locations on any images of humans in the wild.

[0084] FIG. 4 shows a flow chart illustrating a method 440 for inferring virtual keypoint locations, according to various embodiments. As seen in FIG. 4, based on a marker-lesshuman or animal subject captured by a plurality of colour video cameras as sequences of 2D images, at Step 442, for each 2D image captured by each colour video camera, a 2D bounding box is generated using a trained machine learning model. At Step 444, for each 2D image, a plurality of heatmaps with scores of confidence is generated by the trained machine learning model. Each heatmap is for 2D localization of a keypoint of the markerless human or animal subject. The trained machine learning model is trained using at least a training dataset generated by a method 100 of FIG. 1. The trained machine learning model may be trained by a method 220 of FIG. 2. At Step 446, for each heatmap, a pixel with the highest score of confidence is selected, and the selected pixel is associated to the keypoint, thereby determining the 2D location of the keypoint. For each heatmap, the scores of confidence are indicative of probability of having the associated keypoint in different 2D locations in the generated 2D bounding box. At Step 448, based on the sequences of 2D images captured by the plurality of colour video cameras, the respective determined 2D locations are triangulated to predict a sequence of 3D locations of the keypoint, thereby inferring the virtual keypoint locations.

[0085] In other words, the trained machine learning model is used to infer 2D keypoint locations from calibrated cameras (e.g. 8 of them) that simultaneously record a human with frame-level synchronization. Then, each keypoint is triangulated to reconstruct the 3D location of virtual markers.

[0086] In various embodiments, triangulating the respective determined 2D locations at Step 448 may include performing strategic triangulation of one trajectory of the keypoint at a time.

[0087] The strategic triangulation may include: for each 2D image, performing weighted triangulation for 2N-(N+1) combinations of the plurality of colour video cameras to obtain 2N-(N+1 ) candidates of the 3D locations together with respective maxRay Distance, N being the total number of colour video cameras, and for each candidate, the maxRayDistance being a perpendicular distance from each 3D location to a farthest ray used in a corresponding combination; among the 2N-(N+1) candidates in each 2D image, determining whether each candidate is noisy data based on that candidate having less rays with a smaller maxRayDistance than at least one other candidate; eliminating that candidate, if determined to be the noisy data, to obtain remaining candidates; and from theremaining candidates, selecting one candidate in each time frame such that a sequence of the selected candidates forms a 3D trajectory of the keypoint with a lowest cost.

[0088] The 2N-(N+1) combinations exclude a combination with zero colour video camera and N combinations with one colour video camera.

[0089] Each time frame is each moment when N colour video cameras capture N images simultaneously.

[0090] The lowest cost is obtained by applying a cost function given by:where my is a distance multiplier calculated by a monotonically decreasing function of the minimum number of colour video cameras involved in the triangulation of the selected candidates of that keypoint between two consecutive time frames, df is a distance between the selected candidates between the two consecutive time frames, and / = 1, 2, ... , total number of time frames- 1.

[0091] The weighted triangulation may include derivation of each predicted 3D location of the keypoint using a formula: wi is a weight for triangulation or the score of confidence of ray from * colour videocamera, is a 3D location of thecolour video camera associated with the ray,is a 3D unit vector representing a back-projected direction associated with the zthray,is a 3x3 identity matrix.

[0092] The strategic triangulation may be illustrated by the following example.

[0093] Step 1 : By focusing on just one marker, in one specific time frame, eight cameras localize this marker in 2D with the trained machine learning model. However, weighted triangulation does not need to select all eight cameras to contribute to this triangulation - sometimes it may be more accurate to exclude some cameras from the triangulation. Foreight cameras, there are in total 28or 256 combinations to select from. There is one combination with zero camera, and eight combinations with one camera that cannot perform triangulation. For the rest of 256-9 = 247 combinations, all the triangulations are calculated to obtain all the 247 candidates of 3D positions together with maxRay Distance. F0094] Step 2 : using the condition that eliminates some candidates that are guaranteed as not-the-best option, generally reduces the candidates significantly. After applying this condition, maxRayDistance is no longer used.

[0095] Step 3: the whole timeline needs to be considered simultaneously (but one marker at a time). Imagine each time frame has a number of survived candidates floating as a point cloud cluster in 3D. This step is trying to select just one candidate per time frame and link them all together so that the sum of squared jumping distance between consecutive frames are as small as possible. The link from the candidate selection in the first time frame to the candidate selection in the last time frame is the selected trajectory for this marker.

[0096] In various embodiments, each colour video camera may include at least one visible light emitting diodes operable to facilitate retro -reflective markers to be perceived as detectable bright spots. The method 440 may further include extrinsically calibrating the plurality of colour video cameras by based on the retro-reflective markers captured by the plurality of colour video cameras as sequences of 2D calibration images, applying an optimization function to the captured 2D calibration images to fine-tune extrinsic camera parameters of the plurality of colour video cameras.

[0097] While each of the methods described above is illustrated and described as a series of steps or events, it will be appreciated that any ordering of such steps or events are not to be interpreted in a limiting sense. For example, some steps may occur in different orders and / or concurrently with other steps or events apart from those illustrated and / or described herein. Tn addition, not all illustrated steps may be required to implement one or more aspects or embodiments described herein. Also, one or more of the steps depicted herein may be carried out in one or more separate acts and / or phases.

[0098] Various embodiments provide a computer program adapted to perform a method 100 of FIG. 1, a method 220 of FIG. 2 and / or a method 440 of FIG. 4; a non-transitory computer readable medium including instructions which, when executed on a computer, cause the computer to perform a method 100 of FIG. 1 , a method 220 of FIG. 2 and / or amethod 440 of FIG. 4; or a data processing apparatus including means for carrying out a method 100 of FIG. 1, a method 220 of FIG. 2 and / or a method 440 of FIG. 4.

[0099] FIG. 3 shows a system 300 for generating a training dataset to train a machine learning model for inferring virtual keypoint locations, according to various embodiments. The system 300 includes a computer 301 configured to: obtain at least one augmented dataset including projected-marker-based annotated images of subjects; and sample from the at least one augmented dataset and at least one in-the-wild dataset 311 based on a sampling ratio to generate the training dataset. The at least one in-the-wild dataset 311 includes manually annotated images of random subjects under unrestricted conditions. The at least one in-the-wild dataset 311 may be received and / or stored (as denoted by line 313) by the computer 301 . The projected-marker-based annotated images of the subjects include 2D marker-based keypoint locations, augmented keypoints and projected-marker-based bounding boxes. The manually annotated images of the random subjects include 2D manually-annotated keypoint locations and manually-annotated bounding boxes. One or more of the augmented keypoints respectively coincide with one or more of the 2D manually-annotated keypoint locations.

[0100] The system 300 may be configured to perform the method 100 of FIG. 1. The system 300 may include the same or like elements or components as those of the method 100 of FIG. 1, and as such, the like elements may be as described in the same or similar context of the method 100 of FIG. 1, and therefore some corresponding descriptions may be omitted here.

[0101] The system 300 may further include an optical marker-based motion capture system 303 configured to capture a plurality of physical markers over a period of time; and a plurality of colour video cameras 307 configured to capture the subjects over the period of time as sequences of 2D images. Each physical marker may be placed on a bone landmark or a keypoint of each subject, and may be captured as a 3D trajectory.

[0102] The computer 301 may be in communication with the optical marker-based motion capture system 303 (as denoted by line 305) and with the plurality of colour video cameras 307 (as denoted by line 309).

[0103] To obtain the at least one augmented dataset, the computer 301 may be configured to: receive the sequences of 2D images captured by the plurality of colour video cameras307 and the respective 3D trajectories captured by the optical marker-based motion capture system 303; for each physical marker, identify the captured 3D trajectory with a marker label representative of the bone landmark or keypoint on which the physical marker is placed, obtain augmented markers; determine an augmented 3D trajectory with an augmented label representative of each augmented marker; for each physical marker, project the 3D trajectory and for each augmented marker, project the augmented 3D trajectory to each of the 2D images to determine a 2D location in each 2D image; for each physical marker and for each augmented marker, based on the respective 2D locations in the sequences of 2D images and an exposure -related time of the plurality of colour video cameras 307, interpolate a 3D position for each of the 2D images; for each 2D image, based on the respective interpolated 3D positions of the plurality of physical markers and the augmented markers, and an extended volume derived from two or more of the physical markers and / or the augmented markers having an anatomical, functional and / or structural relationship with one another, generate each projected-marker-based bounding box around each subject; and generate the augmented dataset including at least one 2D image selected from the sequences of 2D images, the determined 2D location of each physical marker in the selected at least one 2D image, the determined 2D location of each augmented marker in the selected at least one 2D image, and the projected-marker-based bounding boxes for the selected at least one 2D image.

[0104] To obtain the augmented markers, the computer 301 may be configured to calculate each joint center of the positions of the two or more of the physical markers placed on each subject and / or calculate each 2D wild keypoint location from the positions of the two or more of the physical markers placed on each subject, wherein the 2D wild keypoint location corresponds to one of the 2D manually-annotated keypoint locations in the at least one in- the-wild dataset.

[0105] In other words, the computer 301 may be configured to perform Step 102 (FIG. 1).

[0106] The system 300 may further include a synchronization pulse generator (not shown in FIG. 3) in communication with the optical marker-based motion capture system 303 and the plurality of colour video cameras 307. The synchronization pulse generator may be configured to receive a synchronization signal from the optical marker -based motioncapture system 303 for coordinating the subjects to be substantially simultaneously captured by the plurality of colour video cameras 307.

[0107] In various embodiments, the computer 301 may be further configured to, in each 2D image, draw a 2D radius on the determined 2D location for each physical marker according to a distance with a predefined margin between the colour video camera(s) 307 and the physical marker to form an encircled area, and to apply a learning-based context- aware image inpainting technique to the encircled area to remove a marker blob from the 2D location.

[0108] In other embodiments, the computer 301 may be further configured to, in each 2D image, draw a 2D radius on the determined 2D location for each augmented marker according to a distance with a predefined margin between the colour video camera(s) 307 and the augmented marker to form an encircled area, and to apply a learning-based context- aware image inpainting technique to the encircled area to remove a marker blob from the 2D location.

[0109] For example, the learning-based context-aware image inpainting technique may include a Generative Adversarial Network-based context-aware image inpainting technique.

[0110] In different embodiments, the computer 301 may be further configured to execute a human detection algorithm to detect non-subject humans captured in the 2D images, and to blur the non-subject humans in the 2D images to remove the non-subject humans.

[0111] The computer 301 may be configured to sample from the augmented dataset and the in-the-wild dataset based on the sampling ratio of the augmented dataset and the in-the- wild dataset at A:W, where each of A and W is greater than 0. For example, the computer 301 may be configured to sample from the augmented dataset and the in-the-wild dataset based on the sampling ratio of the augmented dataset and the in-the-wild dataset at 80:20.

[0112] FIG. 5 shows a system 540 for for inferring virtual keypoint locations, according to various embodiments. The system 540 includes a plurality of colour video cameras 543 configured to capture a marker-less human or animal subject as sequences of 2D images; and a computer 541 in communication with the plurality of colour video cameras 543 (as denoted by line 545).

[0113] The system 540 may be configured to perform the method 440 of FIG. 4. The system 540 may include the same or like elements or components as those of the method 440 of FIG. 4, and as such, the like elements may be as described in the same or similar context of the method 440 of FIG. 4, and therefore some corresponding descriptions may be omitted here.

[0114] Further, the computer 541 and the plurality of colour video cameras 543 may be described in the same or similar context to the computer 301 and the plurality of colour video cameras 307 of FIG. 3.

[0115] The computer 541 is configured to receive the sequences of 2D images captured by the plurality of colour video cameras 543; for each 2D image captured by each colour video camera 543, generate, using a trained machine learning model, a 2D bounding box; for each 2D image, generate, using the trained machine learning model, a plurality of heatmaps with scores of confidence; for each heatmap, select a pixel with the highest score of confidence, and associate the selected pixel to the keypoint to determine the 2D location of the keypoint; and based on the sequences of 2D images captured by the plurality of colour video cameras 543, triangulate the respective determined 2D locations to predict a sequence of 3D locations of the keypoint, thereby inferring the virtual keypoint locations.

[0116] Each heatmap is for 2D localization of a keypoint of the marker-less human or animal subject, and the trained machine learning model is trained using at least a training dataset generated by a method 100 of FIG. 1. For each heatmap, the scores of confidence are indicative of probability of having the associated keypoint in different 2D locations in the generated 2D bounding box.

[0117] To triangulate the respective determined 2D locations, the computer 541 may be configured to perform strategic triangulation of one trajectory of the keypoint at a time.

[0118] To perform the strategic triangulation, the computer 541 may be configured to: for each 2D image, perform weighted triangulation for 2N-(N+1) combinations of the plurality of colour video cameras to obtain 2N-(N+1) candidates of the 3D locations together with respective maxRayDistance, N being the total number of colour video cameras, and for each candidate, the maxRayDistance being a perpendicular distance from each 3D location to a farthest ray used in a corresponding combination, wherein the 2N-(N+1) combinations exclude a combination with zero colour video camera and N combinations with one colourvideo camera; among the 2N-(N+1) candidates in each 2D image, determine whether each candidate is noisy data based on that candidate having less rays with a smaller maxRayDistance than at least one other candidate; eliminate that candidate, if determined to be the noisy data, to obtain remaining candidates; from the remaining candidates, select one candidate in each time frame such that a sequence of the selected candidates forms a 3D trajectory of the keypoint with a lowest cost.

[0119] In various embodiments, each colour video camera 543 may include at least one visible light emitting diodes operable to facilitate retro-reflective markers to be perceived as detectable bright spots. The plurality of colour video cameras 543 may be configured to capture the retro-reflective markers as sequences of 2D calibration images. The plurality of colour video cameras 543 may be extrinsically calibrated by applying an optimization function to the captured 2D calibration images to fine-tune extrinsic camera parameters of the plurality of colour video cameras 543.

[0120] Examples of the methods 100 (FIG. 1), 220 (FIG. 2), 440 (FIG. 4) and the systems 300 (FIG. 3), 540 (FIG. 5) will be described in more details below.

[0121] Anatomical priors and data-driven methods are integrated to capture human movement from multi-view RGB video sequences. First, a high-quality dataset annotated with surface anatomical bone landmarks (e.g. 40 surface anatomical bone landmarks) is created to train a 2D keypoint detection model. This dataset (referred to as RRIS40) uses precise marker positions from marker-based mocap to generate pixel-accurate 2D anatomical landmarks. This allows for the data collection to be scaled to millions of images. Both the marker-based and markerless systems are carefully calibrated to ensure optimal spatial and temporal alignment. Next, markers are removed from images to prevent the keypoint detection model from learning anatomical landmarks from the markers’ distinct features. A weighted triangulation method that considers confidence scores is proposed to improve the accuracy of 3D markers estimation from multiple 2D anatomical landmarks. The system outputs the 3D positions of virtual markers, which may be used to derive joint centers. This makes it compatible with existing biomechanical analysis workflows and provides orientation information for body segments.

[0122] Since the RR1S40 dataset involves human-in-mocap-lab images of subjects with similar clothing and barefoot conditions, the robustness and performance of the model areaffected when subjects wear different clothing and footwear. To overcome this issue, the keypoint detection model is trained using a mixture of the RR1S40 dataset and diverse human-in-the-wild images. This helps to improve the generalization ability of the model across different clothing and footwear scenarios, making it more useful in real-world applications.

[0123] Markerless Motion Capture Datasets

[0124] To improve the precision of keypoint annotation, anatomical landmarks that are well-established in marker-based mocap may be used to train a deep learning model to output the marker positions. To integrate anatomical priors into markerless mocap, it requires synchronized and calibrated multi-view RGB videos that are paired with corresponding marker -based mocap data.

[0125] Table 1 shows an overview of such datasets that may be used for training and benchmarking purposes.

[0126] Table 1: Comparison of datasets with synchronized multi -view RGB videos and marker-based motion capture (mocap) system.

[0127] The HumanEva dataset is one of the earliest datasets that was made publicly available to develop and evaluate human pose estimation algorithms. The Human3.6M dataset increased the size of the dataset significantly and included additional data from a depth camera and 3D body scans of all subjects. Both datasets were created to capture natural-looking image data that may be used to train realistic human sensing systems. Although 3D marker positions are not available for both datasets, the 3D joint locations from a 3D human model are provided. However, joint locations obtained from 3D human models differ significantly from those obtained from 3D anatomical landmarks as shown in FIG. 6 depicting a cropped image 660 from the Human3.6M dataset with 3D modelbased keypoints projected onto the 2D image. The model fitting done in this dataset may introduces error. In this case (FIG. 6), the knee joint centers are shifted significantly from the midpoint of the two knee markers. Similarly, the hip joint centers are around the samelevel as LASIS and RASIS markers on the pelvis (Anterior Superior Iliac Spine) instead of being lower.

[0128] The GPJATK dataset is mainly created for gait evaluation and identification, thus it only includes walking motions. Similarly, the ENSAM dataset is adapted for clinical gait analysis and includes the motion of pathological cases. To reduce errors related to marker misplacement for lower limbs, biplanar X-ray images were also acquired. The study found that training pose estimation method on the Human3.6M dataset and then fine-tuning it on the ENSAM training set significantly reduces joint position error compared to not finetuning on the ENSAM training set.

[0129] The MoVi dataset is a large multimodal dataset with different combinations of optical mocap data, video data, and inertial measurement units (IMU). It includes 90 subjects who performed 21 everyday actions and sports movements. Besides generating a 3D human mesh model, it also releases 3D marker positions, but for some of the images, there are misalignments between the projected and actual markers, as shown in FIG. 7, depicting cropped images 760 from the MoVi dataset with 3D markers projected onto the 2D image plane. The inconsistency in camera calibration and synchronization caused misalignment between projected (white pixels) and actual markers.

[0130] It is aimed in this work to create a high-quality dataset by precisely calibrating both the marker-based and markerless systems. Similar to the MoVi dataset, the subjects wore minimal clothing to minimize marker movement relative to the body. Although this limits the variety of subjects’ appearance, it is necessary for ensuring accurate marker data, as errors may arise from marker placements on regular clothing. The RRIS40 dataset includes over 200 subjects performing various activities. These subjects are part of an Asian-centric movement database.

[0131] Pipelines

[0132] There are two main pipelines for this work.

[0133] The first pipeline is for data collection, pre-processing, and model training. The first pipeline (proposed method and / or system, respectively) may be described in similar contexts to the method 100 and the system 300 for generating a training dataset to train a machine learning model for inferring virtual keypoint locations (FIG. 1 , FIG. 3) and themethod 220 for training a machine learning model for inferring virtual keypoint locations (FIG. 2).

[0134] The second pipeline applies the trained model on multi-view videos to detect 2D keypoints and triangulate them to obtain 3D virtual marker trajectories. The second pipeline (proposed method and / or system, respectively) may be described in similar contexts to the method 440 and the system 540 for inferring virtual keypoint locations (FIG. 4, FIG. 5)

[0135] In relation to the first pipeline, FIG. 8 shows a flow chart 800 depicting a simplified overview of the training data selection and training method used to obtain a final model that infers surface anatomical bone landmarks from images of humans in random environments and clothes, according to various examples. The human-in-mocap-lab dataset may refer to the at least one augmented dataset, the images therein may include the projected-marker-based annotated images (see Step 102 of FIG. 1) and the 2D annotation therein may refer to at least the 2D marker-based keypoint locations and the augmented keypoints. The human-in-the-wild dataset may refer to the at least one in-the-wild dataset, the images therein may include the manually annotated images (see Step 104 of FIG. 1) and the 2D annotation therein may refer to at least the 2D manually-annotated keypoint locations.

[0136] The data pool with a combined set of keypoints (i.e. from the human-in-mocap-lab dataset and the human-in-the-wild dataset), based on a sampling ratio (also see Step 104 of FIG. 1), may be used for training (see Step 222 of FIG. 2). The trained model (i.e. the trained machine learning model) predicts the whole set of combined keypoints (i.e. the plurality of anatomical keypoint 2D location predictions, and the plurality of wild keypoint 2D location predictions - see Step 232 of FIG.2). To arrive at the final model, all the wild kepoint predictions (i.e. the plurality of wild keypoint 2D location predictions) may be discarded.

[0137] It should be noted that the following details are about one specific experiment that applied this technique. However, this technique could also be used with different machine learning models / architectures, different sets of keypoints / markers, or different datasets that are similar.

[0138] FIG. 9 shows a schematic view representing the data collection and training pipeline 900, according to one example of the first pipeline. Part (A) refers to the experimental setup, Part (B) refers to the training data collection and pre-processing, and Part (C) refers to the model training.

[0139] Part (A) - Sensing Hardware and Camera Configuration

[0140] A marker-based optical mocap system (e.g. the optical marker-based motion capture system 303 of FIG. 3) and multiple RGB video cameras (e.g. colour video cameras 307 of FIG. 3) are used to collect the RR1S40 dataset. The mocap system includes 16 Arqus A12 and Miqus M3 (Qualisys AB, Gothenburg, Sweden). The model of the RGB video cameras is See3CAM_24CUG (e-con Systems India Pvt Ltd, Chennai, India) with a resolution of 1920 x 1200 pixels. Each camera is equipped with a varifocal lens, which allows for adjustable focal length, to change the size of field of view for optimal coverage of the capture area.

[0141] The mocap system captures data at a rate of about 200 Hz. To ensure frame -level synchronization between the mocap system and the video cameras, an electronic circuit is used to receive synchronization pulses from the mocap system, generate synchronized pulses at a rate of about 50 Hz, and transmit shutter pulses via copper wires to all the video cameras.

[0142] All the video cameras have been positioned at a height of 170 cm above the ground, facing towards a central capture area as shown in FIG. 10 depicting a top-view map 1080 of all the camera placements and the capture volume used in the RRIS40 test set. This uniform height ensures that the training images are captured from a consistent range of perspectives. Additionally, most tripods can reach a height of 170 cm without requiring any additional framework for mounting the cameras, making deployment easier.

[0143] To ensure precise calibration, each video camera is fitted with three 1-watt white LEDs around the lens. These LEDs enable a regular video camera, which can only detect light in the visible spectrum, to see a round retro-reflective marker as a bright spot in the captured image. When the mocap system detects this marker in 3D space and the video camera simultaneously detects it in 2D on the image, a 2D-3D correspondence pair is formed. A sufficient number of these correspondence pairs collected throughout the capturevolume may be used to accurately calculate camera pose (extrinsic parameters) and finetune intrinsic parameters.

[0144] An important camera setting to consider is the exposure time, which is adjusted to minimize motion blur in the captured image. When recording video of human subjects, an exposure time of 3.888 ms is chosen to ensure that even during very fast movements, the edges of the human silhouette remain sharp. However, during calibration where the target object is a retro-reflective marker that can move faster than a human body, the exposure time is reduced to 0.996 ms. This results in a very dark capture environment, but the reflection from the retro-reflective marker is still bright enough to be detected.

[0145] Part (B) - Training Data Collection and Pre-processing

[0146] The training data contains images from the video cameras, annotated with 2D anatomical landmarks, and bounding box of the target subject. This section details how the data is collected and pre-processed before training a deep learning model.

[0147] A total of 40 markers are used in this work, which is a subset of the marker set for RRlS’s Asian-centric movement database. Marker clusters are not included as their placement is inconsistent across subjects and their large size may cause difficulty in the marker removal step later. The 40 markers consist of 4 on the head, 4 on the thorax, 4 on the pelvis, 7 on each upper limb, and 7 on each lower limb. Tire placement of these markers is carried out by trained personnel and standardized according to anatomical landmarks as shown in FIG. 11 (depicting a schematic view 1180 of marker placements).

[0148] The annotation in this group includes 40 anatomical keypoints as follows :-1. RTEMP Head Over right temple2. RHE AD Head Middle of the right eyebrow3. LHEAD Head Middle of the left eyebrow4. LTEMP Head Over left temple5. RACR Right scapula Superior aspect of acromion6. RHLE Right elbow Lateral epicondyle7. RHME Right elbow Medial epicondyle8. RRSP Right wrist Radial styloid process9. RUSP Right wrist Ulnar styloid process

[0149] During the quality control step, any marker that is found to be shifted or placed incorrectly are removed and marked as unlabeled.

[0150] After post-processing of the mocap data, instead of directly using the nearest marker-based 3D sample, the 3D marker trajectories with a sampling rate of about 200 Hz are linearly interpolated from the two nearest samples at the middle timing of the video camera exposure and projected to each video camera to obtain 2D positions with pixel - level accuracy. An example of this projection is shown in FIG. 12 depicting a cropped image 1280 with all the crosses display projected 2D marker positions. Pixel-level precision during fast movements like jumping requires careful spatial and temporal alignment between the mocap system and the video cameras.

[0151] Training images captured from the video cameras always contain visible marker blobs. Thus, a deep learning model may learn to rely on the gray blobs of the visible markers and use them as a key feature to locate the marker itself. This overfitting may lead to a degradation in performance when there is no marker on the subject during actual markerless mocap.

[0152] Therefore, it is important to prepare the training data as if there is no marker on the subject. First, the pixels occupied by the marker are to be identified. This is done automatically by taking the 2D projection and drawing a 2D radius around it. The radius size depends on the distance between the camera and the marker with an additional margin to cover the base and the shadow of the marker. Next, DeepFillv2 image inpainting technique that leverages a Generative Adversarial Network (GAN) is used to replace the pixel color in the target area by being aware of the surrounding context. The result is shown in FIG. 13 depicting cropped images 1380 for an example of marker removal before (left) and after (right) using GAN-based context-aware image inpainting.

[0153] Work on hand tracking has found that raw image data containing hand markers may affect the training process as the markers provide extra features. Thus, it was proposed for a marker removal network (MR-Net) including two stages: marker synthetization and marker removal. While the MR-Net works well for hand context, where hands are usually bare, training on humans with varying clothing may be challenging.

[0154] Additionally, the work suggests that applying a CycleGAN-based method to the whole image may result in unnatural artifacts. Therefore, in this work, image inpainting is applied only to the pixels surrounding the marker region.

[0155] When using multiple video cameras to capture a subject, it is common to have other non-subject humans present in the field of view. These individuals do not have markers and cannot be labeled, which may confuse the deep learning model. Therefore, non-subject humans are automatically detected using default human detection from Detectron2 and blurred within a bounding box with smooth edges. It is important to note that all the pixels within the bounding box of the target subject are not blurred.

[0156] Apart from 2D anatomical landmarks, 2D bounding box (e.g. the projected-marker- based bounding boxes) around each target subject is also required. This rectangular bounding box does cover not only the projected marker positions but also the full silhouette of all body parts. For example, even though there is no marker on the finger, the elbow, wrist, and hand markers are used to approximate the possible volume that the finger reaches. The 3D points on the surface of this volume are then projected onto each camera to approximate the bounding box.

[0157] Part (C) - Neural Network Architecture and Training

[0158] The model (e.g. the machine learning model) uses the keypoint detection variant of Mask R-CNN with a Feature Pyramid Network (FPN) as the feature extraction backbone. The network is based on Detectron2 implementation as it is except for the number of output keypoints. The architecture is designed to produce a bounding box around each human subject and generates a 2D heatmap for each keypoint (i.e. 2D heatmaps indicating the keypoint positions) within its respective bounding box.

[0159] To allow robust detection of anatomical keypoints across various outfits and environments, the training data combines annotated images from two sources. The first source contains around 5.39 million human-in-mocap-lab unique images that are sampled from the RRIS40 training set which is a human-in-mocap-lab dataset. The annotation in this group includes 40 anatomical keypoints as described in Part (B) above.

[0160] The second source contains 56,599 human-in-the-wild images (e.g the in-the-wild dataset - see Step 104 of FIG. 1 ) with annotations from the COCO-WholeBody dataset.For this group, a subset of 27 wild keypoints is used for training. The keypoints are 12 joint centers from the hips, knees, ankles, shoulders, elbows, and wrists, 5 head keypoints from the nose, eyes, and ears, 6 foot keypoints from the big toes, little toes, and heel centers, and 4 hand keypoints from the index and little finger metacarpophalangeal. roi6ii Since none of the 40 anatomical keypoints overlap with the 27 wild keypoints, the keypoint head of the Mask R-CNN is adjusted to predict all 67 keypoints simultaneously. For the RRIS40 training set, 12 keypoints containing the joint centers from the hips, knees, ankles, shoulders, elbows, and wrists are augmented in the annotation if possible. These joint centers may be calculated from anatomical marker positions based on existing anatomical studies. The hip joint centers are calculated using the Bell and Brand hip joint in CODA pelvis coordinate system from C-Motion's guideline. The shoulder joint centers are calculated using the Rab Upper Extremity Model from CMotion's guideline. For the knee joint, the midpoint between FLE and FME markers is used. For the ankle joint, the midpoint between FAL and TAM markers is used. For the elbow joint, the midpoint between HLE and HME markers is used. For the wrist joint, the midpoint between RSP and USP markers is used.

[0162] For the keypoints that do not exist on each training data and cannot be augmented accurately, they are counted as unlabeled.

[0163] For model training, a batch size of 8 images is used and each epoch consumes a random mix of 1 round of the human-in-mocap-lab dataset and 24 rounds of the human-in- the-wild dataset, such that the sampling ratio between the human-in-mocap-lab dataset and the human-in-the-wild dataset is approximately 80:20. The learning rate starts from 0.02 and decreases to 0.002, 0.0002, and 0.00002 at iterations 425,100, 850,100, and 3,050,100 respectively. Before iteration 1,250,000, only the keypoint head is unfrozen, while all the weights and biases are unfrozen thereafter. After the training is completed at iteration 3,100,000 (which consumes about 4 epochs of sampled human-in-mocap-lab data and about 96 epochs of human-in-the-wild data), only the 40 anatomical keypoints are used in the subsequent triangulation step, while the 27 wild keypoints are discarded.

[0164] FIG. 14 shows a schematic view representing the 3D virtual marker tracking pipeline 1440, according to one example of the second pipeline. Part (D) refers to the triangulation.

[0165] Part (D) — Triangulation (Strategic Triangulation)

[0166] After obtaining the predicted 2D anatomical keypoints from multiple images of a subject without markers, the next step is to convert them into 3D virtual markers through triangulation.

[0167] To do this, a 2D location on an image from a single camera may be represented as a 3D ray that points out from the camera’s origin. The triangulation formula then calculates a virtual intersection point of all these rays, resulting in a 3D point. In an ideal scenario, the distance between the 3D point and each ray would be relatively small, less than 10 cm, making it easy to accept the solution. However, in reality, some cameras may fail to capture certain points due to occlusion or confusion between the left and right sides. To ensure that the triangulation method remains robust even if individual camera predictions are inaccurate, the data from cameras that deviate significantly from the consensus or fail to maintain continuity of marker trajectory may be rejected.

[0168] Triangulating one marker trajectory at a time is as follows.

[0169] For each frame, perform weighted ray-distance-based triangulation (refer to Part (D) — Triangulation (Weighted Ray-distance-based Triangulation) below) using all the possible combinations of cameras.

[0170] Among the triangulation combinations in a frame, for a combination with n rays, if at least one combination with greater than n rays with a smaller maxRayDistance exists, the combination with n rays is eliminated. This helps to remove noisy data. The term maxRayDistance means the perpendicular distance from the 3D triangulated point to the farthest ray used in that triangulation combination.

[0171] From the remaining combinations, choose one combination in each frame to find the trajectory with the lowest cost. The cost, as discussed earlier on, is calculated using the formula:

[0172] m / is the distance multiplier with a default value of 1. If at least one of the selected triangulated points has only two contributed rays, my becomes 1.5 to penalize the selection of combinations with too few rays.

[0173] By using this cost function design, unstable trajectories and triangulation combinations may be automatically rejected. Searching for the optimum trajectory may be done in polynomial time with dynamic programming.

[0174] Part (D) — Triangulation (Weighted Ray-distance-based Triangulation)

[0175] It is common for a neural network that locates keypoints in a 2D space to provide a confidence score for each point.

[0176] For example, the keypoint detection version of Mask R-CNN generates a heatmap of confidence within the bounding box for every keypoint. The 2D location with the highest confidence in the heatmap is selected as the answer. In this case, the confidence score at the peak corresponds to the score for that 2D keypoint prediction. Usually, this score is ignored in a normal triangulation. However, the weighted triangulation formula allows the utilization of the score as the triangulation weight to enhance the accuracy of the triangulation.

[0177] As discussed earlier on, the triangulated 3D position (P) may be:

[0178] The directional vector (U, from Q[) of each back-projected ray is calculated by undistorting the 2D observation using cv2. undistortPointsIter to the normalized coordinate. Next, forms a 3D directional vector in the camera reference frame \x_undistorted, y_undistorted, 1]Tand rotates the direction to the global reference frame using the camera orientation. Lastly, normalizes the vector to get the unit directional vector

[0179] As this formula is derived by minimizing the weighted sum of the square of the distance between the triangulated point and all the rays, the prediction with a lower confidence has a lesser influence in the triangulation. This allows the triangulated point to be closer to the ray with a higher predictive confidence resulting in better overall accuracy.

[0180] Post Inference Procedure for Evaluation

[0181] The triangulated virtual markers are then compared to the actual marker position retrieved from a marker-based motion capture system that was running in parallel to get the error for quantitative comparison. Two test datasets are used in this experiment.RRIS40 test set contains around 0.1 million frames from 10 unseen subjects. GPJATK dataset contains around 19 thousand frames from 32 unseen subjects walking in an unseen environment.[01821 (I) Evaluation on RRIS40 Test Set

[0183] A straightforward way to evaluate the accuracy of the proposed system’s virtual marker trajectories is to compare them against the actual marker trajectories from the marker-based mocap system. The RRIS40 test set of around 0.1 million frames from 10 subjects is processed for this comparison. The subjects (5 males, 5 females), were on average 29 years old (range: 21-63 years old), mean height was 1.67 m (range: 1.48-1.83 m), mean body mass was 63.9 kg (range: 47.0-91.2 kg), and mean BMI was 22.5 kg / m2(range: 18.4-29.3 kg / m2). The proposed method outputs 40 virtual anatomical markers through strategic triangulation without using additional outlier removal, gap filling, low- pass filter, or inverse-kinematic model fitting.

[0184] A total of 4,081,976 Euclidean errors have been calculated and their cumulative distribution plot 1680 is shown in FIG. 16 illustrating overall accuracy profile of the proposed method with three different neural network models from all the 40 virtual markers. At every error threshold up to 100 mm, a ratio of error samples within that threshold is plotted. This plot is used to determine the quality of measurements in all evaluations, based on the ratio of the area under the curve (AUG) up to a cutoff of 100 mm. An AUC of 1.0 indicates perfect measurement without any error. The AUC and the error distribution for specific markers are shown in FIG. 15 for benchmarking results of the proposed method with three different neural network models (DWPose-1 and RTMPose-x are trained for ablation study) on the RRIS40 test set using ground truth from marker-based mocap. Bottom of FIG. 15: Box plots of the Euclidean errors of virtual markers (lower is better). The whiskers of the box plots extend to the minimum and maximum errors without any other samples beyond the maximum boundary. Top of FIG. 15: The ratio of AUC for each virtual marker (higher is better). Overall, the mean Euclidean error is 13.23 mm and the median Euclidean error is 10.92 mm.

[0185] Among all the anatomical keypoints, the pelvis markers which are the AS1S and PSIS, are the least accurate. This trend is consistent with a previous study which alsoshowed that the error in marker placement on the pelvic bone landmarks is larger than any other body parts. This suggests that the high levels of error are partially caused by the displacement of the markers in the train or test data.F0186] (II) Evaluation against an Anatomical-marker-based Motion Capture System on RRIS40 Test Set

[0187] Among the markerless multi-view mocap systems, OpenCap is the only one that provides 3D anatomical marker positions in its outputs. However, unlike the proposed method which trains a network to infer 2D anatomical landmarks from the image for direct triangulation, OpenCap uses OpenPose to infer 2D non-anatomical keypoints such as joint centers for triangulation and then uses another network to augment 3D anatomical landmarks from the triangulated 3D non-anatomical keypoints.

[0188] Since the anatomical keypoints of OpenCap do not completely overlap with those of RRIS40, only 23 keypoints are compared against the marker trajectories from the marker-based mocap system in the RRIS40 test set. To remove the variation of calibration methods, the same set of intrinsic and extrinsic camera parameters are used in this comparison.

[0189] FIG. 17 shows benchmarking results on the RRIS40 test set. Bottom of FIG. 17: Error comparisons of overlapping virtual anatomical markers between the propsoed method and OpenCap (lower is better). The proposed method consistently produces significantly lower errors than OpenCap. Top of Fig. 17: The ratio of area under the curve (AUC) for each virtual marker (higher is better).

[0190] FIG. 18 shows a plot 1880 depicting overall accuracy profile of the proposed method and OpenCap from all the 23 overlapping virtual anatomical markers with a total of 2,337,253 Euclidean error samples from 10 test subjects from the RRIS40 test set.

[0191] Based on the results in FIGS. 17 and 18, the proposed method of direct 2D prediction and triangulation of anatomical landmarks is significantly better than OpenCap’ s 3D augmentation method in all comparisons. The proposed method displays an overall error of 14.90 ± 9.76 mm, which is lower than OpenCap’s overall error of 33.22 ± 22.01 mm.

[0192] (III) Evaluation against Joint-center-based Pose Estimation Tools

[0193] To compare against a wider range of methods, it is necessary to evaluate joint center locations. Although the proposed system does not directly output 3D joint centers, these joint centers may be calculated from virtual anatomical markers based on existing anatomical studies.

[0194] For example, the hip joint centers are calculated using the Bell and Brand hip joint in CODA pelvis coordinate system from CMotion’s guideline. The shoulder joint centers are calculated using the Rab Upper Extremity Model from CMotion’s guideline. For the knee joint, the midpoint between FLE and FME markers is used. For the ankle joint, the midpoint between FAL and TAM markers is used. For the elbow joint, the midpoint between FILE and HME markers is used. For the wrist joint, the midpoint between RSP and USP markers is used. The ground truth joint centers from marker -based mocap system are calculated using the same formula.

[0195] The proposed model is compared with three other tools: OpenCap, MediaPipe, and Detectron2. For OpenCap, its default OpenPose 2D keypoint detection, its own triangulation algorithm, and its default filter are used to obtain all the 3D joint center sequences. For MediaPipe, since it does not provide a multi-view pipeline, the proposed 2D keypoint detection model is replaced by MediaPipe’s pre-trained model. The subsequent step is performed on 2D joint centers using our triangulation algorithm. For Detectron2, the proposed 2D keypoint detection model is replaced by the pre-trained model from Detectron2’s repository. This pre-trained model has the same neural network architecture as that of the proposed method, which is the keypoint detection variant of Mask R-CNN. The differences are the set of keypoints, the source of training data, and annotation. Detectron2’s pre-trained model is trained on Microsoft’s COCO dataset with 17 hand-annotated 2D keypoints.

[0196] FIG. 19 shows benchmarking results on the RRIS40 test set. Bottom of FIG. 19: Error comparisons of different methods based on joint centers (lower is better). The proposed method has significantly lower errors than other methods in every key joint center. Top of FIG. 19: The ratio of area under the curve (AUC) from each virtual marker (higher is better).

[0197] FIG. 20 shows a plot 2080 depicting overall accuracy profile of the proposed method, Detectron2, OpenCap, and MediaPipe from all the 12 key joint centers with a total of 1,169,671 Euclidean error samples from 10 test subjects from the RRIS40 test set.

[0198] From the results shown in FIGS. 19 and 20, the proposed method has an overall joint center error of 13.71 ± 10.43 mm, which is significantly lower than Detectron2 (25.16 ± 24.11 mm), OpenCap (28.50 ± 22.14 mm), and MediaPipe (29.25 ± 19.53 mm). The results suggest that including anatomical marker-based annotations in the training data leads to improved accuracy.

[0199] (IV) Evaluation on GPJATK Dataset

[0200] Among all the datasets in Table 1 above, GPJATK is the only dataset that provides raw marker location data from a marker -based mocap system, along with synchronized RGB video data from more than two calibrated viewpoints. Therefore, this dataset is chosen to validate the proposed method and check whether the learned anatomical keypoint may be transferred across different camera models, placements, and capture environments.

[0201] Although the RR1S40 marker set shares 15 common markers with the GPJATK marker set, the placement position might differ slightly. For example, the shoulder marker for the RRIS40 marker set is positioned on the Acromion bone landmark, whereas the GPJATK marker set might place it on the Acromioclavicular joint, which is around one centimeter away.

[0202] FIG. 21 depicts benchmarking results on the GPJATK dataset on 15 overlapping markers. Bottom of FIG. 21 : Error comparisons of different methods based on joint center (lower is better). The errors from the proposed method are shown in two different variants. The first variant is trained using a mixture of the RRIS40 dataset and the wild dataset. The second variant is trained using the RRIS40 dataset alone. Because the triangulated results from the second variant contain some gaps in marker trajectories, those gaps are always filled with linear interpolation. OpenCap results are included as references. Top of FIG. 21: The ratio of the AUC from each virtual marker (higher is better). OpenCap does not provide XPRO, T10, LFCC, and RFCC to compare against those marker data in the GPJATK dataset. In FIG. 21, it can be seen that the proposed method outperforms Open- Cap in all the overlapping markers. However, the performance is worse than the benchmarkresults on the RRIS40 test set. Several factors may explain these performance differences such as the placement of markers, a reduction in the number of cameras from 8 to 4, a decrease in camera resolution from 1920 x 1200 to 960 x 540, differences in the capture environment, or the use of a different calibration method. Additionally, the images captured around the head and shoulder area of the subjects are blurred to remove personal identifiers.

[0203] (V) Impact of Mixing Wild Images without Anatomical Keypoints in the Training Data

[0204] One common challenge of using training data from a single environment is the low transferability of the model to a new environment. To determine whether the mixing of human-in-the-wild data may improve the model’s transferability, another model is trained solely on the RRIS40 dataset and used as a baseline for comparison.

[0205] The comparison of both models on the RRIS40 test set is shown in FIG. 22. Bottom of FIG. 22: Error comparisons across all anatomical markers (lower is better). The errors from our method are shown in two different variants. The first variant is trained using a mixture of the RR1S40 dataset and the wild dataset. The second variant is trained using the RRIS40 dataset alone. The # mark means that the errors from the first variance are significantly smaller than the errors from the second variant. Top of FIG. 22: The ratio of the AUC from each virtual marker (higher is better). Although the model trained on mixed data shows a significantly smaller overall error, the difference is not decisive as only 21 out of 40 markers display significantly lesser errors.

[0206] In contrast, when both models are compared on the GPJATK dataset which is an unseen environment, a clear difference is shown in FIG. 21. The model trained only on inlab data produces much noisier results, as it struggles with human detection. In some frames, the model fails to detect the subject in one or more cameras, leading to a lower number of available triangulation rays. This ultimately results in either inaccurate triangulation or not enough rays to triangulate.

[0207] Incorporating wild images in the training data effectively and advantageously eliminates overfitting from the human detection portion of the network, as demonstrated by the results obtained from the GPJATK dataset. Furthermore, this mixing of data doesnot appear to have any adverse impact on the accuracy of the system in the seen environment like the RR1S40 test set.

[0208] (VI) Qualitative Evaluation on Assistive Outfits

[0209] Markerless mocap has a large potential for uses in rehabilitation research and assistive robotics. This is especially true in scenarios where marker-based mocap systems face challenges. For instance, when a subject is equipped with a walking aid, a robotic device, an exoskeleton, or a safety harness, marker attachment becomes difficult. This is because the markers cannot be attached to all the bone landmarks as they would be blocked by the obtrusive wearable pieces of equipment. Therefore, the proposed markerless mocap system is tested to assess its stability and performance qualitatively when a subject wears unseen outfits.

[0210] In this experiment, a subject is recorded while walking with different types of walking aids which include a walking frame, an exoskeleton, a balance assistant robot, and an overhead body weight support system. FIG. 23 shows snapshots of video records with 2D projections of 3D virtual markers while the subject is walking with different walking aids. Visually, all the triangulated 3D virtual markers are stable and well-aligned with the corresponding body parts. FIG. 24 depicts shows images from the COCO dataset, with overlying 2D keypoints inferred by the proposed trained machine learning model.

[0211] (VII) Impact of Camera Reduction

[0212] To determine the impact of removing certain cameras, the entire RRIS40 test set is reprocessed 247 times with all possible camera combinations. In other words, there are 247 average errors from 247 possible camera combinations. The average error from each combination is plotted in FIG. 25. In the plot 2580 of FIG. 25, each black point represents an average error from one combination of cameras when the results have been compared to marker-based mocap positions 4,081,976 times. When the number of cameras is limited, bad placement of the cameras may increase the overall measurement error up to 3.6 times. As expected, the range of average error increases as the number of cameras decreases. Since, the range of the average error is much larger when there are only two or threecameras, identifying the configurations that contribute to the best and worst results can be helpful in optimizing the camera arrangement in many scenarios.

[0213] Table 2 is produced to analyze the characteristics of the best and worst camera placement combinations.

[0214] Table 2: the best and the worst camera configurations where the position of each camera index is shown in FIG. 10.

[0215] The worst configuration in the case of two cameras is when the cameras are positioned directly opposite each other. This arrangement is not ideal because the 2D keypoints extracted from cameras 3 and 4 provide very little information on the depth, that is, the global’s X component. Even a slight 2D error in at least one of the cameras may significantly shift the triangulated position along the depth axis of both cameras. This issue may be avoided if the two cameras are positioned so that their depth axes are angled close to 90 degrees. That is why the best configuration in the case of two cameras is the one with cameras 0 and 7.

[0216] (VIII) Impact of Lower Image Resolution

[0217] The original video image with a resolution of RRIS40 test set is always scaled down to 1280 x 800 for the first layer input of the neural network. To determine the impact of reduced image resolution, the image is scaled down to three additional resolutions: 320 x 200, 640 x 400, and 960 x 600 before scaling it back to 1280 x 800 for the network input. FIG. 26 depicts the comparison of overall Euclidean errors with varying image resolutions. The results in the plot 2680 of FIG. 26 show a gradual increase in overall Euclidean errorwith lower image resolution. This decline in performance is likely due to the loss of spatial accuracy in 2D prediction, which is expected as the back-projected ray from each camera may be sensitive to even a few pixels shift in 2D.[02181 (IX) Impact of Neural Network Model Selection

[0219] In this ablation study, the Detectron2 is compared with two recent models: RTMPose-x and DWPose-1. The training is done on the MMPose platform. Similarly to the original way of RTMPose-x training, the first and second stages of training with different augmentation pipelines run for 3 and 1 epochs respectively. Then, DWPose-1 first- stage distillation training runs for 4 epochs using the trained RTMPose-x as the teacher and RTMPose-1 as the student architecture. Then, the second-stage distillation is trained for 1 epoch. For all the training in this section, all the models are trained on one NVIDIA Titan RTX GPU using the same amount and mixture of datasets as the Detectron2 training but with a batch size of 72. To perform evaluation, RTMDet-m is chosen for the human detection step before passing the bounding box to RTMPose-x and DWPose-1 for 40 keypoint detection with the same detection input size of 384 x 288. The full inference pipeline (inclusive of RTMDet-m) for RTMPose-x and DWPose-1 require computation of 56.35 and 48.44 GFlops / image respectively which are comparable to the Detectron2 with an input size of 1280 x 800 operating at 51.89 GFlops / image.

[0220] The results in FIGS. 15 and 16 show that all three models are closely comparable. The mean Euclidean errors for the Detectron2, DWPose-1, and RTMPose-x are 13.23, 13.62, and 13.63 mm respectively. It is observed that Detectron2 may have falling-to-the- edge issues when the foot markers are close to an edge of the bounding box as shown in FIG. 24. However, RTMPose-x and DWPose-1 do not exhibit such issues, which may be due to their finer bin resolution. These findings show the potential of newer network architectures and offer opportunities for future exploration.

[0221] The proposed system is designed for multi-view setups that capture the entire human body, which is commonly used in 3D mocap applications like gait analysis. It should be noted that its performance may vary when dealing with heavily cropped human images or when tracking poses that are uncommon or absent in the training datasets. For example, actions like lying on the ground are absent from the RRIS40 dataset as thephysical markers may be easily shifted or occluded during this action. Additionally, the use of the inpainting technique for marker removal may introduce unique artifacts that may bias the training dataset. Since the RRIS40 dataset mainly consists of Asian subjects, it may limit the adaptability of our model to other demographic groups. Variations in height, gender, and body shape within a normal range do not significantly affect the model’s generalizability.

[0222] The proposed method and system contribute to advancing the state-of-the-art in markerless motion capture (mocap), addressing the limitations of existing methods, and facilitating applications in clinical biomechanics and sports science. An important advantage of markerless mocap is that it reduces the time and manpower required for subject preparation and data postprocessing. This makes the mocap workflow more efficient and practical for real-world applications

[0223] This highlights the potential of facilitating wider integration of markerless mocap into biomechanics research, with at least the following further advantages.

[0224] From the user's perspective, having surface anatomical marker keypoints as an output from a markerless human motion capture system allows users to continue using marker-based biomechanical calculation workflow and established formulation from the literature with minimal to no adjustment. Users can enjoy a well-established marker-based workflow without having to deal with issues caused by markers such as long subject preparation time, the need for skilled marker placement personnel, marker drops, unnatural movement, and long manual data post-processing time.

[0225] Having a model that only works robustly in the motion capture room that is used for training data collection has very limited usability. Making it work robustly in other unseen capture environments largely boosts the usability of the system. The same idea works for better variation of clothing support.

[0226] From the aspect of data annotation efficiency, the average time taken to annotate an image using a marker -based motion capture system is much lower than the time taken to manually annotate a human-in-the-wild image with the same marker set. Being able to skip the annotation of the human-in-the-wild photo set with 40 or more surface anatomical markers can save a lot of time.

[0227] When comparing the proposed method to other methods of augmenting those surface anatomical keypoints in 3D from the existing set of wild keypoints, the experiments show that the accuracy of the proposed method is significantly better.

[0228] While the invention has been particularly shown and described with reference to specific embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention as defined by the appended claims. The scope of the invention is thus indicated by the appended claims and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced.

Claims

CLAIMS1. A method for generating a training dataset to train a machine learning model for inferring virtual keypoint locations, the method comprising: obtaining at least one augmented dataset comprising projected-marker-based annotated images of subjects; and sampling from the at least one augmented dataset and at least one in-the-wild dataset based on a sampling ratio to generate the training dataset, wherein the at least one in-the-wild dataset comprises manually annotated images of random subjects under unrestricted conditions, the projected-marker-based annotated images of the subjects comprise 2D markerbased keypoint locations, augmented keypoints and projected-marker-based bounding boxes, the manually annotated images of the random subjects comprise 2D manually- annotated keypoint locations and manually-annotated bounding boxes, and one or more of the augmented keypoints respectively coincide with one or more of the 2D manually-annotated keypoint locations.

2. The method as claimed in claim 1, wherein obtaining the at least one augmented dataset comprises: based on a plurality of physical markers captured by an optical marker-based motion capture system, each as a 3D trajectory, wherein each physical marker is placed on a bone landmark or a keypoint of each subject, and the subjects substantially simultaneously captured by a plurality of colour video cameras over a period of time as sequences of 2D images, for each physical marker, identifying the captured 3D trajectory with a marker label representative of the bone landmark or keypoint on which the physical marker is placed, obtaining augmented markers, each augmented marker being calculated from positions of two or more of the physical markers placed on each subject; determining an augmented 3D trajectory with an augmented label representative of each augmented marker;for each physical marker, projecting the 3D trajectory and for each augmented marker, projecting the augmented 3D trajectory to each of the 2D images to determine a 2D location in each 2D image; for each physical marker and for each augmented marker, based on the respective 2D locations in the sequences of 2D images and an exposure-related time of the plurality of colour video cameras, interpolating a 3D position for each of the 2D images; for each 2D image, based on the respective interpolated 3D positions of the plurality of physical markers and the augmented markers, and an extended volume, generating each projected-marker-based bounding box around each subject, wherein the extended volume is derived from two or more of the physical markers and / or the augmented markers having an anatomical, functional and / or structural relationship with one another, and generating the augmented dataset comprising at least one 2D image selected from the sequences of 2D images, the determined 2D location of each physical marker in the selected at least one 2D image, the determined 2D location of each augmented marker in the selected at least one 2D image, and the projected-marker-based bounding boxes for the selected at least one 2D image, wherein for each physical marker, the marker label is arranged to be propagated with each determined 2D location such that in the augmented dataset, each determined 2D location of each physical marker contains the corresponding marker label, wherein for each augmented marker, the augmented label is arranged to be propagated with each determined 2D location such that in the augmented dataset, each determined 2D location of each augmented marker contains the corresponding augmented label, wherein the 2D marker-based keypoint locations comprise the determined 2D locations of the plurality of physical markers, and wherein the augmented keypoints comprise the determined 2D locations of the augmented markers.

3. The method as claimed in claim 2, wherein the augmented markers comprise each joint center of the positions of the two or more of the physical markers placed on eachsubject and / or each 2D wild keypoint location calculated from the positions of the two or more of the physical markers placed on each subject, wherein the 2D wild keypoint location corresponds to one of the 2D manually-annotated keypoint locations in the at least one in- the-wild dataset.

4. The method as claimed in claim 2 or 3, wherein the plurality of physical markers each being captured as the 3D trajectory and the subjects being substantially simultaneously captured as the sequences of 2D images over the period of time are coordinated using a synchronized signal communicated by the optical marker-based motion capture system to the plurality of colour video cameras.

5. The method as claimed in any one of claims 2 to 4, further comprising: after projecting the 3D trajectory to each of the 2D images to determine the 2D location in each 2D image, in each 2D image and for each physical marker, drawing a 2D radius on the determined 2D location according to a distance with a predefined margin between the colour video camera and the physical marker to form an encircled area, and applying a learning-based context-aware image inpainting technique to the encircled area to remove a marker blob from the 2D location.

6. The method as claimed in any one of claims 2 to 4, further comprising: after projecting the augmented 3D trajectory to each of the 2D images to determine the 2D location in each 2D image, in each 2D image and for each augmented marker, drawing a 2D radius on the determined 2D location according to a distance with a predefined margin between the colour video camera and the augmented marker to form an encircled area, and applying a learning-based context-aware image inpainting technique to the encircled area to remove a marker blob from the 2D location.

7. The method as claimed in claim 5 or 6, wherein the learning-based context-aware image inpainting technique comprises a Generative Adversarial Network-based context- aware image inpainting technique.

8. The method as claimed in any one of claims 2 to 4, further comprising: detecting, using a human detection algorithm, non-subject humans captured in the 2D images; and blurring the non-subject humans in the 2D images to remove the nonsubject humans.

9. The method as claimed in any one of claims 2 to 8, wherein the exposure -related time involves a middle of exposure time to capture each 2D image or a part thereof using each colour video camera.

10. The method as claimed in any one of claims 1 to 9, wherein the in-the-wild dataset comprises COCO dataset, COCO-Wholebody dataset, MPII dataset, Al challenger dataset, Leeds Sports Pose dataset, or Relative Human dataset.1 1 . The method as claimed in any one of claims 1 to 10, wherein the in-the-wild dataset is obtained by capturing images of the random subjects, manually identifying a plurality of wild keypoints in each image with each manually-annotated bounding box surrounding each random subject, and labeling the plurality of wild keypoints in each image to generate the 2D manually-annotated keypoint locations in each manually annotated image.

12. The method as claimed in any one of claims 1 to 11, comprising sampling from the augmented dataset and the in-the-wild dataset based on the sampling ratio of the augmented dataset and the in-the-wild dataset at A:W, where each of A and W is greater than 0.

13. The method as claimed in any one of claims 1 to 12, comprising sampling from the augmented dataset and the in-the-wild dataset based on the sampling ratio of the augmented dataset and the in-the-wild dataset at 80:20.

14. The method as claimed in any one of claims 1 to 13, wherein the unrestricted conditions comprise random environments and the random subjects with random clothes.

15. A method for training a machine learning model for inferring virtual keypoint locations, the method comprising:(i) feeding a training dataset generated by a method as claimed in any one of claims 1 to 14 into a neural network;(ii) generating, by the neural network, outputs for one or more target subjects;(iii) comparing the outputs and annotations in the training dataset to obtain differences thereof;(iv) adjusting weights and / or biases in the neural network based on the obtained differences; and(v) repeating (i) to (iv) until the outputs substantially match the annotations or for a predetermined number of iterations to obtain the trained machine learning model, wherein the substantially matched outputs or the outputs after completing the predetermined number of iterations comprise a plurality of anatomical keypoint 2D location predictions, a plurality of wild keypoint 2D location predictions, and a bounding box prediction surrounding each target subject.

16. The method as claimed in claim 15, further comprising discarding the plurality of wild keypoint 2D location predictions from the trained machine learning model.

17. A method for inferring virtual keypoint locations, the method comprising: based on a marker-less human or animal subject captured by a plurality of colour video cameras as sequences of 2D images, for each 2D image captured by each colour video camera, generating, using a trained machine learning model, a 2D bounding box; for each 2D image, generating, by the trained machine learning model, a plurality of heatmaps with scores of confidence, wherein each heatmap is for 2D localization of a keypoint of the markerless human or animal subject, andthe trained machine learning model is trained using at least a training dataset generated by a method as claimed in any one of claims 1 to 14; for each heatmap, selecting a pixel with the highest score of confidence, and associating the selected pixel to the keypoint, thereby determining the 2D location of the keypoint, wherein for each heatmap, the scores of confidence are indicative of probability of having the associated keypoint in different 2D locations in the generated 2D bounding box; and based on the sequences of 2D images captured by the plurality of colour video cameras, triangulating the respective determined 2D locations to predict a sequence of 3D locations of the keypoint, thereby inferring the virtual keypoint locations.

18. The method as claimed in claim 17, wherein triangulating the respective determined 2D locations comprises performing strategic triangulation of one trajectory of the keypoint at a time.

19. The method as claimed in claim 18, wherein the strategic triangulation comprises: for each 2D image, performing weighted triangulation for 2N-(N+1) combinations of the plurality of colour video cameras to obtain 2N-(N+1 ) candidates of the 3D locations together with respective maxRayDistance, N being the total number of colour video cameras, and for each candidate, the maxRayDistance being a perpendicular distance from each 3D location to a farthest ray used in a corresponding combination, wherein the 2N- (N+l) combinations exclude a combination with zero colour video camera and N combinations with one colour video camera; among the 2N-(N+1) candidates in each 2D image, determining whether each candidate is noisy data based on that candidate having less rays with a smaller maxRayDistance than at least one other candidate; eliminating that candidate, if determined to be the noisy data, to obtain remaining candidates; from the remaining candidates, selecting one candidate in each time frame such that a sequence of the selected candidates forms a 3D trajectory of the keypoint with a lowest cost, wherein each time frame is each moment when N colour video cameras capture Nimages simultaneously, and wherein the lowest cost is obtained by applying a cost function given by:where ni> is a distance multiplier calculated by a monotonically decreasing function of the minimum number of colour video cameras involved in the triangulation of the selected candidates of that keypoint between two consecutive time frames,is a distance between the selected candidates between the two consecutive time frames, and / = 1, 2, total number of time frames- 1.

20. The method as claimed in claim 19, wherein the weighted triangulation comprises derivation of each predicted 3D location of the keypoint using a formula:wheregiven thatis a weight for triangulation or the score of confidence ofray from colourvideo camera, is a 3D location of the colour video camera associated with theray,is a 3D unit vector representing a back-projected direction associated with theray, is a 3x3 identity matrix.

21. The method as claimed in any one of claims 17 to 20, wherein each colour video camera comprises at least one visible light emitting diodes operable to facilitate retro- reflective markers to be perceived as detectable bright spots, and the method further comprises extrinsically calibrating the plurality of colour video cameras by based on the retro-reflective markers captured by the plurality of colour video cameras as sequences of 2D calibration images,applying an optimization function to the captured 2D calibration images to finetune extrinsic camera parameters of the plurality of colour video cameras.

22. A computer program adapted to perform a method as claimed in any one of claims 1 to 21.

23. A non-transitory computer readable medium comprising instructions which, when executed on a computer, cause the computer to perform a method as claimed in any one of claims 1 to 21.

24. A data processing apparatus comprising means for carrying out a method as claimed in any one of claims 1 to 21.

25. A system for generating a training dataset to train a machine learning model for inferring virtual keypoint locations, the system comprising: a computer configured to: obtain at least one augmented dataset comprising projected-marker-based annotated images of subjects; and sample from the at least one augmented dataset and at least one in-the-wild dataset based on a sampling ratio to generate the training dataset, wherein the at least one in-the-wild dataset comprises manually annotated images of random subjects under unrestricted conditions, the projected-marker-based annotated images of the subjects comprise 2D marker-based keypoint locations, augmented keypoints and projected-marker- based bounding boxes, the manually annotated images of the random subjects comprise 2D manually-annotated keypoint locations and manually-annotated bounding boxes, and one or more of the augmented keypoints respectively coincide with one or more of the 2D manually-annotated keypoint locations.

26. The system as claimed in claim 25, further comprising: an optical marker-based motion capture system configured to capture a plurality of physical markers over a period of time, wherein each physical marker is placed on a bone landmark or a keypoint of each subject, and is captured as a 3D trajectory; and a plurality of colour video cameras configured to capture the subjects over the period of time as sequences of 2D images, wherein to obtain the at least one augmented dataset, the computer is configured to: receive the sequences of 2D images captured by the plurality of colour video cameras and the respective 3D trajectories captured by the optical marker-based motion capture system; for each physical marker, identify the captured 3D trajectory with a marker label representative of the bone landmark or keypoint on which the physical marker is placed, obtain augmented markers, each augmented marker being calculated from positions of two or more of the physical markers placed on each subject; determine an augmented 3D trajectory with an augmented label representative of each augmented marker; for each physical marker, project the 3D trajectory and for each augmented marker, project the augmented 3D trajectory to each of the 2D images to determine a 2D location in each 2D image; for each physical marker and for each augmented marker, based on the respective 2D locations in the sequences of 2D images and an exposure-related time of the plurality of colour video cameras, interpolate a 3D position for each of the 2D images; for each 2D image, based on the respective interpolated 3D positions of the plurality of physical markers and the augmented markers, and an extended volume derived from two or more of the physical markers and / or the augmented markers having an anatomical, functional and / or structural relationship with one another, generate each projected-marker-based bounding box around each subject; andgenerate the augmented dataset comprising at least one 2D image selected from the sequences of 2D images, the determined 2D location of each physical marker in the selected at least one 2D image, the determined 2D location of each augmented marker in the selected at least one 2D image, and the projected-marker- based bounding boxes for the selected at least one 2D image, wherein for each physical marker, the marker label is arranged to be propagated with each determined 2D location such that in the augmented dataset, each determined 2D location of each physical marker contains the corresponding marker label, wherein for each augmented marker, the augmented label is arranged to be propagated with each determined 2D location such that in the augmented dataset, each determined 2D location of each augmented marker contains the corresponding augmented label, wherein the 2D marker-based keypoint locations comprise the determined 2D locations of the plurality of physical markers, and wherein the augmented keypoints comprise the determined 2D locations of the augmented markers.

27. The system as claimed in claim 26, wherein to obtain the augmented markers, the computer is configured to calculate each joint center of the positions of the two or more of the physical markers placed on each subject and / or calculate each 2D wild keypoint location from the positions of the two or more of the physical markers placed on each subject, wherein the 2D wild keypoint location corresponds to one of the 2D manually- annotated keypoint locations in the at least one in-the-wild dataset.

28. The system as claimed in claim 26 or 27, further comprising a synchronization pulse generator in communication with the optical marker-based motion capture system and the plurality of colour video cameras, wherein the synchronization pulse generator is configured to receive a synchronization signal from the optical marker -based motion capture system for coordinating the subjects to be substantially simultaneously captured by the plurality of colour video cameras.

29. The system as claimed in any one of claims 26 to 28, wherein the computer is further configured to, in each 2D image, draw a 2D radius on the determined 2D location for each physical marker according to a distance with a predefined margin between the colour video camera and the physical marker to form an encircled area, and to apply a learning-based context-aware image inpainting technique to the encircled area to remove a marker blob from the 2D location.

30. The system as claimed in any one of claims 26 to 28, wherein the computer is further configured to, in each 2D image, draw a 2D radius on the determined 2D location for each augmented marker according to a distance with a predefined margin between the colour video camera and the augmented marker to form an encircled area, and to apply a learning-based context-aware image inpainting technique to the encircled area to remove a marker blob from the 2D location.

31. The system as claimed in claim 29 or 30, wherein the learning-based context-aware image inpainting technique comprises a Generative Adversarial Network-based context- aware image inpainting technique.

32. The system as claimed in any one of claims 26 to 28, wherein the computer is further configured to execute a human detection algorithm to detect non-subject humans captured in the 2D images, and to blur the non-subject humans in the 2D images to remove the non-subject humans.

33. The system as claimed in any one of claims 25 to 32, wherein the in-the-wild dataset comprises COCO dataset, COCO-Whol ebody dataset, MPTT dataset, Al challenger dataset, Leeds Sports Pose dataset, or Relative Human dataset.

34. The system as claimed in any one of claims 25 to 33, wherein the in-the-wild dataset is obtained by capturing images of the random subjects, manually identifying a plurality of wild keypoints in each image with each manually-annotated bounding box surroundingeach random subject, and labeling the plurality of wild keypoints in each image to generate the 2D manually-annotated keypoint locations in each manually annotated image.

35. The system as claimed in any one of claims 25 to 34, wherein the computer is configured to sample from the augmented dataset and the in-the-wild dataset based on the sampling ratio of the augmented dataset and the in-the-wild dataset at A:W, where each of A and W is greater than 0.

36. The system as claimed in any one of claims 25 to 35, wherein the computer is configured to sample from the augmented dataset and the in-the-wild dataset based on the sampling ratio of the augmented dataset and the in-the-wild dataset at 80:20.

37. The system as claimed in any one of claims 25 to 36, wherein the unrestricted conditions comprise random environments and the random subjects with random clothes.

38. A system for inferring virtual keypoint locations, the system comprising: a plurality of colour video cameras configured to capture a marker-less human or animal subject as sequences of 2D images; and a computer configured to: receive the sequences of 2D images captured by the plurality of colour video cameras; for each 2D image captured by each colour video camera, generate, using a trained machine learning model, a 2D bounding box; for each 2D image, generate, using the trained machine learning model, a plurality of heatmaps with scores of confidence, wherein each heatmap is for 2D localization of a keypoint of the markerless human or animal subject, and the trained machine learning model is trained using at least a training dataset generated by a method as claimed in any one of claims 1 to 14; for each heatmap, select a pixel with the highest score of confidence, and associate the selected pixel to the keypoint to determine the 2D location of the keypoint, wherein foreach heatmap, the scores of confidence are indicative of probability of having the associated keypoint in different 2D locations in the generated 2D bounding box; and based on the sequences of 2D images captured by the plurality of colour video cameras, triangulate the respective determined 2D locations to predict a sequence of 3D locations of the keypoint, thereby inferring the virtual keypoint locations.

39. The system as claimed in claim 38, wherein to triangulate the respective determined 2D locations, the computer is configured to perform strategic triangulation of one trajectory of the keypoint at a time.

40. The system as claimed in claim 39, wherein to perform the strategic triangulation, the computer is configured to: for each 2D image, perform weighted triangulation for 2N-(N+1) combinations of the plurality of colour video cameras to obtain 2N-(N+1) candidates of the 3D locations together with respective maxRayDistance, N being the total number of colour video cameras, and for each candidate, the maxRayDistance being a perpendicular distance from each 3D location to a farthest ray used in a corresponding combination, wherein the 2N- (N+l ) combinations exclude a combination with zero colour video camera and N combinations with one colour video camera; among the 2N-(N+1) candidates in each 2D image, determine whether each candidate is noisy data based on that candidate having less rays with a smaller maxRayDistance than at least one other candidate; eliminate that candidate, if determined to be the noisy data, to obtain remaining candidates; from the remaining candidates, select one candidate in each time frame such that a sequence of the selected candidates forms a 3D trajectory of the keypoint with a lowest cost, wherein each time frame is each moment when N colour video cameras capture N images simultaneously, and wherein the lowest cost is obtained by applying a cost function given by:where is a distance multiplier calculated by a monotonically decreasing function of the minimum number of colour video cameras involved in the triangulation of the selected candidates of that keypoint between two consecutive time frames, df is a distance between the selected candidates between the two consecutive time frames, and total number of time frame- 1.

41. The system as claimed in claim 40, wherein the weighted triangulation comprises derivation of each predicted 3D location of the keypoint using a formula:wheregiven thatis a weight for triangulation or the score of confidence of ray from colourvideo camera,is a 3D location of the colour video camera associated with theray,i is a 3D unit vector representing a back-projected direction associated with the* ray, is a 3x3 identity matrix.

42. The system as claimed in any one of claims 38 to 41, wherein each colour video camera comprises at least one visible light emitting diodes operable to facilitate retro- reflective markers to be perceived as detectable bright spots, and wherein the plurality of colour video cameras is configured to capture the retro- reflective markers as sequences of 2D calibration images, and the plurality of colour video cameras is to be extrinsic ally calibrated by applying an optimization function to the captured 2D calibration images to fine-tune extrinsic camera parameters of the plurality of colour video cameras.

Citation Information

Patent Citations

  • Pose estimation systems and methods trained using motion capture data

    US20230119559A1

  • Method and system for generating a training dataset for keypoint detection, and method and system for predicting 3D locations of virtual markers on a marker-less subject

    WO2022265575A2