Machine learning system for marker-based motion capture
Patent Information
- Application Number
- PCT/GR2024/000026
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-11
- Filing Date
- 2024-09-09
- Publication Date
- 2025-08-14
AI Technical Summary
Existing marker-based motion capture systems face challenges in real-time processing and accuracy due to noise and variability in marker placements, requiring complex post-processing and high-end equipment.
A machine learning system that uses neural networks and parametric body models to estimate body landmarks from raw marker position data, incorporating a novel technique to balance training data distribution and reduce bias, allowing for real-time refinement of results.
The system achieves robust and accurate real-time motion capture with reduced noise and improved accuracy, enabling efficient processing of marker-based motion capture data using affordable and scalable technologies.
Smart Images

Figure GR2024000026_14082025_PF_FP_ABST
Abstract
Description
TITLE OF THE INVENTION
[0001] MACHINE LEARNING SYSTEM FOR MARKER-BASED MOTIONCAPTUREFIELD OF THE INVENTION
[0002] This disclosure relates generally to the field of motion capture technologies, and more specifically to real-time marker-based motion capture systems. In particular, the sys- tem processes input image data to estimate pre-defined body landmarks in real-time, be- longing either to the body’s surface or inside it. The system Hist estimates the subjects body shape using the estimated body landmarks, and then uses it to calculate the subject s articulation in time, delivering a full ID time- varying capture which generally relates it to the wider field of performance capture. While the system relies on machine learning to es- timate the body landmarks from input marker position measurements, it does not require marker position data for training. Instead, it leverages a parametric body model represen- tation to simulate marker position data and captures the necessary training body model parameters using a marker-less motion capture method. Further , using generative models, the system utilizes a. novel technique to balance the model’s training with respect to bias and the underlying data distribution. Finally, the parametric body model standardizes the system’s outputs with respect to its representation, allowing for the real-time results to be automatically refined using a post-processing step that relies on a noise adaptive and robust bundle fitting method.BACKGRO UND OF THE INVENTION[00031 Motion capture (MoCap) is the technology that digitizes the motion of rigid or articulated physical entities or objects. These can be rigid ( c.g. objects) or articulated ( e.y. people, animals). The technology finds numerous applications, with examples being the transferring of the digitized motion information into virtual enviromnents tor driving simulation (such as synthetic pre-populated worlds for rendering data for autonomous driv- ing) or increasing the realism of the depicted content (such as in video games), content compositing effects (such as virtual effects in the film industry), analysing the statistical properties of the raw motion and features deriving from it (such as health or bio-mechanics research), and the like.
[0004] There are different, ways to implement motion capture technology, and they axe typically categorized with respect to the equipment ( e.y. sensors, accessories) that are de- ployed and used to capture the subjects' (or objects’) motion. From a sensing perspective, inertial measurement units and cameras are used, with the former being body-worn, typically attached to a suit that the performer wear's, and the latter usually being aimed at a suit that lias markers attached to it instead of sensors. The markers can be either active, emitting self-identifying signals that are decoded by the cameras, or passive, relying on physical prop- erties (f. c. retro-reflective material) to improve their capturing potential by the cameras. To achieve that, cameras are equipped with infrared projectors and limit their capturing spec- trum to match the emitted light’s properties. The systems relying on markers are referred to as marker-based MoCap systems, whereas there also exist camera-based systems that do not require the subjects to wear suits and markers and simply operate on the image data provided by the cameras. These systems are referred to as marker-less systems and exploit intermediate representations estimated by algorithms operating on raw image data.[00051 MoCap technology has benefited from the last decade’s da ta-driven breakthroughs mostly due to significant research on human-centric visual understanding that focuses on unencumbered capture using raw color (u. marker-less) inputs. The gold standard of Mo- Cap technology referred to as “optical” - still uses markers attached to suits for robustand accurate captures and has received little attention in the modern machine learning literature.
[0006] The remainder of this section will discuss marker- based MoCap related art using machine learning for the purpose of assisting reader comprehension and understanding of the advantages of the invention.[0007, Machine learning based MoCap systems using markers mainly focus on processing (raw) archival MoCap data for the purposes of directly labeling them [G / torbani Ac et al.. 2010], [Black et al.. 2023\ or labeling them through regression [Han et a I., 2018[\ solving the skeleton’s joints [Chen et al. , 2021] or transforms [Holden, 2018] and also integrating commodity sensors that introduce higher noise levels \ Chert zito fir et al. , 2021],
[0008] As even high-end systems produce output with varying noise levels, be it either information- (swaps, occlusions and ghosting), or measurement-related (jitter, positional shifts), all aforementioned approaches exploit the plain nature of raw marker representation to add synthetic noise during training. Still, for data-driven systems, the variability of marker placements (i.e. layouts) comprises another challenge that needs to be overcome.[00091 It is possible to address this implicitly, relying on the learning process, or address this quasi-explicitly, considering them as input to the model. Another way to overcome this involves fitting the raw data to a parametric model after manually [ / / an et al... 2018. Loper et al., 20140 or automatically [ Ghorba.nl N. et al. , 2019\. [ Black et al , 2023\ labeling and / or annotating correspondences, standardizing the underlying representation.[0010| Solving the joints’ positions from marker data is a cascade of numerous (sometimes optional) steps. The markers need to be labeled, ghost markers H,. e. measurement outliers) need to be identified, occlusions (b e. missing measurements) should be predicted and then a skeletal or rigid bodies structure needs to be lit to the observed marker data. Various works address errors at different stages of MoCap solving, with contemporary ones relying on smoothness and bone-related (angles, offsets and lengths) constraints.
[0011] Other approaches resort to existing data for initialization [Schubert et al. , 2015[or marker cleaning [Aristldov. et al. , 2018\. MoSli [Loper et al.. 2014] instead, employed a parametric human body to solve raw marker data and estimate post* articulation arid joint positions, even accounting for marker layout inconsistencies and / or soft tissue motion.
[0012] Nonetheless the advent of modern deep data-driven technologies have stimu- lated new approaches for MoCap solving. A label-via-regression approach [Han et al, 2018\ used a deep model to regress marker positions and then perform maximum assignment matching for labeling the input markers. Labeling has also been formulated as permutation learning [Ghorbani S. et al., 201.9], albeit with constraints on the input, which can be relaxed by adding an extra ghost category to label [Ghorbani N. et al., 2019], [Black el al. , 20231
[0013] However, labeling assumes that the raw data are of a certain quality as the raw measurements are then used to solve for the joints’ transforms or extra processing steps are required to denoise the input. As a result, end-to-end data-driven approaches that can simultaneously denoise and solve have been a parallel line of research.
[0014] Even though it is possible to perform end-to-eiid cleaning and solving using a single feed-forward network in [Holden, 2018[, tire process naturally' benefits from using two cascaded autoencoders [Paella et al., 2019], with the first operating on marker data and cleaning them for the subsequent joint regressor. The staging from markers to joints has been shown to improve accuracy even when using convolutional neural networks but this comes at the cost of run-time performance [Chatzitofis et al.. 2021]. Graph convolutional networks are another implementation candidate [Chen et al... 2021] as they allow for encoding topological features like an explicit marker layout and the skeleton hierarchy, two important variations that were only implicitly handled in the aforementioned end-to-end solvers.
[0015] Learning to solve MoCap marker data requires supervision provided by collecting data using professional high-end MoCap systems. Marker data can be easily (re-)syuthesized when higher level information is available [e.g. marker- to- joint offsets, meshes), but it is also possible to use synthetic model fits to real data, even when using consmner-grade sensors, to acquire and then synthesize data.
[0016] Statistical parametric models \ Loper et al.. 2015] (SMPL), [Osman et al. , 202' 0\ (STAR), [Osman el al. , 2022] (SUPR). | Wang el, al., 2021] (PanoMan), [L a el. al.. 202' 0\ (GHUM), | Yan et al. , 2021] (DeepDaz) are more expressive alternatives than simple1skinned meshes as apart from realistic' shape variations, deformation corrective factors can also be leveraged to improve the fitting result. They have been used to synthesize training data before lor photorealistic rendering j Varol el al.. 2017] and simulating sensor measurements \Kam}mann et al , 2021], but still, these crucially rely on preceding high-end MoCap acqui- sition.
[0017] These low-dimensional parametric model representations can be fit to marker data, with notable examples being the AMASS dataset [Mahmood et al.. 2019] that consolidates an assortment of MoCap datasets using the SMPL family of models, and FrtTL) [Lleraru et al. , 2021] using the GHUM model. The potential of acquiring such data using less expensive capture solutions is very important, as long as it is feasible to train high quality models. Still, one also needs to take into account, ’ lie nature of human performance data and their collection processes. As seen in both AMASS arid Fit3D. they contain significant redundancies and suiter from the long-tail distribution effect.
[0018] Rare poses are challenging for regression models to predict, a challenge stemming mainly from tire combined effect of the selected estimators and stochastic optimization with mini-ba.tcbes. Various solutions have been surfacing in the literature, some tailored to the nature of the problem [Rong et al., 2022], leveraging a prototype classifier branch to initialize the learned iterative refinement, and others adapting works from imbalanced classification to the regression, domain.
[0019] Traditional approaches fall into either the re-sampling or re-weighting category, with the former focusing on balancing the frequency of sample's and the lat ter on properly adjusting the parameter optimization process.
[0020] Re-sampling strategies involve common sample under-sampling [ Torgo et al. , 2015], rare sample over-sampling by synthesizing new samples via interpolation | Torgo et al.,2013], re-sampling them after perturbing them with noise, and hybrid approaches that simul- taneously under- and over-sample \J3raneo el al.. 2017], Yet interpolating high-dimensional samples like human pose is problematic or even defining the rare samples that need to be re-sampled.>[0021 Utility-based [ Torgo et al., 2007] - or otherwise, cost-sensitive regression assigns different weights > or relevance > to dilferent samples. Defining a utility function is also important to re-sampling strategies for regression | Torgo el al., 2015] , Recent approaches employ kernel density estimation | Ste al., 2021]. adapt evaluation metrics as losses\Silva el al. , 2022] or resort to label / feature smoothing and binning [ Yang et al., 2021].
[0022] Another family of methods that are now explored can be categorized as contrastive, with \ Gong et al. , 2022] regularizing training to enforce feature and output space proximity. A balanced version of the mean squared error (MSE) [Ren el al., 2022] is also a contrastive- like objective that employs intra-batch minimum error sample classification using a cross- entropy term that corresponds to an L2 error from a likelihood perspective.[00231 However, most approaches rely on binning the output space, which when consid- ering human pose, is a challenging task,[0U24| Moreover, capturing motion on a per-frame basis suiters from noisy estimates or fitting to noisy observations. While temporal filtering is one option to resolve the temporal jitter and inconsistency of per-frame captures [Ingwersen at al.. 2023], it is very challenging to design filters to handle occlusions and outliers. Most approaches integrate temporal constraints, like joint smoothness \Arnab et al., 2012] , | Ye et al. , 2023], or velocity constancy [Loper et al., 2014], [Zanfir et al., 2018], [Mahmood et al.. 2019], when solving tor the pose and motion, in addition to exploiting causality by solving for a munber of frames simultaneously. A DCT smoothness term lias also been used when fitting the SMPL body to multiple images [Huang et al. , 2017]. More recent approaches rely on data-driven priors, like the feature space of a velocity autoencoder [Huang et al. , 2022], the latent space of poses [Huang et al., 2021]. motion [Sami et al. , 2023], or the latent space and transition state of an autoregressivemotion prior [Kempe et al. , 2021].A common theme to all prior approaches is that motion smoothness is achieved by the addition of one or more objectives to the optimization problem. This incurs the cost of additional parameters that need tuning, a case tha t is especially complex when considering the high number of objectives that are involved, when fitting a parametric body model to i iu ages.
[0025] Adding such temporal constraints is more effective when solving for a bundle of frames simultaneously. The AMS prior used by [Huang et al., was added a t he second stage of optimization when a group of 30 frames was solved simultaneously Similarly, joint smoothness was imposed on a group of frames solved simultaneously {Saini et al.. 2t)2S\, [ he et al. , 2023] . Apart from solving a small bundle, it has also been shown that it is possible to solve entire videos simultaneously [Arnab et al. , 2010], exploiting the entire temporal context, even with very complex objective functions {Huang et al., 2022] involving human- object interactions. Still, this requires good initialization, which can be achieved using a regressor model [Arnab et al. , 2013] or a per-frame body fitting stage {Huang et al.. 2022]. This type of staged optimization is standard in the ill-posed monocular case [I’avlakos et al., 2019], but also finds widespread use in multiview settings [ Huany el al., 2017], [Saint el al. , 2023], When considering videos, multiple stages are used to initialize estimates for each separate frame, and then solve bundles using smoothness constraints [Huang et al. , 2017], | Saini et al.. 2023], [Ye et al.. 2023] , [Arnab et al. , 2019], be it either via direct regression models or solving an optimization problem for each frame,
[0026] Nonetheless, this is rather costly as multiple passes are required over flic videos. Additionally, the temporal nature of the data is only considered at tire constraints level, but not at the parameters level, with each frame's state being solved separately in the bundle.BRIEF SUMMARY OF THE INVENTION
[0027] The present invention disdoses a marker-based motion capture system that; esti- mates an object ’s or subject’s known landmarks, and additionally its pose and shape vising image-based sensor (s), essentially delivering a four-dimensional ('41.)) digital representation of the object’s motion and / or subject’s performance, meaning its three dimensional (3D) time-varying motion across the time dimension.[0028i The system operates in two modes, either in real-time, where tlie input is processed at interactive rates, offering a live visualization of the output, as well as post-capture, where (die reakdime outputs are refined.[0029J The disclosed invention uses neural network parameter databases and a body model database. Appropriately matching combinations of neural networks and body models are used at rimtime to capture the object / subject.
[0030] In one embodiment the system only produces pre-defined landmarks in real- time using raw marker position information as acquired and estimated by the image- based sensor(s). The system processes the observed unstructured landmarks producing higher quality, labeled estimates.
[0031] hi another embodiment, the system processes raw unstructured landmarks de- rived from marker position measurements and estimates the corresponding body surface landmarks, essentially, completing, refining and intruding the marker input estimates, to produce outputs that cannot be easily measured in the necessary accuracy and precision.
[0032] In a further embodiment, it additionally solves for unobserved landmarks like internal landmarks (e.f / . human joints) or body surface landmarks, using only extruded marker position measurements as input.
[0033] In a further preferred embodiment, it also estimates the subject’s pose and shape, by fitting it to the estimated landmarks through optimization.
[0034] In a further preferred embodiment, the subject’s shape is initially calibrated, by fitting the parametric body model to the estimated landmarks through optimization. The calibrated shape is then fixed and used when solving for the subject s pose to produce the41) digital representation of the captured performance.
[0035] Advantageously, the disclosed invention uses a machine learning model in the form of a neural network to process the unstructured marker position input and can hallucinate (body) occluded markers, remove erroneous estimates that do riot reflect the actual markers placed on the subject (ghostings), de-noise the input markers (refined positioning). This makes the system robust to contemporary MoCap system tracking artifacts.
[0036] Another advantageous feature of this invention is that the same neural network machine learning model can lie trained to solve for internal landmarks that represent, the captured object’s or subject’s inner (skeletal) structure
[0037] Yet another advantageous fea ture of this invention is that the same neural network machine learning model can be trained to solve for body surface landmarks that represent the captured object’s or subject’s surface.
[0038] In a most advantageous embodiment, this invention uses a neural network machine learning model that is trained to solve for body surface landmarks and the body surface landmarks that represent the captured object’s or subject’s skeletal structure and body surface sinrultaneously.[00391 The invention discloses a method for training the machine learning model usefl by the system using data from various sources. Tliis is an advantageous feature of the invention as the system can be fed with data acquired by more practical and affordable solutions without sacrificing quality or performance.[00401 lu one embodiment the system can use data from contemporary MoCap systems.[00411 In a preferred embodiment, the system can use data acquired from (a) low-cost camera(s) using marker-less MoCap technology.[00421 In a further preferred embodiment, the system can use data acquired from different sources including contemporary MoCap systems and affordable markerdess systems. Tliis makes the system scalable with respect, to the data that will be used to continuously improve it. It also makes the system flexible as it can lie specialized to newly acquired data subsetscollected by different systems than the previously collected data.[00431 An advantageous feature that is also disclosed in this invention is a method to better exploit the available data to improve the training of the machine learning model by reducing the bias of the collected data The method adjusts the training data distribution to reduce the effect of its long-tailed nature, as well as make it robust to the significant bias depicted in the captured poses Advantageously this is accomplished with no additional data and automatically.[0044j Yet another advantageous feature that is also disclosed in this invention is a method used by the system to fit the body shape and pose to noisy estimates, a condition absent from the prior art, as it conventionally relied on high-quality estimates from expensive high-end systems. This makes the present invention adjustable to various sensing options and more importantly, affordable options.
[0045] A further preferred embodiment includes the robustly trained neural network to estimate the landmarks and tire noise-aware fitting process to produce high-quality landmark, pose and shape estimates, reconstructing a four-dimensional digital avatar representation of the captured performance.
[0046] Advantageously, the present invention relies on body model parameters. This al- lows for the acquisition of human motion data using marker-less capture technology, therefore solving the cold-start problem of delivering the said system. One would require an existing motion capture system to acquire model training data, hindering scalability, increasing costs and time, and requiring complex processes to standardize the data. By using marker-less capture and fitting body model parameters to landmarks from image data, data acquisition is significantly simplified, scalable, and defaulted to a standardized representation, reducing post-processing effort .[00471 Another advantageous feature of this invent ion is that it applies to systems which provide minimal optical sensor data and operate in environments with minimal constraints.
[0048] Yet another advantageous feature of the disclosed invention is a method that ex-plicitly models measurement and / or inferred uncertainty when fitting human body parame- ters to 31) or 21) landmarks, allowing for increased robustness to input landmark noise.
[0049] Further, another advantageous feature of the disclosed invention is tin’ way tlie neural network combined with the parametric body model standardizes the landmark rep- resentation, and by extension, the MoCap output data.
[0050] The final advantageous feature exploits the standardized output to perform a post-capture generative refinement process that employs past and future information to efficiently coned erroneous estimates and produce smooth motion, even in highly challenging deploy merit environments .BRIEF DESCRIPTION OF THE DR AWINGS
[0051] FIG. 1 is a schematic illustration of an apparatus for data-driven motion cap- ture solving noisy landmarks using a machine learning model, optical / image sensors, a low- dimensional body model and a. noise-aware fitting method.
[0052] FIG, 2 illustrates a process of training the neural network presented as a module in the apparatus of FIG, 1.
[0053] FIG. 3 illustrates a process of capturing, processing, and storing real data for training the neural network presented as a module in the apparatus of FIG. 1. The real data are used as input in the training process of FIG. 2
[0054] FIG. 4 is a schematic representation of the autoencoding generative model trained for generating synthetic data. The synthetic data are used as input in the training process of FIG. 2
[0055] FIG. 5 illustrates a process of selecting latent samples from a shaped manifold, transforming them to end up with new latent samples on the same manifold, reconstructing them using the decoder of the generative model, and storing the reconstructed samples in the synthetic data database.
[0056] FIG. 6 illustrates the application of an exponential relevance function on the reconstruction error. The error values are colorized using binary colormap where black indicates the maximum error, whereas the input data are visualized in transparent-gray color for comparison purposes.
[0057] FIG. 7 illustrates the process of refining a recording produced by the real-time motion capture system using a generative model to jointly optimize over a bundle ol frames.[00581 FIG. 8 is a block diagram of an exemplary computing environment within which various embodiments of the invention may Ire in ip loin cut cd and upon which various embod- iments of the invention may lie employed.DETAILED DESCRIPTION OF THE INVENTION AND DRAWINGSDescription will now be given with reference to the attached Figs. 1-8. It should be understood that these Figures are exemplary m nature and in no way serve to limit the scope of the invention.
[0059] The disclosed invention uses a parametric model 'A, whose parameters are re- trieved from a database 108, that jointly encodes pose articulation and shape as its MoGap representation. After reconstructing the shape, a set of pre-defined joints inside the shape are used together with a set of articulation parameters to deform and pose the shape using a skimiing method. Typically, skinning artifacts are accounted for post-deformation using offsets calculated from the available (shape and pose) parameters. Different variants ex- ist. all data-driven, some relying on stochastic representations [Xu d al. , 2020\, others on explicit ones | Loper et al.. 201o\, \ Osman et al. , 2020], with a notable exception using an artist made one [ han et a,l. , 2021] arid all typically employ linear blend skinning arid pose corrective factors [Loper el al. , 2015], [Xu el all.. 20201 to overcome its artifacts. The present invention is not bound to a specific implementation of ® and can be implemented wi th any of the available variants as long as the low-dimensional representation can be transformed toa 3D surface defining one (e.p. mesh, point cloud). As the pose corrective deformations can be considered as optional, the present invention can be implemented purely with an artist- made linear blendshape representation for the shape, and an associated kinematic tree > derived either manually or automatically LZfomm et al.. 2007\ accompanied by bi-harmonic skinning weights \ Jacobson et al.. 2011],|0060J In tins invention, ® is generally defined as a function (v, f) = fof Pfo. T h where \ vA} are the vertices v G Rv'x3and faces f t>, ,’Jof a mesh surface that is defined by 5 blendshape coefficients P G R‘n posed by P pose parameters 0 G SO(3pb and globally positioned by the transform T > G SE(3).
[0061] After reconstructing a posed shape, represented as a mesh, it is possible to use linear functions r expressed as matrices X to extract body landmarks := r(v ) . f( x v. with f t R1"'5and Jj E This way, surface points 2' can tie extracted using delta(vertex picking) or barycentric (face interpolation) functions, and joints G using weighted average functions. Since markers are extruded by the marker size d they correspond to t',z= f + rfft x n), with n being the vertices' surface's orientation (t. e. the normal's direction),[00621 While the following description and illustrations will focus on parametric human body models, a skilled person will realize that the present invention s system and methods can be implemented with any parametric or articulated model, representing for example animals [Zufft et al. , 2017], or articulated objects [Xue et al. , 2021\ .[00631 FIG. 1 illustrates a machine learning system 100 that uses image-based sensors 101 to acquire raw unstructured marker positions 102 as input to a neural network 103 and estimates parameters associated to the parametric body model representation 104, be it either explicit 104 or implicit 106. Explicit refers to the landmark 104 or mesh-based representation comprising vertices v and f, or representations extracted by the mesh-based representation like marker and joint landmarks 104, whereas implicit refers to the parametric representation 106 comprising the shape [3, pose 0 and transform T parameters,
[0064] While in this particular illustration 100 the intermediate representation output bythe neural network 103 is explicit (i.e. landmarks 104). and the final optimized representation is implicit (i e. body model parameters 106), these can be interchanged. Another example would be to output an implicit representation from the neural network 103, optimize it for improving its accuracy, and then extracting the explicit representation from the relined implicit one. Generally, the invention refers to input unstructured landmark estimates 102 derived from image-based / optical sensing data 101, a machine learning model 103 that receives these as input and whose parameters are selected from a database of parameters 107. a parametric body model of an articulated object / subject whose paraineters are selected from a database of parameters 108, and an optimization process 105 that relines t lie estimated body model representation derived from the neural network, to produce fitted parameters that better describe the observations. These are streamed to a data storage system 111 (i.e. in memory, a database or a tile).
[0065] In addition, the fitting process 105 can benefit from a pre-calibrated shape jj 109 as the optimization 105 would only need to solve for the pose parameters & and the transform Tz106 for each time instance / , compared to also solving for the shape, given that the shape is largely constant during a single performance.
[0066] Further, the fitting process 105 can also exploit higher level priors lor the pose, as expressed by the prior encoded by generative models 110, which encode high dimensional distributions. This is a step beyond fixed discrete constraints (e.g. joint limits) towards smoother distribution constraints (c.y. the normal distribution encoded by variational am toencoders or iiormalizi ng flow models).
[0067] The stored captured performances 111 can lie refined 600 through the use of generative models 110 lay exploiting the completeness of temporal information, producing artefact-free, natural, smooth motions.[00681 FIG. 2 illustrates the neural network 103 training process that estimates the net- work’s parameters and adds them to the database 107. The neural network 103 that is used in this invention receives as input training data the parameters of a body model fB. Thesecan be either parameters acquired from real world captures 201 or synthetic (f. e, computer- generated) parameters 202. Be it from real world captures or synthetically generated, the resulting representation is of synthetic nature, and as a result can be augmented 203, and / or corrupted 205 with artifacts and noise, essentially perturbing the training data to increase their diversity and simulate realistic acquisition conditions. The data augmentation and / or corruption can take place in both representations of the body model ®, explicit ( v,f) or implicit (p, 6,T).[00691 FIG. 2 shows the training process that begins by loading samples from a database of body model parameters, augment s and / or corrupts them, extracts the representation that will be used to feed into the neural network as training samples 204, performs a stochastic optimization iteration to refine its parameters arid repeats this process for a pre-specified number of iterations or unt il deemed as appropriate to stop 206. The resulting neural network 103 parameters are then stored into a model parameter database 107 along with appropriate metadata describing the model and the training details. The latter include the extracted representation 204 that was chosen as input to the neural network.[0070| In this particular illustration, a set of body surface landmarks is considered, ex- truded by a pre-defined distance d, to simulate marker positions. However, there is no specific restriction to this representation as it can be accompanied by any other transforma- tion acting on the explicit or implicit body representation. One example would be a different rotation parameterization of the pose parameters used to define the body s articulation.[0071 i Apart from describing ihe machine learning model 103 that will overcome the challenges associated to unstructured marker input and the defects of marker-based cap- tures ( e.g. marker swaps, occlusions, ghosting) , the disclosed system benefits from a training method 200 that delivers robust inference performance overcoming the long-tailed distribu- tion of training data, and additionally a model fitting method that is robust to varying input noise levels 105, Further, as the system uses a standardized representation in the form of the body model ®, it benefits from training data acquired by multiple sources and differenttechniques. Therefore, in addition to acquiring data using expensive MoCap systems, the neural network can be trained with data 201,300 acquired from affordable systems 301 using marker-less MoCap technology 302.
[0072] Overall the present invention is an economical MoCap system as it can be deployed using consumer-grade, affordable sensors 101. and can exploit data acquired by lower-cost MoCap technology 300.201 to continuously improve its performance by training unproved neural network variants 200. All methods and neural network training schemes can be implemented by means of a computer having a processor and storage as described below The following text will describe the system’s machine learning model robust training scheme and the model tilting optimization and refinement approaches in detail.Balanced Regression .
[0073] Relevance functions are the drivers of utility regression [Branco et al., 20BI] , [ 7'oryo el al. , 2015\, [ Torgo el al.. 20071. \ Toryo el al., 2013\ and have been also guiding the re- / over- / inter-sample selection / generation. Instead of defining relevance or sample selection based on an explicit formula or set of rules, this invention employs representation learning to learn it directly from the data.
[0074] Autoencoding generative machine learning models 207 \Kingma el al. , 201B\- [Rezende et al.. 2015\ jointly learn a reconstruction model 401 as well as a generative sampler 202:with varying constraints on the input 0 and latent z = E(0),z t R1spaces, a typical one being a Gaussian likelihood prior imposed on the latent space. An encoder 15(0) maps input 0 to a latent space z which gets reconstructed to 0+by a generator (7[z). Using a sampling function 5 to sample the latent space' it is also possible to generate novel output samples 0' 202.
[0075] This invention exploits the hybrid nature of such models 207 as an imbalancedregression solution that simultaneously over-samples the distribution at the tail 202 and adjusts the optimization by re-weighting rarer samples 208.Rel e wince b tin ct i on .
[0076] Autoencoding models are expected to reflect the bias of their training data, with redundant / rare samples being easier / harder to properly reconstruct respectively. This bias in reconstructability can be used to assign relevance to each sample as those more challenging to reconstruct properly are more likely to be tail samples. A relevance function p is defined: (2) using a reconstruction error £:with ( " ) denoting tuiit normalization using the input joints’ bounding box diagonal, e the norrnalized-RMSE over the reconstructed and original joints, and c a scaling factor controlling flic relevance p.
[0077] Using landmark positions preserves interpretable semantics in p and (7 as they are unidirectionally interchangeable (linear mapping) with the pose 0 given fixed shape p. FIG. 6 shows exemplary input-reconstruction pairs as scored 'by the relevance function in Eq. (2).
[0078] The relevance function Eq. (2) 208 is used as a weight for each sample during the supervised loss calculation. Compared to traditional utility-based regression | Torgo ct al.. 2015\, \ Targo et al.. 20u0\, \ Torgo et al. , 201 dj whose relevance functions downweight common samples, the present- invention instead boosts the loss and by extension gradients derived from each sample that is riot reconstructed faithfully by the autoencoding synthesis model of Eq. ( 1 ] 207.Controlled Synthesis.
[0079] Even though the tail samples are not reconstructed faithfully , the generative nature of modern synthesis models shape manifolds that map inputs to the underlying factors of data variation, effectively mapping similar poses to nearby latent codes which can be traversed across the latent space’s dimensions. Considering this, the present invention defines a controlled sampling scheme for synthesizing new tail samples 500 (see lug. 5).
[0080] Using the relevance function from Eq. (2) 208 tail samples IT can be automatically identified through statistical thresholding and create a sei; of anchor latent codes. Then the following sampling function is used:(4)
[0081] Eq. (4) 501. samples from a normal distribution centered around two random an- chors i. j from A using a standard deviation s, and blends them 502 using spherical linear interpolation g with a uniformly sampled blending factor b G 0(0, b"). Non-linear interpola- tion between samples avoids dead manifold regions as not all directions lead to meaningful samples and increases our samples’ plausibility. This is one specific embodiment of using a generative autoencoding model to balance the neural net work's training,[00821 Advantageously, a process comprising of a latent space sample selection 501 of S samples and then a function to transform 502 them represents a family of embodiments of the disclosed method. The sampling function can be a random sampler from a normal distribution Ag(0,I) G KL, or a discrete anchor sampler, or the aforementioned random anchor neighbor sampler, while tire transformation function can be any interpolation function or even the identity function when S = 1. Generally, the sampling function 501 produces 5 > 1 latent space samples, and the transforming function 502 generates a single sample out of the S sampled ones.
[0083] This controlled tai] sampling process 500 is schematically illustrated in FIG. 5 andbenefits the present invention by over-sampling the training data distribution with samples that are synthesized. This presents several advantages as the diversity ol tire generated samples is beneficial to the neural network’s training, but also, its controllability umskews the long-tailed distribution of motion capture data samples.
[0084] Advantageously, the present invention is designed to work with any autoencoding generative model 207, or even a cascade of autoencoding generative models, capable of generating and reconstructing data. To validate our method, a Variational AutoEncoder (VAE) is employed, which is trained with tlie following training setting: ( O ) (6) (7) < 8;where z G f?2’2is the 32-dim latent code, R G 50(3) is tire 3 x 3 rotation matrix tor each joint, while R is the 3 x 3 output of the decoder. V, V corresponds to the predicted and ground truth vertices, inchoating that the reconstruction term incorporates both angular and positional errors. The manifold shaping is regularized using the Kullback-Leibler (KL1 divergence with a Charbormier penalty functionto prevent posterior collapse and to decompose the factors of variation more effectively. Eq. (6,7) follow file VAE training scheme - e.g., trading of reconstruction quality with learning a Gaussian-like manifold, while Eq. (7,8) force the model to construct a valid rotation latent space The VAE training scheme is complemented by the weight-decaying Adarn optimization method that penalizes large weights and prevents overfitting. Two metrics are used to choose the best-performing model for tail-pose generation and regression regularization; (i) the Frechet inception distance for quantifying the generative model’s ability to synthesize high-quality data, and (ii) a diversitymetric for evaluating the ability of the model to generate diverse data, l ire model with the beat diversity-quality trade-off has been selected for 500 and 200.
[0085] Table 1 illustrates an example and summarizes the regression results on the THu- man2.0 dataset. The THuman2.0 dataset is a collection of diverse poses captured from 5 subjects. Table 1 reports the Root Mean Square Error (RMSE) that measures the Euclidean distance between predicted and ground truth landmarks, arid the 31) variant of the Percent- age of Correct Keypoints (PCK) that measures the percentage of predicted landmarks that lie within the following radius y of the ground truth position: a) y = 0.01m denoted as PCK1. b) y = 0.03m denoted as PCK3, and c) 7 = 0.07m denoted as PCK7. Table 1 shows a per- formance comparison between (i) 103 trained with a regression loss function and denoted as ‘ Baseline’ , (ii) 103 trained with the Balanced Mean Squared Error (BMSE) [foc / i el al.. ill)22\ loss function and denoted as ‘BMSE’, and (iii) 103 trained with a regression loss, samples using 500, regularization with 208, and denoted as ‘fours’. The ‘Ours’ model performs better than the other two in all metrics, indicating that the combination of controlled synthesis and regularizing with a relevance function can regress diverse poses from datasets that do not present heavy pose bias but do not comprise solely rare poses.
[0086] Table 2 illustrates another example and summarizes the regression results on acustom set of rare poses that include ‘high kicks', ‘crossed legs’, ’crossed arms , and •crouch- ing'. The results are the average of the model performance on all 4 custom tail sets. Table 2 reports the R.MSE and the PCK for the 3 different radius values as presented above. 'Table 1 shows a performance comparison between (1) 103 trained with a regression loss function and denoted as ‘Baseline’, (it) 103 trained with the BMSE loss function and denoted as ‘BMSE’. and (Hi) 103 trained with a regression loss, samples using 500, regularization with 208, and denoted as ‘Ours; The ‘Ours’ model performs better than the 'Baseline’ under all metrics while achieving better RMSE. PCK3, and PCK7 performance compared to the ‘BMSE’. indicating that it learns to regress the overall pose robustly although losing some accuracy in specific body parts.
[0087] A skilled practitioner will realize that any neural network satisfying the properties of autoencoding (i.e. data sample reconstruction) and generation (i.c. data sampling) can be used to implement the aforementioned balanced regression training scheme. Additionally, said person will also understand that the autoencoding generative model can only be used for either relevance / importance assignment or synthetic data sampling when used to balance the training of a regression neural network.Real-time Landmark Estimation.
[0088] Compared to pure labeling \ Ghorbani N. et al., 2021\, | Black at al,, 202tJ\, \ Ghor- bani S. at al. , 2019\ or pure solving approaches [Holden, 2018\, \ Ghan ct al. , 2021\ the present invention’s neural network model 103 simultaneously denoises, estimates and hallucinates landmarks. With raw marker positions as input 102. different options are available with re- spect to the neural network that- will be used. Transformer-based models can purely operate on the unstructured raw marker position input whereas convolutional neural networks need to first convert it to a structured representation. The benefit of the former is the infinite input granularity, whereas the corresponding limitation that is imposed when transforming the raw marker position representation into a structured grid is controlled by the use of well developed output representations like heatmaps. Heatmaps open up richer supervision andcoordinate decoding techniques than the pure regression alternative that transformers oiler Nevertheless, the present invention can benefit from both architectures, and is not limited to these only as the training regime 200 is agnostic to the neural network architecture at hand.[0089) The goal is to transform the unstructured, noisy and potentially in- / over-com plcte marker positions into pre-determined at the model training phase body associated land- marks. The controlled sampling 500 and relevance function 208 weighting are applied prior and post neural network pass, meaning that any black-box neural network 103 that satisfies these relaxed constraints in additional to real-time run-time performance is a possible iuiplenientatiou candidate.[0090J While some methods regress surface marker and others joint landmarks, the present invention benefits from jointly regressing surface (sub- vertex), outer (extruded mark- ers) and inner (joints) landmarks as this boosts performance and adds utility with respect to body fitting and joint rotation solving.[00911 As shown in FIG. 2, during training the mini-batches randomly contain synthetic samples 202 acquired from Eq. (4) using the model formulated by Eq. (1) (right) 207 and each sample is weighted using Eq. (2) after reconstructing it via Eq. (1) (left) 208.[00921 Advantageously, the present invention is designed to work with any neural network 103, or even a casca.de of neural networks, capable of inferring markers, surface points and joints from an input markers’ point cloud. To validate the model, one of the most commonly used architectures was employed that have been demonstrated to achieve state-ol-the-art. results in various applications, and is also more efficient from a run-time performance point of view and allows us to regress high-resolution heatmaps. A modified version of the UlNet [Ronneberger et al.. 2015\ neural network architecture was impleinented for 103 to predict 53 markers and 18 joints landmarks simultaneously.
[0093] The model consists of five blocks, with each block consisting of 32, 64. 128, 256, and 512 features, respectively. Each encoder block is comprised of two convolution blocks, witha kernel size of 3. a stride and a padding of 1, followed by the ReLU activation and batch normalization luuctions, which are commonly used in deep neural networks to overcome vanishing and exploding gradients. Downscaling across blocks is achieved via anti-aliased max pooling [Zhang et al... .260,9], which helps to reduce aliasing and produce smoother results, '1'he bottleneck of the model consists of a single convolution block, utilizing the same parameters as the encoder blocks. Tlie decoder mcludes the same convolution blocks, and the output of each block is combined with the corresponding encoder’s output after (bi-) linear upsampling. Finally, the prediction layer consists of a convolution block with a kernel size of 1, a stride of 1 , and padding of 0. with Re LU activation. 'Fins layer produces the final output, of the model, which is used to estimate the predicted landmarks,
[0094] The raw unstructured input markers' positions Nt102 are normalized to [0, 1] and rasterized using orthographic projection to produce two orthogonal normalized depth maps. Assuming prior knowledge of the gravity direction along the y axis, these orthogonal ami orthographic views denoted as ,ry and vz share the y axis. Then, using a variant of marginal heatmap regression [Nibali at al.. 2019\. [ Ye at al., 2022] the model is supervised using center-of-mass regression [Nibali et al... 2018g [Sun et al.. 2018] to estimate the landmarks’ normalized positions taking the average expectation [AhViah e,l al.. 2019\ . [ he el al. , 2022] for the y axis.
[0095] Advantageously, instead of requiring the model to predict the extra viewpoint as done for monocular images, the 3D nature of the input data combined with the orthogonal view depth map rendering, allows for explicit prediction of each input view that will be tlien fused, reducing the domain and learning gap.[00961 The model is supervised by:where Ljg is the X— weighted Jensen-Shannon divergence [Menendez et al. , 1997] betweenthe normalized ground truth and soft-max normalized predicted heatmaps, while th. is the robust Welsch penalty function Holland el ai 19771, with the support parameter V. between the rasterized normalized landmark f coordinates
[0097] Overall, A / v accelerates mini-batch training 206 while facilitates higher levels of sub-pixel accuracy, even though the heatmaps H are reconstructed using the normalized - un- quantized coordinates [Zhang el all.. 2020\, discretization artifacts can never be completely removed.[00981 lb train the neural network, the synthetic nature of tin' low-dimensional body model parameters allows us to generate the training data on-the-fly. This data generation pipeline aims to closely approximate real- world MoCap settings by accounting for subject body shape variations, and input-related issues such as ghost markers, marker occlusions, and marker noise. While an exemplary pipeline will lie followingly described, it is not restricted to these functions, as any transformative operations to reduce the synthetic-to-real data gap, and produce meaningful data variations are eligible.[00991 Initially, body model ® parameter variations 203 are performed beginning with a random controlled shifting of the shape P coefficients:using a uniform distribution around the normal standard deviation.Additionally, a subset of the shape parameters p are randomly replaced: Cmwhere S is a set of n' indices sampled uniformly from the set of indices.A random left-right flip augmentation is then applied to account: for right-handedness.
[0100] Then, a number of synthetic corruptions 205 are also applied, with the following being non-restrictive, indicative examples. To simulate marker occlusions, we follow the process below. Then, we randomly select a subset of markers for occlusion by determining the number of markers to be occluded, denoted as k. We draw a random sample from a discrete uniform distribution to determine Ay k U(m,n'), m < k < n1< n, where is the uniform distribution over the range of integers . . . . fo}, and nzdefines the maximum number of markers to be occluded. Next, we draw' another random sample from a uniform distribution to determine the indices of the markers to be occluded, i.e.. m = \ ni\ . m-2 , . . . , ng)' 1 , ri) , k < n where ‘Llkl , n) is the uniform distribut ion over the markers set; of indices. The resulting vector m contains the indices of t he markers to be occluded and is used to exclude these markers from f".As a next step, the ghosting of markers is emulated by extracting samples from a Gaussian distribution with mean and standard deviation values equivalent to the original marker positions. In more detail, t he median position for each spatial dimension of the marker positions py- is first computed, (?. e. the median value for the / -th spatial dimension of the marker positions), and the sample covariance matrix A, Then, samples g ~ (yfofol) are drawn and appended to the original markers' positions k'kFinally, to simulate marker noise, a set of markers are randomly selected to have noise added on them With N being the number of markers to shift, and M being the maximum allowable shift distance, we randomly sample from a uniform distribution to determine the indices of the markers to which the noise will lie added / Ukl jk). For each indexrandom offset vector o ~ U( M,M) is generated, and add this offset to the original marker position to obtain the noisy positionThe proposed augmentation 203 and corruption 205 pipeline is randomly applied in each epoch, with different probabilities of activating each augmentation or corruption function for each sample inside a mini-batch. By carefully controlling the generation of synthetic MoCap data, it is possible to generate realistic data suitable for training machine learningalgorithms to perform MoCap.Table 3 shows quantitative results using standard metrics of different trained variants of this model. Each row corresponds to a different training dataset , and all tests were conducted on the ACCAD part of the AMASS dataset [Mahmood ct al.. 2019[, winch was never seen during training arid corresponds to an unbiased test set. The training data of the first 3 rows (denoted as MoCap) correspond to data acquired using high-quality MoCap systems, while tiie last row corresponds to data acquired by multi-view marker-less MoCap systems (referred to as Real). The different datasets were equalized in t erms of total minutes of MoCap, and different parts of the AMASS dataset were selected for the MoCap variants, with the necessary temporal down-subsampling to them to result m a total duration of approximately 9 minutes of MoCap, which is the amount of MoCap data available in the Real data. The latter used the GeneBody [Cheng el al.. 2022[ and THuman2.0 [ Fw et al. , 2021} while the former used three different dataset part combinations. MoCapv\ used EK11T, HumanEva. MoSh, arid Soma for training. MoCapv-2 used CNRS and HumanEva, and MoCapoe only used HumanEva for training. The results demonstrate that the model trained on real data performs on par or better than the MoCap data variants in some cases. Advantageously, this validates the disclosed invention’s capacity to deliver a high-quality MoCap system using affordable marker-less MoCap data acquisition means when relying on a standardized low- dimensional representation for the acquision and training data generationBody Model Fitting.
[0101] Once the neural network 103 is trained and its parameters stored in the model pa- rameter database 10 / , they can be loaded on demand to satisfy’ specific capturing conditions.One example is having N models trained with N distinct marker layout's and, depending on the marker layout attached on the captured subject’s suit, the corresponding neural network parameters can be loaded. Then the neural network 103 can be presented with the raw marker position data 102 as estimated by a sensing setup comprising K > 1 optical sen- sors / cameras 101, The neural network 103 will correspondingly predict M landmarks 104. with M corresponding to the training scheme of the respective model.However this is not limited to a single model as the training scheme presented in FIG.2 is generic so as to support cascaded regression One model can first predict one set of landmarks that will be used as input to another model trained to infer another set of landmarks. For example, in ano >ther embodiment a primary neural network can predict de- noised markers and another neural network can solve these predicted de-noised markers to joint positions. In another embodiment, a primary neural network can predict the bodysurface points corresponding to the extruded markers, and then another secondary neural network can solve for the joint positions.
[0102] At the end of this neural network processing cascade comprising 1, ...,C steps, a. set of de-noised and completed known landmarks f.extt R£xoare available. The present invention additionally discloses a method 105 for fitting the body to these estimates m order to obtain the body model’s parameters ([10,T) 106 that reconstruct an explicit representation in the form of an articulated skeleton and mesh surface.
[0103] Estimating the body parameters [p, 0, T) 106 is a non-linear optmiwauori problem 105 with the standard solution being MoSh [Loper et al., 2014\ and its successor MoSh i I [Mahmood at al, 2019\ that also leverages a low-dimensional parametric body model. How- ever, MoSh(d r) also solves for the marker layout which in our case is known a-priori as the neural network 103 estimating fesl104 was loaded from a database 107 containing multiple trained models, each with a pre-defined configuration that was matched to current capture context. Therefore, solving for marker layout is irrelevant for our system, significantly dif- ferentiating it troin Hie MoSh variants that focus on processing archival MoCap data.Furthermore, MoSh assumes the estimates are of high-quality or. otherwise, low signal-to-noise ratios, thereby captured by expensive high-end MoCap systems. The present invention discloses a MoCap system that instead relaxes this assumption to support addi- tional sensing options 101, including affordable consumer-grade alternatives that produce input data of various distinct noise levels. The solution to fitting to noisy data is robust optimization but typical approaches that involve robust kernels / estimators require confident knowledge about the underlying data distribution. This is not easily available in practice, and moreover, it varies with different sensing options.
[0105] More importantly, the present invention presents a system that integrates a data-driven model in the form of a neural network 103 that produces these estimates 104.Therefore, albeit the neural network 103 is trained to de-noise its inputs 102, it introduces another form of information noise, whose distribution is very challenging to model.
[0106] One solution to this is using an adaptive robust optimization variant, with an ex-2019 that adapts to the underlying distribution and interpolat.es / generalizes many known variants by adjusting their shape and scale jointly. Still, optimizing a fitting function with this objective has not been shown to be a practical alternative as this type of objective lias also been used in the context of stochastic optimiza- tion for training neural networks.
[0107] Instead, this invention discloses a likelihood-based formulation to formulate a noise-aware fitting objective that is adaptive and optimizes the Gaussian uncertainty region o e iRz' jointly with the data and prior terms:where L is the numbers of landmarks being optimized. A set of prior terms are also used:both being an 1:2 statistical penalty from the mean shape and pose, given their Gaussian modelling, due to P(JA statistical modelling for the former, and latent- space Gaussianity for the latter. In practise, given an autoencoding generative model prior for the pose 0, solving can by either performed directly on the pose parameters 9. with the prior objective enforced post-encoding the pose parameters, or solving can use the latent space z winch will get decoded into pose parameters 0 and have the prior objective enforced directly on the solved latent code. Essentially, tire pose 0 and latent code z are considered interchangeable by the functioiis in Eq. (1). The values for A.p and XQ can be empirically set but can also be discovered with simulated noise experiments. The data term is formulated as:which simultaneously optimizes the uncertainty region as given a, as well as the landmark fileting; 1:2 term.[QI 08] Staged annealed optimization is performed but with only 2 stages as a marker layout: optimization stage is not needed. The first stage optimizes over J3' , 0*, TE while the second stage fixes (3 and T and optimizes 0*, o". This ensures proper initialization for the uncertainty region optimization, effectively escaping local minima and ill-formed solutions. For the second stage the prior terms are also relaxed to half their values as a good local optima is reached in the first stage where the increased prior terms help in making the objective function more well-behaved {i.c. convex).
[0109] 'table 4 presents results for the disclosed fitting method compared to tire standard approach presented in MoSh(d I-) but without the marker layout optimization step which is not necessary. Artificial noise 205 is added to the marker positions generated by the aforementioned data extraction approach 204 and additionally, these inputs are fed to the neural network to estimate the required landmark positions, adding the neural network estimator’s noise. These different noise sources are combined as indicated in the 'Table, andadvantageously, the disclosed invention’s method manages both noise sources, presenting with clear improvements compared to the standard fitting approach.
[0110] Additional prior constraints p can be added with their corresponding weights A,,. such as joint limit constraints, or priors offered by data-driven per joint distributions as en- coded by a normalizing flow autoencoding generative model \Rezende et al.. S015\. These are indicative priors considered as implementation details, together with the optimization algo- rithm which can include quasi-Newton solving, or Gauss-Newton, or Levenberg Marquardt algorithms for real-time solving. Preferably, to reduce complexity, solving would minimize the joint landmarks in addition to a small subset of surface landmarks for solving the limbs rotations.
[0111] ll) further reduce complexity, a shape calibration process can be used to only perform kinematic optimization when solving, as the shape would be considered fixed and known a-priori lay reading the object / subject’s parameters from a database, or by solving it prior to capturing the performance / motion.
[0112] Extracting the shape can be the result of optimizing Eq. ( 12) over a single, or multiple1, frames. Multiple frames are used to remove the effects of noise on the optimized shape coefficients. In this case, apart from optimizing a single shape coefficient parameter vector, it is also possible to optimize all frames separately, and then aggregate the resulting shape coefficients of each frame into a single parameter vector, typically taking the mean.
[0113] Still, in this case, any errors stemming from properly fitting tire pose would alsoaffect the quality of the shape solution. Thus, a more effective approach is to solve for the shape only by formula ting the problem in a local coordinate system agnostic manner, prac- tically eliminating the dependency on estimating each joint’s rotation. Using the estimated landmarks ('eslan optimization problem formulated, using the following objective:essentially minimizing the difference between the distance of a set of landmark pairs. As long as the set of pairs P a, re selected so that the parent p and child c are rigidly moving landmarks, this objective is pose parameter 0 agnostic. Especially in the case that the landmarks (;es;comprise both markers and joints, high quality shape estimation is possible.Post-capture Refinement.
[0114] The present invention leverages the standarcUzation of tire output body mid motion representation as provided by the parametric body to automatically post-process the body fit data produced by the real-time system. Advantageously, the post-processing 600 accesses a performance recording / segment and uses a motion prior as encoded by a generative model 111 to refine 601 the outputs before storing the newly processed motion 602 (in memory, a hie or to a dataset). The goal of the refinement process is to correct potential erroneous outputs of the real-time system, remove noise and produce smoother output motions. This is achieved by exploiting past and future information, or technically by solving lor a bundle of frames simultaneously" 601.[01151 In more detail, solving for a single frame minimizes the following objective:.P with the optimization performed over a latent code z € KZ' expressing a prior over the pose parameters 0, as explained in the body fitting section. This can be extended to solving lora bundle of frames simultaneously when considering a temporal window(T := 7j:with (he shape (3 omitted from the optimization and considered lined.[0116| Prior works have solved for large temporal windows simultaneously using the pose parameters 9 directly and added temporal constraints, with \ Mahmood et al... 3019\ solving 30 frames simultaneously using a DOT prior after first solving each frame separately, and [Arnab et al. , 2019\ solving over the entire sequence using a joint smoothness objective after having initial estimates from a data-driven model. Other works consider smoothly varying motion priors and add a finite difference latent space smoothness objective [Huang et ah , 2021\ , or operate directly on position trajectories [Hang, et al, 2020\. These however are not efficient as they require extra passes, are not scalable when considering finite differencing, and require the tuning of an additional hyperparameter lor the explici t smoothness objective.[OUT] Tliis invention relies on solving over a scalable temporal window using latent space reconstruction via interpolation, and a set of keyframes, defining the starting and ending points of a smooth pose transition segment. This greatly reduces the optimized parameter space as two latent keyframes at time instances 0 and T can be used:with the prior teiiii imposed only on the keyframes, and tin' data term defined over the entire window ‘T:To solve over the entire window and fit our solution to all available data constraints, we reconstruct the intermediate frames via interpolation:with Cfand t' being spherical and linear interpolation functions tliat map / to a closed unit interval QO, 1]) within ‘I to blend the starting I = 0 and end 1 = T points. Specifically for the pose 0, this process crucially relies on a well-trained generator fr that captures an expressive manifold -M that can be smoothly interpolated to reconstruct correspondingly smooth pose space transitions.
[0118] This implicitly lifts smoothness constraints on the manifold instead of an explicit objective, requiring no hyperparameter tuning. The present technique can additionally be implemented in a streaming manner, considering the first keyframe, t = 0, as an initial condition arid thus, fixed. This way, it is both efficient and scalable as it does not require solving over all parameters in a redundant temporal window, and instead minimizes the next key frame’s t = T latent (aide.
[0119] Each reconstructed pose parameter vector & reconstructs a set of landmarks 604 that are used in the data objective which can be an L2 error or the noise-aware objective of Eq, (15), or any other robust estimator, without loss of generality. The objective is formulated to match the estimated landmarks 104 and produce the next solved keyframe and intermediate frames 602, starting from a fixed previous estimate. A matching generative model is used 605 selected from a database of generative models 110, which can be any generative model that can reconstruct a sequence of poses representing a motion segment.
[0120] A skilled practitioner will realize that any number of keyframes can be used with different multivariate interpolation schemes, as long as they traverse the learned manifold’s latent space appropriately to produce natural poses. Similarly, while a static temporal win- dow chunking process was used, this can be straightforwardly adapted to a dynamic / adaptive temporal window chunking process, with varying levels of chunk overlap and cross chunk pa- rameter initialization. One indicative, and by no means restrictive example would be anoverlap of 3 frames between each chunk, with the next chunk’s initial keyframe state initial- ized, prior to solving, using an extrapolating motion prediction model using the solut ion of tlie overlapping frames as input.Computing Environment.
[0121] FIG. 8 depicts an exemplary computing environment in which various embod- iments of the invention may lie implemented and upon which various embodiments of the invention may be employed. The computing system environment is only one example of a suitable computing environment and is not intended to suggest any limitation as to the scope of use or functionality. Numerous other general purpose or special purpose computing system environments or configurations may be used. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use include, but; are not limited to, personal electronic devices such as smart phones and smart; watches, tablet, computers, personal computers (PCs), server computers, handheld or laptop devices, multi- processor systems, microprocessor-based systems, network PCs, minicomputers, mainframe computers, embedded systems, distributed computing environments, like the cloud, that in- clude any of the above systems or devices, and the like.
[0122] Computer-executable instructions such as program modules executed by a com- puter may be used. Generally, program modules include routines, programs, objects, com- ponents, data structures, etc. that perform particular tasks or implement particular abstract data types. Distributed computing environments may be used where tasks are performed by remote processing devices that are linked through a communications network or other data transmission medium. In a distributed computing environment, program modules arid other data may be located in both local and remote computer storage media including memory storage devices.
[0123] With reference to FIG. 8, an exemplary system for implementing aspects de- scribed herein includes a computing device, such as computing device 1000. In its most basic configuration, computing device 1000 typically includes at least one processing unit1002 and memory 1004, Depending on the exact configuration and type of computing device, memory 1004 may be volatile (such as random access memory (RAM)), non-volatile (such as read-only memory (ROM), flash memory, etc.), or some combination of the two, 'Phis most basic configuration is illustrated in FIG. 8 by dashed line 1006. Computing device 1000 may have additional features / ftmctionality. For example, computing device 1000 may include additional storage (removable and / or non-removable) including, but not limited to. magnetic or optical disks or tape. Such additional storage is illustrated in FIG. 8 by remov- able storage 1008 and non-removable storage 1100 Computing device 1000 as used herein may be either a physical hardware device, a virtual device, or a combinat ion t hereof.[01241 Computing device 1000 typically includes or is provided with a variety of computer -readable media. Computer-readable media can be any available media that can be accessed by computing device 1000 and includes both volatile and non-volatile media, removable and non-removable media. By way of example, and not limitation, computer- readable media may comprise computer storage media and communication media.
[0125] Computer storage media includes volatile and non-volatile, removable and non- removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Memory 1004, remova.ble storage 1008, arid non-removable storage 1010 are all examples of computer storage media. Computer storage media includes, but is not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), electrically erasable programmable read-only memory (EEPROM). Hash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disks (DVD) or oilier optical storage, magnetic cas- settes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired informa lion and which can accessed by com- puting device 1000. Any such computer storage media may be part of computing device 1000.
[0126] Comput ing device 1000 may also contain communications connection) s) 1012 thatallow the device to communicate with other devices. Each such communications connection 1012 is an example of communication media. Communication media typically embodies computer-readable instructions, data structures, program modules or other data in a mod- ulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term "modulated data signal" means a signal that lias (me or more of its characteristics set or changed in such a maimer as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, ra- dio frequency (RE), infrared, and other wireless media. The term computer-readable media as used herein includes both storage media and communication media. Computing device 1000 may also have input device(s) 1014 such as keyboard, mouse, pen, voice input device, touch input device, etc. Output device(s) 1016 such as a display-, speakers, printer, etc. may also be included. All these de vices are generally known and therefore need not be discussed in any detail herein except as provided.
[0127] Notably, computing device 1000 may be one of a plurality of computing devices 1000 inter-connected by a network 1018, as is shown in FIG. 8. As may be appreciated, t he network 1018 may be any appropriate network: each computing device 1000 may be con- nected thereto by way of a connection 1012 in any appropriate manner, and each computing device 1000 may communicate with one or more of the other computing devices 1000 in the network 1018 in any appropriate manner. For example, the network 1018 may be a wired or wireless network within an organization or home or the like, and may include a direct or indirect coupling to an external network such as the internet- or the like.
[0128] It should be understood that the various techniques described herein may be implemented in connection with hardware or software or, where appropriate, with a combination of both. Thus, the methods and apparatus of the present ly disclosed subject matter, or certain aspects or portions thereof, may take the form of program code (i.e., instruc- tions) embodied in tangible media, such as universal serial Ims (USB) flash drives. SecureDigital (SD) memory cards, CD-ROMs, hard drives, or any other machine-readable storage medium wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the presently disclosed subject matter.
[0129] In the case of program code execution on programmable computers, t he computing device generally includes a processor, a storage medium readable by I lie processor (including volatile and non- volatile memory and / or storage elements), at least one input device, and at least one output device. One or more programs may implement or utilize the processes described in connection with the presently disclosed subject matter, e.g., through the use of an application-program interface (API ), reusable controls, or t he like. Such programs may be implemented in a high-level procedural or object-oriented programming language to commmiicate with a computer system. However, the prograni(s ) can tie implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language, and combined with hardware implementations. In an embodiment, the system can be developed using C / C-t I- / 'Python. Although exemplary embodiments may refer to utilizing aspects of the presently disclosed subject matter in the context of one or more stand-alone computer systems, the subject matter is not so limited, but rather may be implemented in connection with any computing environment, such as a network 1018 or a distributed computing environment. Still further, aspects of the presently disclosed subject matter may be implemented in or across a plurality of processing chips or devices, arid storage may similarly be effected across a plurality of devices in a network 1018. Such devices might include personal computers, network servers, and handheld devices, for example.[01301 While certain embodiments are described, these embodiments have been presented by way of example only, and are not intended to limit the scope of tire invention. Indeed, Hie novel systems and methods described herein may be embodied in a variety of other forms. Furthermore, various omissions, substitutions, and changes in the form of the system and methods described herein may be made without departing from the spirit of the invention.TT10 scope of the invention includes any equivalents thereof as would be appreciated by one of ordinary skill in the art .REFERENCESAristidou, Andreas, et al., “Self-similarity analysis for motion capture cleaning." Computer graphics forum. 2018Aniab, Anurag, Carl Docrsch, and Andrew Zisscrman. “'Exploiting temporal context lor 3D human pose estimation in the wild.” IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019.Baran, Ilya, and Jovan Popovic. "Automatic rigging arid animation of 3d characters.” ACM Transactions on graphics. 2007.Barron, Jonathan T. ”A general and adaptive robust loss function.” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019,Branco, Paula, et al,, ’’SMOGN: a pre-processing approach for imbalanced regression.” First international workshop on learning with imbalanced domains: Theory and applications. PMLR. 2017.Chatzitofis, Anargyros, et al., ”Democap: low-cost marker-based motion capture" International Journal of Computer Vision. 2021.Chen, Kang, ct al.. "MoCap-Solvcr: A neural solver for optical motion capture data.” ACM Transactions on Graphics (TOG). 2021.Cheng, Wei, et al. "Generalizable neural performer: Learning robust radiance fields for human novel view synthesis" arXiv preprint arXiv:2204.11798. 2022.Fieraru, Mihai, et al., "Ailit: Automatic 3d human -interpretable feedback models for fitness training.” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021 .Ghorbani, Mima, et al, ’’Sonia: Solving optical marker-based mocap automatically.” Proceedings of the lEEE / CVF International Conference on Computer Vision. 2021.Ghorbani, et al.. "Auto-labelling of markers in optical motion capture by permutation learning"Advances in Computer Graphics: 36tli Computer Graphics International Conference. 2019.Gong, Yu, el al., ’’RankSim: Ranking Similarity Regularization for Deep Imbalance'! Regression.’’ arXiv preprint arXiv:2205.15236. 2022.Hau, Shangchen, et al. "Online optical marker-based hand tracking with deep labels " ACM Transactions on Graphics (TOG ). 2018.Holland. Paul W., et al. ’’Robust regression using iteratively reweighted least-squares" Communications in Statistics-theory and Methods. 197'7.Holden, Daniel. ’’ Robust solving of optical motion capture data by denoising" ACM Transactions on Graphics (TOG) . 2018.Huang, Y ., Bogo, F.. Lassner, C., Kanazawa, A , Gehler, P. V.. Romero. J., ... V: Black, M. J. •'Towards accurate marker-less human shape and pose estimation over time." In international conference on 3D vision. 2017.Huang, Y., Taheri. O., Black. M J., & Tzionas, D. "InterCap: Joint Markcrlcss 3D 'Tracking of Humans and Objects in Interaction " In DAGM German Conference on Pattern Recognition. 2022.Huang, B., Shu, Y., Zhang, T.. & Wang. Y. ’’Dynamic multi-person mesh recovery from uncalibrated multi- view cameras." In International Conference on 3D Vision. 2021.Huang, Buzhen, YYian Shu, Tianshu Zhang, and Yangang W’ang. "Dynamic multi-person mesh recovery from uncalibrated multi-view cameras.” International Conference on 3D Vision. 2021. Ingwersen, C. K., Mikkelstrup, C. M., Jensen, J. N., Hannemose. M. R., X: Dahl, A B."SportsPose-A Dynamic 3D sports pose dataset." Proceedings of the IEEE / GVF Conference on Computer Vision and Pattern Recognition. 2023Jacobson, Alec, Ilya Baran, Jovan Popovic, and Olga Sorkinc. ’’Bounded biharmonic weights for real-time deformation.” ACM Transactions on Graphics. 2011.Kaufmann, Manuel, et al. ”Em-pose: 3d human pose estimation from sparse electromagnetic trackers.” Proceedings of the lEEE / CVF International Conference on Computer Vision. 2021.Kingrna, Dicderik P., et al. ” Auto-encoding variational bayos.” arXiv preprint arXiv:1312.61 14 (2013) . In 2nd International Conference on Learning Representations, (ICLR). 2014.Loper, Matthew, et ah, "MoSh: motion and shape capture from sparse markers." ACMTrausactions on Graphics (TOG). 2014.Loper. Matthew, et al. ”HMPL: A skinned umlti-pcrsoii linear model.” ACM transactions on graphics (TOG ). 2015.Mahmood, Naureen, et al. "AMASS: Archive of motion capture as surface shapes.” Proceedings of the 1EEE / CVP international conference on computer vision. 2019.Menendez. M. L., et al. '’The jensen-shannon divergence." Journal of the Franklin Institute. 1997. Nibali, Aiden, et al. ’’Numerical coordinate regression with convolutional neural networks.” arXiv preprint arXivrlSOl .07372. 2018.Nibali, Aiden, et al. ”3d human pose estimation with 2d marginal heatmaps." I EEE Winter Conference on Applications of Computer Vision (WACV). 2019Osman. Ahmed A A, et ah, "Star: Sparse trained articulated human body regressor" Proceedings of the 10 th European Conference on Computer Vision. 2020Osman. Ahmed AA. et al. ”SUPR: A Sparse Unified Fart-Based Human Representation" Proceedings of the 17th European Conference on Computer Vision. 2022.Pavlakos, Georgios, et al. ’’Expressive body capture: 3d hands, face, and body from a single image." Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019.Pavllo, Dario, et al. "Real-time neural network prediction for handling two-hands mutual occlusions.” Computers & Graphics. 2019.Rernpe, I)., Birdal, T., Hcrtzmann. A .. Yang. J ., Sridhar, S., & Guibas, L. J. (2021 ). Humor: 3d human motion model for robust pose estimation. In Proceedings of the 1EEE / CVE international conference on computer vision (pp. 11488-1 1499).Ren, Jiawei, et al. "Balanced mse for imbalanced visual regression." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022.Rezende, Danilo, et al. ’’Variational inference with normalizing hows." International conference on machine learning, (PMLR). 2015.Ronneberger, Olaf, et al. ”U-net: Convolutional networks for biomedical image segmentation.”18th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAH. 2015.Kong, Yti, et al., Chasing tlic tail in monocular 3d human rcconslructiou with prototype memory.” IEEE Transactions on Image Processing. 2022.Saini. i\L, Huang, C. H. 11, Black. M. J., fc Ahmad, A. ’’SmartMocap: Joint Estimation of Human and Camera Motion Using Uncalibrated RGB Cameras." IEEE Robotics ami Automation Letters. 2023.Schubert, Tobias, et al. "Automatic initialization for skeleton tracking in optical motion capture.” IEEE International Conference on Robotics and Automation (ICRA). 21)15.Silva, Anfbal, et al., ’’Model Optimization in Imbalanced Regression.” Proceedings of the 25th International Conference on Discovery Science. 2022.Steininger, Michael, et al. ’’Density-based weighting for imbalanced regression.” Machine Learning. 2021Sun, Xiao, et al. ’’Integral human pose regression.” Proceedings of the European conference on computer vision (ECCV). 2018.Torgo, Luis, et al. ’’Resampling strategies for regression.” Expert Systems. 2015.Toigo, Luis, et al., ’’Utility-based regression.” PKDD. 2007.Torgo, Luis, et al. "Smote for regression.” 16th Portuguese Conference on Artificial Intelligence. 2013.Varol, Gul, et ah ’’Learning from synthetic humans.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017.Xu, Hongyi, et al. ”Ghum & ghuml: Generative 3d human shape and articulated pose models.” Proceedings of the 1EEE / CVF Conference on Computer Vision and Pattern Recognition. 2020. Xue, Han. et al. ' OMAD: Object Model with Articulated Deformations for Pose Estimation and Retrieval." Proceedings of the British Machine Vision Conference. 2021.Van, Hannan, et al. "Ultrapose: Synthesizing dense pose with 1 billion points by human-body decoupling 3d model.” Proceedings of the IEEE / C V’F International Conference on ComputerVision. 2021.Yang, Yuzhe, et al. ’’Delving into deep imbalanced regression." International Conference onMachine Learning, PMLR. 2021.Ye, Hang, et al. "Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection.” 17th European Conference on Computer Vision. (ECCV). 2022.Ye, V., Pavlakos. G.. Malik, J., & Kanazawa, A, Decoupling human and camera motion from videos in the wild. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2023.Yu, Tao, et al. ”Function4d: Real-time human volumetric capture from very sparse consumer rgbd sensors.” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2021 .Zanfir, Andrei, Elisabeta Marinoiu, and Cristian Smiuchisescu. ’’Monocular 3d pose and shape estimation of multiple people in natural scenes-the importance of multiple scene constraints ” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018.Zhang, Feng, et al. ’’Distribution-aware coordinate representation lor human pose estimation.’' Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020. Zhang, Richard. ’’Making convolutional networks shift-invariant again.” International conference on machine learning (PMLR). 2019.Zuffi, Silvia, et al. ”3D menagerie: Modeling the 3D shape and pose of animals.” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPRj. 2017.Hang, Jianwei, et al. ’’Abnormal value processing method for three-dimensional trajectory data and optical motion capture method", WO2020133447A1. WIPO. 2020.Michael J. Black, Xirna Ghorbani, "Method and system for labelling motion-captured points”, EP4163882A2, European Patent Office. 2023.
Claims
CLAIMS1 A computer- implemented machine learning based system (100) operable to perform motion capture using body-worn markers, the system comprising:(a) An optical sensor setup (101) that provides raw unstructured physical marker- position estimates ( 102) said markers attached to the body of either a human or auimal subject or an object;(b) A parametric body model database (108);(c) A neural network (103), selected from a neural network database (107), that uses the raw unstructured physical marker position estimates to estimate a sei of body related landmarks (104), simultaneously performing input measurement denoising, landmark labeling, solving and completion, whereby the selection is associated to: i the parametric body model used to train the neural network so as to match the captured subject / object, and in the configuration of the markers on the body so as to match the neural network’s training input.(d) A shape calibration subsystem (109) that uses the estimated set of subject / object related landmarks to solve for the shape of the subject / objecfs outer 3D surface and inner kinematic skeleton, using a matching parametric body blendshape representation, selected from the parametric body model database:(e) An optimization subsystem (105) that uses the calibrated shape and a pose and / or motion generative model, selected from a generative model database (110), to refine the estimated landmark representation, by fitting the calibrated landmark positions to the estimated ones, producing an improved shape and motion representation (106):(f) A database (111) that stores the captured motion, comprising optimized parametric model parameters, including the shape and pose representations, said pose representations comprising joint positions and / or rotations, whereby Idle captured motion together with the shape and pose representations can reconstruct the subject / object’s 4D motion. The system of claim 1. wherein the parametric body model comprises an artist-made set of blendshapes, rigged manually and artist-painted skinning weights. The system of claim 1. wherein the parametric body model relies on statistical modeling, either deterministic or stochastic, of a set of 3D human body scans. These ca.n include additional blendshapes and joint regressors. The system of claim 1, wherein the neural network estimated landmarks comprise a marker set. The marker set can include the actual markers placed on the subject / object but is not limited to it. Additional markers can be estimated depending on tire training configuration. The system of claims 1, 2, 3. wherein the neural network estimated landmarks comprise the human’s body or animal’s body or object’s joints. The inner joints are solved by the model, producing unobserved landmarks. The system of claims 1. 2, 3, wherein the neural network estimated landmarks comprise surface points that correspond to the positions where markers have beenplaced at. These estimated surface positions are not measurable by the optical system delivering the input. The system of claims 1.
2.
3. wherein the neural network estimated landmarks comprise markers and joints. The neural network solves both for the input markers as well as the human’s body or animal's body or object’s joint positions. The system of claims 1. 2, 3, wherein the neural network estimated landmarks comprise surface points and joints. The neural network solves both for the marker placed surface positions as well as the inner joints. The system of claim 1, with any permutation of body models as defined in claims 2 and 3 and the training and estimated landmark combinations defined in claims 4, 5, 6, 7. or 8. The system of claims 1 to 9, wherein the shape calibration subsystem is used to solve for the pose and shape simultaneously using the estimated landmarks. The system of claims 1 to 9. wherein the shape calibration subsystem uses only pairwise relationships of the estimated landmarks to solve for the shape blendshape coefficients only. I'he system of claims 1 to 11. wherein the optimization (105) optimizes the uncertainty region of the neural network estimated landmarks in addition to the parametric body’s parameters. A method where a neural network (103) is trained to perform balanced regression (200) using an imbalanced training data distribution, the method comprising the singes of:(a) During a first stage, training an autoencoding generative model, comprising an encoder and a decoder, using the imbalanced data wherein:
1. The encoder is used to extract a latent code representation of the samples. ii. The decoder is used to output a reconstruction of the input representation of the samples using the encoded latent code representation,(b) During a second stage, generating synthetic samples using the decoder of the autoencoding generative model to supplement the available data, and assigning importance values to the training data using said autoencoding generative model; i. Whereby, the second stage comprises training a neural network performing regression, using data samples and / or importance values, ii. Whereby the data samples are sourced either from the imbalanced training data, or from tire synthetic samples generated using the decoder of tlic autoencoding generative model or from a combination of both, iii. And whereby an importance value is assigned to each dat a sample with a. function using the autoencoding generative model's inputs and outputs, The method of claim 13, wherein the autoencoding generative model is a non-mvertible variational auto-encoder comprising fully connected, convolutional, or a combination of fully-connected and convolutional layers, or a combination of such invertible variational auto-encoders. The method of claim 13, wherein the autoencoding generative model is an invertible probabilistic model or a chain of invertible probabilistic models that operate differentiable and invertible transformations, with a differentiable inverse. The method according to any of claims 13. 14, 15, wherein the decoder generates synthetic data samples using as input random latent codes drawn from a uniform, normal or any other distribution.The method according to any of claims 13. 14, 15. wherein the decoder generates synthetic data samples rising as input latent codes drawn from a miilbrm, normal or any- other distribution around a number of predefined anchor latent codes used as the distributions' centers with varying scales. The method according to claim 17, wherein a number of synthetic data samples are used to reconstruct new synthetic data samples using multivariate interpolation with random interpolation parameters. The method according to claim 18, wherein two synthetic data samples are used to reconstruct a new' synthetic data sample using spherical linear interpolation with a random interpolation parameter sampled uniformly from a predefined range lower bound by 0 and upper bound by 1. The method according to any of claims 17. 18, 19, wherein the predefined latent code anchors are any of:(a) manually created,(b) manually selected from any available data, selected after extracting a statistical representation of the data using each data sample’s assigned importance value, wherein the selection takes place: i. automatically, or ii. manually(cj or any combination thereof. The method according to any of claims 13 to 20. wherein the importance value is assigned using any7of:(aj the input and reconstruct ed output of the autoencoding generative model, or,(b ) the latent code representation of the autoencoding generative model, or,(c) any combination thereof. The method according to any of claims 13 to 21. wherein the importance value is assigned to the real data samples from tire imbalanced training data, but is fixed to a set value for the generated synthetic data samples. The method according to any of claims 13 to 21, wherein the importance value for all samples is fixed to unity. The method according to any of claims 13 to 23. wherein tire synthetic data samples are not used when training the model, The method according to any of claims 13 to 24. wherein the importance value is adapted as a function of the training step. I’lie method according to any of claims 13 to 25. wherein the importance value is fixed to unity up to a predefined training step, and after that training step is assigned to the data samples. The method according to any of claims 13 to 26. wherein the generated synthetic data samples are used when training the model after a predefined training step. A system operable to train a neural network model (103) to perform balanced regression using training data, comprising data samples, exhibiting an imbalanced training data distribution, the system comprising:(a) Means, operable during a first stage, said means adapted to train an autoencoding generative model (207). said autoencoding generative model comprising an encoder and a decoder, using the imbalanced data wherein: i. The encoder is used to extract a. latent code representation of the samples. ii. The decoder is used to output a reconstruct ion (401) of the input representation of the samples using the encoded latent code representation.(b) Means, operable during a second stage, said means adapted to generate synthetic samples (202) using the decoder of the autoencoding generative model to supplement the training data, said means adapted to assign importance values (208) to the training data using said autoencoding generative model: i. Whereby, during said second stage, said means is adapted to train a neural network performing regression, using data samples and / or importance values, ii. Whereby during said second stage, said means is adapted to source the data samples either from the imbalanced training data, or from tire synthetic samples generated using the decoder of the autoencoding generative model or from a combination of both, iii. And whereby, during said second stage, said means is adapted to assign an importance value to each data sample with a function using the autoencoding generative model’s inputs and outputs. A system according to claim 28, wherein the autoencoding generative model is a non-invertible variational auto-encoder comprising fully connected, convolutional, or a combination of fully-connected and convolutional layers, or a cornbination of such non invertible variational auto-encoders. A system according to claim 28, wherein tlie autoencoding generative model is an invertible probabilistic model or a chain of invertible probabilistic models that operate differentiable and invertible transformations, with a differentiable inverse. A system according to any of claims 28,29, or 30. wherein the decoder comprises means for generating synthetic data samples using as input random latent codes drawn from a uniform, normal or any other distribution. A system according to any of claims 28, 29, or 30, wherein the decoder comprisesmeans for generating synthetic data samples using as input latent codes drawn from a. imiArrm, normal or any other distribution around a number of predefined anchor latent codes used as the distributions' centers with varying scales. A system according to claim 32, wherein a number of synthetic data samples are used to reconstruct new synthetic data samples using multivariate interpolation with random interpola bion parameters. A system according to claim 33, wherein two synthetic data samples are used to reconstruct a new synthetic data sample using spherical linear interpolation with a random interpolation parameter sampled uniformly from a predefined range lower bound by 0 and upper bound by f . A system according to any of claims 32. 33, or 34, wherein the predefined anchor latent codes are any of:(a) manually created,(bj manually selected from any available data,(cl selected after extracting a statistical representation of Hie data using each data sample’s assigned importance value, wherein the selection takes place: i. automatically, or ii. manually(d) or any combination of a or b or c thereof. A system according to any of claims 28 to 35, wherein the importance value assigning means, assigns values using any of:(a) the input and reconstructed output of the autoencoding generative model, or,(b) the latent code representation of the autoencoding generative model, or,(c) any combination thereof. A system according to any of claims 28 to 35. wherein the niiportance valueassigning means, assigns values to the real data samples from the imbalanced training data, wherein said important values are fixed to a set value for the generated synthetic da ta samples. A system according to any of claims 28 to 35, wherein the importance value for all samples is fixed to unity. A system according to any of claims 28 to 37, wherein Hie synthetic data samples are not used when training the model. A system according to any of claims 28 to 38, wherein the importance value is adapted as a. function of the t raining step. A system according to any of claims 28 to 39. wherein the importance value is fixed to unity up to a predefined training step, and after that training step is assigned to tire data samples. A system according to any of claims 28 to 40, wherein the generated synthetic clatri samples are used when training the model after a predefined training step. A system according to any of claims 28 to 42, wherein the trained neural network model is added to the neural network database of any of t-lie systems in claims 1 to 12. A method to process a captured motion, using a generative neural network operable to generate a motion segment using a number of latent code vectors, whereby the captured motion is represented as successive pose frames, said pose frames comprising landmark positions, and / or rotations, the method comprising the stages of:(a) Splitting the captured motion into temporal window chunks: i. each chunk comprising a tune continuous set of pose frames; ii. whereby the chunks can overlap;(b) For each temporal window chunk, initializing a set of latent codes used by the generative neural network to generate a motion segment matching the corresponding temporal window chunk;(cj For each temporal window chunk, optimizing the set of initialized latent codes to better match the pose frames within each temporal window chunk by minimizing a distance between the motion segment generated using (lie generative neural network and the observations comprising the pose frames, said pose frames comprised m the split temporal window chunk: The method according to claim 44, wherein the generative neural network generates a motion segment using a. single latent code. The method according to claim 44, wherein the generative neural network is an autoregressive model that generates a motion segment using a number of latent codes. 'fhe method according to claim 44. wherein the generative neural network generates a motion segment using a number of successive latent codes. The method according to claim 47, wherein additional distance constraints are imposed between the optimized latent codes to ensure their smooth temporal succession. The method according to claim 44, wherein the generative neural network generates a set of pose keyframes using a number of latent codes and then reconstructs the motion segment- using multivariate interpolation between the pose key frames.The method according to claim 49, wherein two pose keyframes are used, the first being the first temporal window frame ami the second being the last temporal window frame within the temporal window chunk, and the motion segment is reconstructed via spherical linear interpolation between the two generated keyframes. The method according to claims 44 to 50, wherein lire generative neural network was trained with latent code distribution constraints, arid such constraints are also used on the optimized latent codes during the optimization process. The method according to claims 44 to 51, wherein the generative neural network is an autoencoder model, comprising an encoder and a decoder, and whose encoder is used to initialize the set of latent codes. The method according to claim 50. wherein a sliding window optimization process is used comprising the following steps:(a) the temporal window chunks are split to overlap by a single frame, whereby the last of the previous chunk coincides with the first of the current chunk, effectively making each pair of successive chunks overlap by a single frame:(b) the first temporal window chunk’s first, keyframe is initialized ami not optimized;(c) the first temporal window chunk’s second keyframe is optimized so that the motion segment matches the temporal window chunk’s pose frames;(df for each temporal window chunk following f lic first, the following steps are repeated until the entire captured motion, comprising all temporal window chunks, is processed: i. the previous temporal window chunk's second keyframe is used as the first keyframe of the current temporal window chunk and is not: optimized, ii. the second keyframe of the current: temporal window chunk is optimized sothat the motion segment matches the temporal window chunk's pose frames. -4 The method according to claim 53, wherein the generative neural network was trained with latent code distribution constraints, and such constraints are also used on the optimized latent codes during the optimization process. 5 The method according to claims 53 or 54, wherein the generative neural network is an autoencoder model, comprising an encoder and a decoder, and whose encoder is used to initialize the set of latent codes. 6 The method according to any of claims 44 to 55. wherein the optimization process additionally optimizes for the uncertainty region of the landmark estimates. A system operable to process a captured motion, using a generative neural network operable to generate a motion segment using a number of latent code vectors, whereby the captured motion is represented as successive pose frames, said pose frames comprising landmark positions, and / or rotations, the system comprising the following means:(a) Means, operable during a first stage, said means adapted to split the captured motion into temporal window chunks: i. each chunk comprising a time continuous set of pose frames; ii. whereby' the chunks can overlap;(b) Means, operable for each temporal window chunk, said means adapted to initialize a set of latent codes used by the generative neural network to generate a motion segment matching the corresponding temporal window chunk:(c) Means, operable for each temporal window chunk, said means operable to optimize the set of initialized latent codes to match the pose frames within eachtemporal window chunk by minimizing a distance between the motion segment generated using the generative neural network and the? observations comprising the pose frames, said pose frames comprising the split temporal window chunk; A system according to claim 57, wherein the generative neural network comprises means for generating a motion segment using a single la rent code. A system according to claim 57, wherein the generative neural network is an autoregressive model tha t comprises means for generating a motion segment using a number of latent (-odes. A system according to claim 57, wherein the generative neural network comprises means for generating a motion segment using a number of successive latent; codes. A system according to claim 60, wherein additional distance constraints are imposed between the optimized latent codes to ensure their temporal succession. A system according to claim 57, wherein the generative neural network comprises means for genera ting a set art pose keyframes using a number of la t ent codes and then reconstructs tire motion segment using multivariate interpolation between the pose keyframes. A system according to claim 62. wherein two pose keyframes are used, the first being the first temporal window frame and the second being the last; temporal window frame within the temporal window chunk, and the motion segment is reconstructed via spherical linear interpolation between the two generated keyframes. A system according to claims 57 to 63, wherein the generative neural network was trained with latent code distribution constraints, and such constraints are also used on the optimized latent codes during the optimization process.A system according to claims 57 to 64, wherein the generative neural network is an autoencoder model, comprising an encoder and a decoder, and whose encoder comprises means to initialize the set of latent codes, A system according to claim 57 to 65, wherein a sliding window optimization process is used, comprising the following means:(a) means operable to split the temporal window chunks such that they overlap by a single frame, whereby the last of the previous chunk coincides wi th the first of the current chunk, effectively making each pair of successive chunks overlap by a single frame;(b) means operable to initialize but not optimize the first temporal window chunk's first keyframe;(c) means operable optimize the first temporal window chunk's second keyframe so that the motion segment matches the temporal window ('hunk’s pose frames;(d) means operable to repeat for each temporal window chunk following the first, until the entire captured motion, comprising all temporal window chunks, is processed: i. the previous temporal window chunk's second keyframe is used as the first keyframe of the current temporal window chunk and is not optimized, ii. the second keyframe of the current temporal window chunk is optimized so that the motion segment matches the temporal window chunk's pose frames. The system according to any of claims 57 to 65, wherein the optimization process additionally comprises means to optimize for idle uncertainty region of the landmark estimates.