System and method for personalizing the operation of avatar movements
Patent Information
- Application Number
- PCT/IB2025/052346
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2025-03-04
- Publication Date
- 2025-10-02
AI Technical Summary
Existing VR-avatar animation systems face challenges in accurately recognizing user motion gestures, particularly when using hand-focused motion capture devices, leading to incomplete avatar control and limited joint capture, and struggle with real-time processing of video streams, necessitating complex equipment or post-editing.
A method involving motion embedding vectors and clustering to personalize avatar movements, using a skeleton encoder to transform user movements into lower-dimensional vectors, associating these with commands, and utilizing K-Nearest Neighbors classification for real-time avatar control.
Enables accurate and real-time avatar control with reduced equipment complexity, allowing individual expressiveness and improved gesture recognition across various motion capture devices.
Smart Images

Figure IB2025052346_02102025_PF_FP_ABST
Abstract
Description
[0001] SYSTEM AND METHOD FOR PERSONALIZING THE OPERATION OF AVATAR
[0002] MOVEMENTS
[0003] FIELD OF THE INVENTION
[0004] The present disclosure relates to virtual reality in general, and to avatar operation, in particular.
[0005] BACKGROUND OF THE INVENTION
[0006] VR-avatar animation is used for a variety of industries, such as social interactions, computing gaming, education, advertising, and other kinds of media production. Avatar animations are built and then triggered via computer algorithms. There are various devices for capturing user’s movements (hereinafter referred to as motion capture device): hand controllers, sensors and transducers, glasses and helmets, video cameras and other equipment for transmitting information about user movements to a computer motion recognition system (hereinafter referred to as motion capture program). At the output of the motion capture program, user’s skeleton tracing is recognized and formed from the captured user's movements, which is a time sequence of coordinates of the set of points of the recognized skeleton (hereinafter referred to as skeleton tracing). Motion capture complex (hereinafter referred to as motion capture complex) is a software-hardware complex (a combination of motion capture program and motion capture device) captures and converts user movements into set of frames, named as skeleton tracing, which consists of the recognized coordinates of the joints of user's skeleton.
[0007] SUMMARY OF THE INVENTION
[0008] The term computing device refers herein to a device that includes a processing unit. Examples of such device are a personal computer, a laptop, a server, a wearable device, a tablet, a cellular device, a cloud device and IOT (internet of things) device.
[0009] The term command for operating movements of an avatar refers herein to a command from a user to switch the avatar animation to a correspondent avatar movement during operation. Such movements can be for example, running, jumping and etc.
[0010] The term user simple movement refers herein to one or more user’s motion gestures that are used for personalizing the avatar movements. Such user’s motion gestures may include full- body gestures, half-body (torso) gestures, fingers gestures, eyes gestures, face mimic, other kinds of physical gestures provided by the user. The user simple movement may also include a combination of user’s motion gestures with user’s interacting with any technical devices. Such interactions include pressing on controller buttons, keyboard keys, as controlling by joystick or by voice commands.
[0011] The user simple movements are typically used for training the system to personalize the commands for operating movements of an avatar.
[0012] The term user complex movement refers herein to a movement that the user performs for operating the avatar in real time. To provide this operating, the user complex movement reflects a user simple movement as well as a movement for avatar navigating in its virtual space and movement for avatar interacting with other avatars or with some virtual objects or with some virtual characters like NPCs (non-playable characters) and CPCs (co-playable characters).
[0013] Examples of virtual space in which the avatar navigates are Virtual reality, Augmented Reality (AR), Merge Reality (MR) and other kinds of realities.
[0014] Such complex movements may include full body movements, half-body (torso) movements, fingers movements, eyes movements, face mimic and any other physical movements and activities performed by the user including user’s interaction with any technical devices such as pressing on controller buttons, keyboard keys, as controlling by joystick or by voice commands.
[0015] The term user’s movement refers herein to a user simple movement or to a user complex movement.
[0016] The term motion capture device refers herein to a device that captures user’s movements in form of initial motion data such as a video stream or sensors data.
[0017] The term initial motion data item refers herein to data captured by a motion capture device during one time frame such as one frame of a video stream.
[0018] The initial motion data items can be transformed frame-to- frame by a motion capture program into a format of skeleton-based motion data items.
[0019] The term skeleton-based motion data item refers herein to initial motion data item that is linked to body skeleton. The skeleton-based motion data item includes positions of numerous user’s body joints such as head, shoulders, legs, hands, knees and feet data items. The skeleton- based motion data item includes static components or dynamic components or a combination thereof. The static body components may include positions of body joints, armatures (body bones), angles between armatures, body planes and their normal vectors. The dynamic body components may include time-derivatives of static positions including velocity (speed), acceleration, jerk and the like.
[0020] The term motion data item refers herein to initial data motion item or to skeleton based motion item.
[0021] The term skeleton tracing refers herein to a data set that includes a time sequence of skeleton-based motion data items that is generated by motion capture program during a session when the user performs his user’s movements to train the system or to operate an avatar in real time.
[0022] The term time segment of skeleton tracing refers herein to a part of skeleton tracing that is limited by a certain first time-frame and by a certain second time frame. Usually in state-of-art technologies of motion capture, skeleton tracing is divided to a continuous sequence of time segments with the same quantity of skeleton-based motion data items.
[0023] The term vector transformation refers herein to the mathematical operations that convert a vector from one coordinate system to another. There are several types of vector transformations, including linear transformations, affine transformations, and non-linear transformations. Linear transformations involve operations that preserve vector addition and scalar multiplication, while affine transformations include translations in addition to linear transformations. Non-linear transformations, on the other hand, can distort the vector space in more complex ways, making them useful in advanced data modelling scenarios.
[0024] The term motion embedding refers herein to a vector transformation of a time sequence of motion data items such that a time sequence of motion data items is compressed into a vector with a lower coordinate dimension.
[0025] The term skeleton embedding refers herein to a vector transformation of a time segment of skeleton tracing such that a time segment of skeleton tracing is compressed into a vector with a lower coordinate dimension.
[0026] The term skeleton encoder refers herein to a type of machine learning system that is configured to perform a vector transformation of a time segment of skeleton tracing into skeleton embedding. Such a skeleton encoder can be based on certain types of neural networks that are designed to perform the vector transformation process.
[0027] The term gesture recognition system refers herein to a system that is based on recognizing various user’s motion gestures from captured user’s movements to operate of avatar movements.
[0028] One technical problem disclosed by the present disclosure is how to ease the user’s operation of avatar movements while maintaining accurate recognition of the user’s motion gestures. In one example of the state-of-the-art solutions such as gesture recognition system, when using a hand focused motion capture device, there are difficulties in providing the avatar control via user movements based on sequence of user’s motion gestures; for example, when operating two hand controllers and helmet, the movements of the avatar's body-joints, excluding the hands and head, are compensated by a combination of inverse kinematics methods and finger pressing and head) not only do not participate in the formation of the control gesture but also interfere with the recognition of hands tracking. As a result, many hand controller-only applications are focused to control of only half-torso avatars.
[0029] In one other example of the state-of-the-art solutions such as finger gesticulation, the user is required to move the fingers only and preferably in the most static position of his body. As a result, such finger-control turned may be effective for clicking on virtual menus and is not effective in full-bodied VR- avatars animations.
[0030] In another state-of-the-art example, the motion capture device has a limited number of captured joints. For example, the motion capture device provides the implementation of the gesture recognition system but does not provide accurate recognition of the captured user movements when generating skeleton tracing and, at the same time, has rather limited number of captured joints (from 5 to 9). This, in turn, leads to the difficulty of recognizing of user’s motion gestures from such a skeleton tracing and the difficulty of reaching a correct switching of types of avatar movements since standard gesture recognition systems are usually trained on the rather narrow range of human models which perform set of universal body gestures for each kind of movements.
[0031] One other technical problem disclosed by the present invention is how to process the video stream of VR-avatar animation in real time. In one example of a state-of-the-art single video camera transmission technologies, the processing of the video stream of VR-avatar animation is performed after the completion of the entire video shooting, and not in real time. In another example, when using a simple video camera, the lack of high-quality motion capture from a simple video camera is compensated via post-editing of input video to improve captured skeleton tracing for future VR-avatar animation. Such a solution causes a time lag between user’s movements and avatar’s movements and cannot be implemented for real time movements transfer. In another example combined with Avatar Grow Legs technology from Meta Reality Labs, when using a helmet Meta Quest, the lack of capturing low-body joints is compensated via implementation of neural network diffusion model that synthesis avatar legs movements by captured upper-body joints tracking and requires a high-performance computer to provide in real time the execution of such computationally intensive process.
[0032] One other technical problem is how to achieve individual expressiveness for the user when controlling an avatar, even on simple motion capture device while avoiding the usage of complex equipment such as the ROKOKO suit or complex equipment room that requires a plurality of cameras placed around the Kinect studio, combining both RGB cameras and built-in infrared projectors with detectors.
[0033] One technical solution is receiving a user selection of a command for operating movements of an avatar; receiving one or more motion data items, associated with a motion capture of the user while the user performing a simple movement to be associated with the command; transforming the one or more motion data items into a motion embedding vector; and generating a cluster, the generating comprises associating, in a data repository, the motion-embedding vector with the command; the cluster being for personalizing the operation of the avatar by the user. One other technical solution is receiving one or more second motion data items, associated with a second motion capture of the user while the user performing a variation of the simple movement to be associated with the command; transforming the one or more second motion data items into a second motion embedding vector; and associating, in the data repository, the second motion embedding vector with the cluster.
[0034] In some aspects of the present invention relates to a non - transitory computer - readable medium comprising instructions which when executed by at least one processor causes the processor to perform the method of the present invention.
[0035] One exemplary embodiment of the disclosed subject matter is a method, the method comprises: receiving a user selection of a command for operating movements of an avatar; receiving one or more motion data items, associated with a motion capture of the user while the user performing a simple movement to be associated with the command; transforming the one or more motion data items into a motion embedding vector; and generating a cluster, the generating comprises associating, in a data repository, the motion-embedding vector with the command; the cluster being for personalizing the operation of the avatar by the user. According to some embodiments the simple movement comprises one member selected from a group consisting of: full body gestures, half-body (torso) gestures, fingers gestures, eyes gestures, face mimic and user’s interaction with a technical device. According to some embodiments, the user interaction includes one member selected from a group consisting of: pressing on controller buttons, pressing on a keyboard key, operating a joystick or performing a voice command.
[0036] According to some embodiments the motion data items being a vector of time segment of skeleton tracing and wherein the motion embedding vector being a skeleton embedding vector. According to some embodiments the method further comprising: receiving one or more second motion data items, associated with a second motion capture of the user while the user performing a variation of the simple movement to be associated with the command; transforming the one or more second motion data items into a second motion embedding vector; and associating, in the data repository, the second motion embedding vector with the cluster. According to some embodiments the method further comprising, utilizing the cluster for classifying an operation of the user while playing with the avatar as operating the command. According to some embodiments the method further comprises the second computing device being one member selected from a group consisting of Virtual Reality (VR) program, Augmented Reality (AR) program and Merge Reality (MR) program .
[0037] One other exemplary embodiment is a method, comprising: receiving a vector of time segment of skeleton tracing, associated with a motion capture of an interaction of a user with a computing device for operating of an avatar;by skeleton encoder, transforming the vector of time segment of skeleton tracing into an input skeleton embedding vector; comparing the input skeleton embedding vector to clusters of skeleton embedding vectors; the clusters being generated during a training period; the training period associating variation of performing a plurality of commands in a plurality of variations of simple movements performed by the user; selecting from a data repository a selected cluster from the clusters; the selected cluster being identified by the comparing process as the nearest to the input skeleton embedding vector; retrieving from the data repository a selected command associated with the selected cluster; and transmitting the selected command to a second computing device for operating the avatar in accordance with the selected command.
[0038] According to some embodiments the method further comprising, the second computing device being one member selected from a group consisting of Virtual Reality (VR) program, Augmented Reality (AR) program and Merge Reality (MR) program.
[0039] According to some embodiments the method further comprising the comparing comprises utilizing K-Nearest Neighbors classification algorithm. According to some embodiments the method further comprises the comparing comprises comparing a centroid of the clusters to the input skeleton embedding. According to some embodiments the method further comprising the comparing being in combination with comparing magnitude of the dispersion of the input skeleton embedding to a distribution of the cluster. According to some embodiments the method further comprising the comparing comprises calculating a closest cluster to a sequence of other input skeleton embedding identified from one or more time segments of skeleton tracing successive to a time segment of the input skeleton embedding. According to some embodiments the method further comprising further comprising generating a log file during the operating the avatar; identifying suspicious skeleton tracing time segments combined with correctness of classifying of a command for operating movements of an avatar, presenting the identifies-suspicious skeleton tracing time segments to the user and updating a data repository associated with the training process in accordance with a response of the user combined with a value of the command for operating movements of an avatar. According to some embodiments the method further comprising the variations comprise one member selected from a group consisting of: a plurality of angles of capturing, a plurality of velocities of performing the movement, a plurality of locations of performing the movement and a plurality of field of views.
[0040] THE BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0041] The present disclosed subject matter will be understood and appreciated more fully from the following detailed description taken in conjunction with the drawings in
[0042] 5 which corresponding or like numerals or characters indicate corresponding or like components. Unless indicated otherwise, the drawings provide exemplary embodiments or aspects of the disclosure and do not limit the scope of the disclosure. In the drawings:
[0043] Fig. 1 shows a block diagram of an environment for personalizing the operation of avatar movements, in accordance with some exemplary embodiments of the subject matter;
[0044] Fig. 2 shows a block diagram of a system for personalizing the operation of avatar movements, in accordance with some exemplary embodiments of the disclosed subject matter;
[0045] Fig. 3 shows a flowchart diagram of a method for personalizing the operation of avatar movements, in accordance with some exemplary embodiments of the disclosed subject matter;
[0046] Fig. 4 shows a flowchart diagram of a method training a system for personalizing the operation of avatar movements, in accordance with some exemplary embodiments of the disclosed subject matter; and
[0047] Fig. 5 shows a flowchart diagram of a method for fine tuning the operation of avatar movements, in accordance with some exemplary embodiments of the disclosed subject matter.
[0048] DETAILED DESCRIPTION
[0049] Fig. 1 shows a block diagram of an environment for personalizing the operation of avatar movements, in accordance with some exemplary embodiments of the subject matter. Environment 100 includes a computing device 101, motion capture device 102, a VR display 103 and a user 104.
[0050] The computing device 101 is configured to train the system according to the user movement and to control the operation of the avatar.
[0051] The computing device 101 includes User’s Personalized Gesture Recognition system 1011, motion capture program 1012 and VR program 1013.
[0052] It should be noted that modules 1011, 1012 and 1013 may also be placed on different computing devices including the motion capture devices 102 and the VR display 103.
[0053] The User’s Personalized Gesture Recognition system 1011 is configured for personalizing the operation of avatar movements. The User’s Personalized Gesture Recognition system 1011 receives user’s skeleton tracing from the motion capture program 1012, recognizes user’s motion gestures and transfers corresponding commands for operating movements of an avatar to the VR program 1013. The User’s Personalized Gesture Recognition system 1011 interacts with the user 104 via the computing devices 101 or / and via the motion capture device 102. The operation of the User’s Personalized Gesture Recognition system 1011 is explained in greater detail in figures 2, 3, 4 and 5.
[0054] The motion capture program 1012 is configured for transforming motion data combined with user’s movements into user’s skeleton tracing. The motion capture program 1012 receives the initial motion data items from motion capture devices 102 and transforms the initial motion data items into skeleton-based motion data items. The motion capture program 1012 connects the skeleton-based motion data items into user’s skeleton tracing. The motion capture program 1012 transfers the user’s skeleton tracing to the User’s Personalized Gesture Recognition system 1011 and to the VR program 1013. Examples of motion capture programs are OptiTrack Motive, Captury Live, MVN Animate, Meta Quest SDK and Sony Mocopi SDK.
[0055] The VR (virtual reality) program 1013 is configured for displaying Virtual Reality including avatar movements on the VR display 103. The VR (virtual reality) program 1013 receives user’s skeleton tracing from the motion capture program 1012 and command for operating movements of an avatar from the User’s Personalized Gesture Recognition system 1011. Examples of VR programs are VR applications such as Second Life or Altspace VR.
[0056] The VR display 103 can be also Augmented Reality (AR) display, Merge Reality (MR) display or a display of other kinds of realities where avatar displaying can be implemented. Correspondingly, the VR program 1013 can be also AR or MR program or to another program associated with other kinds of realities where avatar displaying can be implemented.
[0057] The motion capture device 102 is configured for capturing of user’s movements performed by the user 104 in form of initial motion data items and for transferring the initial motion data items to the motion capture program 1012. Examples of such devices are MOCAP set with helm and 2 hand controllers and 2-4 sensors like HTC Vive or MOCAP device based on webcam.
[0058] The VR display 103 is configured for displaying Virtual Reality to the user 104. Usually, the display is embedded in VR-helmet or in VR glasses.
[0059] The user 104 operates his avatar movements in Virtual Reality by performing his movements that are captured by the motion capture device 102. Additionally, the user 104 can directly interact with the User’s Personalized Gesture Recognition system 1011 via one or more computing devices 101. Finally, the user 104 sees his avatar movements as a part of Virtual Reality on the VR display 103.
[0060] The operation by the user is explained in greater detail in figures 3,4 and 5.
[0061] Fig. 2 shows a block diagram of t h e system for personalizing the operation of avatar movements, in accordance with some exemplary embodiments of the disclosed subject matter. User’s Personalized Gesture Recognition system 1011 is typically downloaded to a computing device. User’s Personalized Gesture Recognition system includes user front-end interface 201, data repository 202, skeleton encoder 203 and computation module 204. User’s Personalized Gesture Recognition system 1011 interacts with the user and with external programs including the motion capture program (not shown in this figure) and the VR program (not shown in this figure).
[0062] The user front-end interface 201 is configured for interacting with the user during training mode, as explained in greater detail in figure 4, and during fine-tuning mode, as explained in greater detail in figure 5.
[0063] The data repository-202 is configured for storing users’ personal data such that each command for operating movements of an avatar is associated with a cluster of user skeleton embeddings encoded from user movements during training mode and during fine-tuning mode. The cluster includes user skeleton embedding from various positions and angels. The avatar movements include, for example: walking, running, jumping and etc. The data repository may be any vector database such as the Annoy library and the Faiss or any other types of databases oriented to vector data or a combination thereof. The data repository stores one or more such clusters per each user.
[0064] The skeleton encoder 203 is configured for transforming the user’s time segment of skeleton tracing into skeleton embeddings. The operation of the skeleton encoder 203 is explained in greater detail in figure 3. The encoder can be implemented by various types of neural networks such as Skeletal Convolutional Neural Networks (S-CNN, see / 5 / ), Skeletal Recurrent Neural Networks (S-RNN, see / 6 / ), Graph Convolutional Networks for Skeletons (GCN-S, see / 7 / ), Skeletal Transformer Networks (see / 8 / ). The encoder can be implemented by other types of neural networks with the same deep feature skeleton encoding functionality and some combinations and modifications of above-mentioned types of neural networks can be used as Skeletal-based Neural Networks.
[0065] The computation module 204 is configured for associating, during training mode, user simple movements with avatar movements. The associating process is performed by receiving , via front-end interface 201 , user’s selection of command for operating movements of an avatar and by receiving user’s skeleton tracing via external motion capture program. Based on the received data, the computation module transforms time segments of the skeleton tracing into skeleton embeddings by using the skeleton encoder 203. The computation module 204 downloads these skeleton embeddings to the data repository 202 in association with a selected command for operating movements of an avatar. The computation module 204 is configured for clustering, in the data repository 202, the skeleton embeddings are associated with a certain command.
[0066] The computation module 204 is further configured for analyzing user’s operation of avatar movements during performing mode when user operates the avatar in Virtual Reality in real time. The analyzing is by
[0067] • Receiving user’s skeleton tracing from external motion capture program which captures user complex movements and transforms them into corresponded skeleton tracing.
[0068] • Transforming via the skeleton encoder 203 each received time-segment of skeleton tracing into input skeleton embedding
[0069] • Comparing this input skeleton embedding with clusters of skeleton embeddings stored in the data repository 202 for selecting the cluster that is closest to the input skeleton embedding.
[0070] • Retrieving the command for operating movements of an avatar associated with the closest cluster.
[0071] • Transferring this command for operating movements of an avatar to the external VR program.
[0072] The comparing process can be implemented, for example, by K-Nearest Neighbors classification algorithm to detect in the data repository a set of K skeleton embeddings, that are most close to the input skeleton embedding in terms of Euclidean distance. Since each data repository embedding corresponds to a command for operating movements of an avatar, the system selects the most probabilistically plausible user command for operating movements of an avatar, the one to which are corresponded the largest number from selected K embeddings. The K- Nearest Neighbors classification algorithm may be mathematically represented as:
[0073] Assume there is - training sample, set of skeleton embeddings from the data repository, N - number of embeddings in training sample clusters of skeleton embeddings, C - number of clusters in training sample and each cluster is corresponded to one of commands for operating movements of an avatar; Assume there is some distance functio Euclidean distance, for example. Assumeu as the input skeleton embedding. The system finds k of the closest embeddings to u, in terms of distance to objects of the training sample with known clusters
[0074] Finally, ) is associated with a proper command for operating movements of an avatar.
[0075] The comparing process can also be implemented by types of classification algorithms or other types of comparison processes which allow to select from the data repository skeleton embeddings and / or their vector derivatives more similar to the input skeleton embedding and to choose the command for operating movements of an avatar which the most probabilistically plausible associated with selected skeleton embeddings and / or their vector derivatives.
[0076] The comparing process can be performed in various skeleton embedding vector transformation processes. Examples of such transformations are described herein.
[0077] In one other embodiment the comparing process is implemented by comparing the input skeleton embedding with the centroids of the clusters of skeleton embeddings. In such a case the cluster whose centroid is closest to the input skeleton embedding is selected: represent the set of clusters of user skeleton embedding where each h cluster is associated with a certain command for operating the avatar. be a set of skeleton embeddings from the data repository representing particular cluster corresponded to one of commands for operating movements of an avatar.
[0078] Let epresents the input skeleton embedding.
[0079] The centroid vector c of cluste The Euclidean distance d betw
[0080] The input skeleton embedding is deemed as belonging to particular cluster if its distance from the centroid c of the cluster is minimal across all other centroids of the clusters Mathematically, the condition is:
[0081] The command for operating the avatar is identified as the command associated with the cluster of the
[0082] In one other embodiment the system performs a combination of the comparing method of centroids of clusters of skeleton embeddings described herein before with the magnitude of the dispersions of the distributions of such skeleton embeddings from the clusters.
[0083] In this case, when classifying the input skeleton embedding, the distances to the centroids are calculated with a correction coefficient that considers the specified dispersion of the skeleton embedding distribution.
[0084] In some embodiments such vector derivatives of skeleton embeddings are calculated and are saved in the data repository in advance for increasing computational performance. In such a case, when implementing the classification procedure, only these vector derivatives themselves are unloaded from the depository without skeleton embeddings.
[0085] In some cases, there is a problem to identify a command when the time segment of the skeleton tracing from which the input skeleton embedding is transformed represents a switching from one command to another, for example from running to jumping.
[0086] One other problem is how to classify a prolonged simple movement corresponded for a plurality of time segments. An example of such command is jumping when a user performs a sequence of short different body gestures combined with preparing to jump during the 1sttime segment, with jumping during the 2ndtime segment and with jump finishing during the 3rdtime segment. In this prolonged simple movement, the system associates with the same command “jumping” all skeleton embeddings transformed from three-time segments captured from this sequence of the simple movements associated with jumping body gestures. However, for a user it’s rather difficult to perform enough contrast simple movements when all body gestures for jumping up command are different from all body gestures for jumping down command.
[0087] Both above-described problems can be solved by a prolonged classification procedure. The prolonged classification classifies the command of operating of avatar movements associated with the cluster that is the closest to a sequence of input skeleton embeddings. The sequence of input skeleton embeddings includes the input skeleton embedding and one or more skeleton embeddings calculated from time segments of the skeleton tracing successive to the time segment of the input skeleton embedding.
[0088] Such a prolonged classification procedure may be implemented as described herein: Let training sample, set of skeleton embeddings from the data repository, N - number of embeddings in training sample
[0089] L clusters of skeleton embeddings, C - number of unique clusters in training sample and each cluster is corresponded to one of commands for operating movements of an avatar.
[0090] Let d be a distance function -Such a distance function can be, for example an Euclidean distance.
[0091] Assumeu as the input skeleton embedding. The system calculates k the closest embeddings to u, in terms of distance to objects of the training sample with known as sequence of adjacent input skeleton embeddings. Where n size of embedding vector, and 1 number of skeleton embeddings in sequence.
[0092] Le — probability of skeleton embedding u be a label , calculated by frequency of elements across
[0093] Let probability of sequence of skeleton embeddings expressed as exponentiation of likelihood probability for consistent float multiplication in computer calculations.
[0094] The new input skeleton embedding sequenc s deemed belonging to particular cluster if its probability of sequence of skeleton embeddings e a labe s maximal. Mathematically, the condition can be expressed as:
[0095] The system defines the command of avatar movements operating as classifie
[0096] In some embodiments the computation process can take into account hansition states (when it occurs a switching from one command for operating movements of an avatar to another or and when it occurs above described prolonged classification procedure) through the use of inertial prediction algorithms and the implementation of special transition command for operating movements of an avatar into the chain of commands for operating movements of an avatar that are performed by the avatar when such transitional situations are detected.
[0097] The computation module 204 is further configured for associating, during fine-tuning mode, user complex movements with avatar movements. To provide this fine-tuning process the computation module 204 identifies, during performing mode, skeleton tracing time-segments with suspicious command for operating movements of an avatar; that is to say, command for operating movements of an avatar in which there is not enough confident level in the identification dew to probability of correct recognition below a defined threshold. The system logs these skeleton tracing segments in combination with corresponding time-segments of VR stream. In fine-tuning mode, the computation module 204 receives from front-end interface 201 user’s feedback with correct command for operating movements of an avatar and updates skeleton embeddings in the data repository 202 based on this feedback.
[0098] The user interacts with User’s Personalized Gesture Recognition system 200 directly via the user front-end interface 201 and indirectly via external motion capture program and external VR program.
[0099] It should be noted that Virtual Reality can be also Augmented Reality (AR), Merge Reality (MR) or other kinds of realities where avatar displaying can be implemented. Correspondingly VR program and VR stream can be also AR, MR program and AR, MR stream or another program and stream associated with other kinds of realities where avatar displaying can be implemented.
[0100] Fig. 3 shows a flowchart diagram of a method for personalizing the operation of avatar movements, in accordance with some exemplary embodiments of the disclosed subject matter. Blocks 305 illustrates the training mode.
[0101] At block 305 the user trains the system to recognize simple movements. The training results in the form of user-trained skeleton embeddings that are stored in the personal user’s data repository. The method for training the system is explained in greater detail in figure 4.
[0102] Blocks 310, 315, 320, 325, 330, 335, 340 and 345 illustrate the performing mode. The performing mode is the mode in which the user operates the avatar in Virtual Reality in real time after the training and in which the system logs data for fine-tuning mode.
[0103] At block 310 the system receives the user’s skeleton tracing via external motion capture program.
[0104] At block 315 the system transforms time segments of the skeleton tracing into skeleton embeddings by using the skeleton encoder.
[0105] At blocks 320 and 325 the system classifies the input skeleton embeddings by comparing with user-trained skeleton embeddings from personal user’s data repository.
[0106] At block 320 the system selects a subset of user trained skeleton embeddings that are the nearest to the input skeleton embedding. The result is a subset of skeleton embeddings from user’s personalized data repository that are the nearest to input skeleton embedding. The selecting of the subset may be performed by the K-nearest neighbours’ method. The term nearest refers herein to the shortest distance utilizing distance methods such as the Euclidian distance method.
[0107] At block 325 the system divides the selected subset of the nearest skeleton embeddings into subgroups. Each subgroup includes skeleton embeddings that belong to a certain cluster associated with one command for operating movements of an avatar. The system selects the subgroup with the largest number of skeleton embeddings. The system selects the command for operating movements of an avatar that is associated with this subgroup.
[0108] At block 330 the system transmits the command for operating movements of an avatar to the external VR-program which performs the user’s avatar animation in VR-application.
[0109] At block 335 the system log is recorded in full time with the recorded stream of user’s avatar animation.
[0110] At block 340 the system identifies-suspicious skeleton tracing time segments. The suspicious segments are segments that include user complex movements in which the confidence level of identifying the command for operating movements of an avatar is below a threshold associated with the probability of correct recognition. Each time segment is identified as suspicious, the skeleton tracing is recorded in the system log in combination with the corresponded time segment of VR stream which is executed by VR program.
[0111] At block 345, the system uploads the log in the data repository.
[0112] Block 350 illustrates the fine-tuning mode. At block 350 the system performs fine tuning. At fine tuning mode the system presents to the user the suspicious segments and the user arbitrates correctness of these segments by approving the command for operating movements of an avatar (positive reward shaping) or by updating a wrong recognized command for operating movements of an avatar to the correct one (negative reward shaping with correct cluster). The method for fine tuning the system is explained in greater detail in figure 5.
[0113] It should be noted that Virtual Reality can be Augmented Reality (AR), Merge Reality (MR) or other kinds of realities where avatar displaying can be implemented. Correspondingly VR program and VR application can be also AR, MR program and AR, MR application or to another program and application associated with other kinds of realities where avatar displaying can be implemented.
[0114] Fig. 4 shows a flowchart diagram of a method for training a system for personalizing the operation of avatar movements, in accordance with some exemplary embodiments of the disclosed subject matter.
[0115] According to some embodiments, the system is trained to personalize simple movements by providing a user with an option to associate a certain command for operating movements an avatar with user simple movements; for example, the user may associate a simple movement for running with the command causing the avatar to run. According to some embodiments, the user trains the system with a plurality of variations of user simple movements each associated with the movement of the avatar. The user’s simple movements are performed by the user in different time segments. Each such user simple movements may correspond to the same personalized gesture but may differ in properties such as angle of capturing the movement, velocity of performing the movement, location of performing the movement, field of view, performance of the user and etc. The system associates each of the plurality of user simple movements with the selected command for operating movements of an avatar for allowing the user to personalize the operation of the avatar movements according to his personalized gestures.
[0116] Referring now to the drawing:
[0117] At block 400 the user calibrates the system settings by adjusting the parameters, among other things, to his body features.
[0118] At block 405 the user selects a command for operating movements of an avatar. The selection is for associating a plurality of similar user’s simple movements with the command for operating movements of an avatar.
[0119] At block 410 the system receives the selection of the user.
[0120] At block 415 the user trains the system to associate a certain personalized gesture with the selected command for operating movements of an avatar by performing several times the user simple movements. In each time the simple movement may be performed in another location or may be captured from a different angle or point of view or may be performed in a different manner.
[0121] At block 420 the system receives user’s skeleton tracing via external motion capture program.
[0122] At block 425 based on the received data, the system transforms time segments of the skeleton tracing into skeleton embeddings by using the skeleton encoder.
[0123] At block 430 the skeleton embeddings are uploaded in a data repository in association with the selected command for operating movements of an avatar. According to some embodiments the data repository is a personalized nearest neighborhood database.
[0124] It should be noted that each user may repeat the process for all the commands for operating movements of an avatar that he wishes to activate.
[0125] Fig. 5 shows a flowchart diagram of a method for fine tuning the control of an avatar, in accordance with some exemplary embodiments of the disclosed subject matter.
[0126] According to some embodiments the system enhances the data repository during the fine- tuning mode for providing user arbitrage of suspicious recognition cases based on this system log.
[0127] At block 500 the system receives from the log file that is stored in data repository the next time segment of the skeleton tracing-that is marked in the system log as suspicious. The log file is recorded during the performing mode. The suspicious segments are segments that include user movements in which the confidence level of identifying is below a threshold as explained in greater detail in figure 3.
[0128] At block 505 the system retrieves from the data repository the corresponding time segment of previously recorded (in performing mode) VR stream of user’ s avatar animation. The VR stream can be also changed to Augmented Reality (AR) , Merge Reality (MR) stream or to another stream associated with other kinds of realities where avatar displaying can be implemented.
[0129] At block 510 the system presents the time segment to the user. The presenting may be by playing the content of the segment in combination with displaying of the suspicious command for operating movements of an avatar. The presenting is for approving or disapproving the correctness of command for operating movements of an avatar.
[0130] At block 515 the system receives the response from the user.
[0131] At block 520, which occurs in case of positive approving, the system uploads the skeleton embeddings encoded from the time-segment of skeleton tracing that was presented to the user into the data repository in association with the approved command for operating movements of an avatar.
[0132] At block 525, which occurs in case of negative response from the user, the user corrects the command for operating movements of an avatar and the system uploads the skeleton embeddings encoded from the time-segment that was presented to the user into the data repository in association with the user’ s corrected command for operating movements of an avatar; additionally, the system detects, in data repository, existing skeleton embeddings that causes the wrong recognition and devalues them with a penalty coefficient. Such devaluation may also lead to the removal of these skeleton embedding from the data repository in subsequent iterations. The removal may occur when this stored skeleton embedding is marked many times by penalty coefficient and the accumulated penalty coefficient reaches a threshold.
[0133] The removal process may be mathematically represented as:
[0134] Assume there i training sample, set of skeleton embeddings from the data repository, N - number of embeddings in training sample, - clusters of skeleton embeddings, C - number of unique clusters in training sample and each cluster is corresponded to one of commands for operating movements of an avatar. Assume there is some distance function Euclidean distance, for example. Assume, there is a task to classify new object u which correspond to the input skeleton embedding. Assume there is a vector Φ of accumulated penalty for all embeddings in training sample X, as follows:
[0135] Each element of vector Φ has been updating during lifetime of the data repository. Assume - true cluster of new object u. Update of Φ vector appears in fine-tuning mode, when the user can give feedback about true onstant factor, which participates in Φ ’s update procedure in following way described below.
[0136] A there is vecto of accumulated penalties for embeddings from
[0137] Whe updated vector of accumulated penalties for vector
[0138] Assume, there is constant penalty threshold updated accumulated penalty of embedding then delete embedding
[0139] As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms. "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0140] It should be noted that, in some alternative implementations, the functions noted in the block of a figure may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.
Claims
AMENDED CLAIMS received by the International Bureau on July 21 , 2025 (21 .07.2025) CLAIMSWhat is claimed is:
1. A method, the method comprises: receiving a user selection of a certain command for operating movements of an avatar; receiving a plurality of motion data items, associated with a plurality of motion captures of said user while said user performing a plurality of simple movements to be associated with said command; said plurality of simple movement being distinguished from each other in an at least one capturing property; transforming said plurality of motion data items into a plurality of motion embedding vectors; and.associating, in a data repository, said motion-embedding vectors with said certain command and with said user ;_said associating being for personalizing user choice of which simple movement to associate with which command ; said associating being for personalizing said operation of said avatar by said user, said personalization including user choice of which simple movement to associate with which command.
2. The method of claim 1, wherein said simple movement comprises one member selected from a group consisting of: full body gestures, half-body (torso) gestures, fingers gestures, eyes gestures, face mimic and user’s interaction with a technical device.
3. The method of claim 2, wherein said user interaction includes one member selected from a group consisting of: pressing on controller buttons, pressing on a keyboard key, operating a joystick or performing a voice command.
4. The method of claim 1 , wherein said motion data items being a vector of time segment of skeleton tracing and wherein said motion embedding vector being a skeleton embedding vector.
5. The method of claim 1, further comprising, utilizing said cluster for classifying an operation of said user while playing with said avatar as operating said command.
6. A method, comprising:while operating an avatar receiving a vector of time segment of skeleton tracing, associated with a motion capture of an interaction of a user with a computing device for operating of an avatar; by skeleton encoder, transforming said vector of time segment of skeleton tracing into an input skeleton embedding vector; comparing said input skeleton embedding vector to clusters of skeleton embedding vectors; said clusters being generated prior to said operating said avatar and comprises ; said generating comprises associating variation of simple movements with a command selected by said user and with said user; said plurality of variations of simple movements being performed by said user prior to said operating said avatar for personalizing user choice of which simple movement to associate with which command; selecting from a data repository a selected cluster from said clusters; said selected cluster being identified by said comparing process as the nearest to said input skeleton embedding vector; retrieving from said data repository a selected command associated with said selected cluster; and transmitting said selected command to a second computing device for operating said avatar in accordance with said selected command.
7. The method of claim 6, wherein said second computing device being one member selected from a group consisting of Virtual Reality (VR) program, Augmented Reality (AR) program and Merge Reality (MR) program.
8. The method of claim 6, wherein said comparing comprises utilizing K-Nearest Neighbors classification algorithm.
9. The method of claim 6, wherein said comparing comprises comparing a centroid of said clusters to said input skeleton embedding.
10. The method of claim 6, wherein said comparing being in combination with comparing magnitude of the dispersion of said input skeleton embedding to a distribution of said cluster.
11. The method of claim 6, wherein said comparing comprises calculating a closest cluster to a sequence of other input skeleton embedding identified from one or moretime segments of skeleton tracing successive to a time segment of said input skeleton embedding12. The method of claim 6, further comprising generating a log file during said operating said avatar; identifying suspicious skeleton tracing time segments combined with correctness of classifying of a command for operating movements of an avatar, presenting said identifies-suspicious skeleton tracing time segments to said user and updating a data repository personalised for said user in accordance with a response of said user combined with a value of said command for operating movements of an avatar.
13. The method of claim 6, wherein said variations comprise one member selected from a group consisting of: a plurality of angles of capturing, a plurality of velocities of performing said movement, a plurality of locations of performing said movement and a plurality of field of views.
14. The method of claim 6, wherein said generating comprises associating a certain simple movement of a user with a selection of a certain command of operating an avatar; said associating being for personalizing user choice of which simple movement to associate with which command, receiving a user selection of a command for operating movements of an avatar; receiving a plurality of motion data items, associated with a plurality of motion capture of said user while said user performing a plurality of simple movement to be associated with said command; said plurality of simple movement being distinguished from each other in an at least one capturing property; transforming said plurality of motion data items into a plurality of motion embedding vectors; and generating a cluster, said generating comprises associating, in a data repository, said plurality of motion-embedding vectors with said command and with said user; said cluster being for personalizing said association of said selected command with said user simple movements for personalizing said operation of said avatar by said user.