Real-time style for virtual environment motion

By matching the generated feature dataset with the stored motion dataset, and combining scaling techniques and synthesis methods, the limitations of user motion tracking in virtual reality systems are solved, enabling real-time motion stylization of the user avatar and improving user experience and interaction.

CN115244495BActive Publication Date: 2026-01-23MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180020366.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-10
Filing Date
2021-03-01
Publication Date
2026-01-23
Estimated Expiration
2041-03-01

AI Technical Summary

Technical Problem

Existing virtual reality systems have limitations in user motion tracking and mapping, which prevents users from fully performing virtual avatar actions when space or physical constraints are present, thus affecting user experience and interaction.

Method used

By receiving user input data, generating a feature dataset, and matching it with a stored motion capture dataset, full-body motion is generated using scaling techniques and synthesis methods to achieve real-time motion styling of the user avatar, including one-to-one mapping, position-based scaling, and trajectory-based scaling techniques.

Benefits of technology

It improves users' freedom of movement and natural interaction experience in the virtual environment, reduces user fatigue and risk of injury, and enhances users' control and sense of agency over their virtual avatars.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115244495B_ABST
    Figure CN115244495B_ABST
Patent Text Reader

Abstract

Examples of the present disclosure describe systems and methods for providing real-time motion patterns in a virtual reality (VR), augmented reality (AR), and / or mixed reality (MR) environment. In aspects, input data corresponding to a user interacting with a VR, AR, or MR environment can be received. The input data can be characterized to generate a feature set. The feature set can be compared to a stored motion data set that includes motion capture data representing one or more motion patterns for performing an action or activity. Based on the comparison, the feature set can be matched to feature data for one or more motion patterns in the stored motion data. The one or more motion patterns can then be performed by a virtual avatar or virtual object in the VR / AR / MR environment.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Virtual reality (VR) systems provide simulated environments in which users can interact with virtual content. VR systems use motion data corresponding to a user's real-world movements to manipulate an avatar of the user in the simulated environment. Typically, the avatar's movements and motions are limited to the user's personal capabilities and / or the physical limitations of the user's real-world environment. As a result, many users experience a significant decrease in the experience when using such VR systems.

[0002] It is with respect to these and other general considerations that aspects disclosed herein have been made. Also, while the examples are amenable to variations and alterations, some examples will be discussed for the sake of clarity. SUMMARY

[0003] Examples of the present disclosure describe systems and methods for providing real-time motion styling in virtual reality (VR), augmented reality (AR), and / or mixed reality (MR) environments. In aspects, input data corresponding to a user's interaction with an environment of a VR, AR, or MR can be received. The input data can be characterized to generate a feature set. The feature set can be compared to a stored set of motion data that includes motion capture data representing one or more motion styles for performing an action or activity. Based on the comparison, the feature set can be matched to feature data for one or more motion styles in the stored motion data. The one or more motion styles can then be performed by a virtual avatar or virtual object in the VR / AR / MR environment. For example, the motion style can be applied to a virtual avatar that accurately represents the position, shape, and size of the user's body and / or limbs. The virtual avatar can perform the motion style in real-time such that the motion style closely matches the user's interaction.

[0004] This summary is provided to introduce some concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used in limiting the scope of the claimed subject matter. Additional aspects, features, and / or advantages of examples will be set forth in the description that follows, and in part will be apparent to those with skill in the art based upon the disclosure, or can be learned by practice of the disclosure. Examples BRIEF DESCRIPTION OF DRAWINGS

[0005] Non-limiting and non-exhaustive examples are described with reference to the following figures.

[0006] Figure 1 An overview of an example system for providing real-time motion styling in VR, AR, and / or MR environments described herein is illustrated.

[0007] Figure 2 FIGURE illustrates an example input processing unit for providing real-time motion patterns in VR, AR, and / or MR environments described herein.

[0008] Figure 3 FIGURE illustrates an example method for providing real-time motion patterns in VR, AR, and / or MR environments described herein.

[0009] Figure 4 is a block diagram illustrating example physical components of a computing device with which various aspects of the present disclosure can be practiced.

[0010] Figure 5A and 5B is a simplified block diagram of a mobile computing device with which various aspects of the present disclosure can be practiced.

[0011] Figure 6 is a simplified block diagram of a distributed computing system in which various aspects of the present disclosure can be practiced.

[0012] Figure 7 FIGURE illustrates a tablet computing device for performing one or more aspects of the present disclosure. DETAILED DESCRIPTION

[0013] Various aspects of the present disclosure are described more fully below with reference to the accompanying drawings, which form a part hereof, and which show specific example aspects. However, different aspects of the present disclosure can be implemented in many different forms and should not be construed as limited to the aspects set forth herein; rather, these aspects are provided as a breadth of disclosure to provide an overall understanding of the aspects of the present disclosure and the scope thereof. The aspects can be practiced as methods, systems, or apparatuses. Accordingly, the aspects can take the form of a hardware implementation, a software implementation, or an implementation combining software and hardware aspects. The following detailed description is, therefore, not to be taken in a limiting sense.

[0014] VR enables users to transport themselves into virtual environments simply by donning a VR headset and picking up a VR controller. As the user moves their head and hands during a VR session, the motion captured by the VR tracking system (e.g., via sensors in the headset, controller, and other components of the VR system / application) is applied to the visual rendering presented to the user. Many VR systems display a virtual avatar that is a graphical representation of the user. In such systems, the motion detected by the headset and controller is mapped to corresponding parts of the avatar. The mapping process often involves a direct coupling of the user’s motion to the motion of the avatar, which enables the user to maintain full agency over the avatar. As used herein, direct coupling refers to a one-to-one coupling of the user’s physical movements and the user’s avatar’s physical movements. As a result, direct coupling enables natural interaction by the user since hand-eye coordination and proprioception are accurately represented by the visual avatar. In many VR systems that employ direct coupling, many parts of the user’s body can not be tracked due to missing sensors, occlusions, or noise. As a result, poses for untracked body parts are interpolated by the VR system using, for example, inverse kinematics, animation that adapts to the application context, or pre-captured data.

[0015] While direct coupling is simple, many VR systems can be limited in their use of direct coupling. For example, in some scenarios, such as in action games, the user’s physical capabilities in the physical (e.g., real) world can be limited due to spatial constraints of the user’s physical environment, the user’s physical limitations, or some combination thereof. Such physical limitations can adversely affect the user’s virtual avatar, which can be represented by an athlete, a warrior, a dancer, etc. For example, when participating in a tennis application of a VR system, the user can not have enough space in the user’s physical environment to simulate one or more actions (e.g., a serve, a full backhand swing, movements associated with court coverage, etc.). Or, the user can not be able to perform these actions due to physical injury or lack of physical training. In either case, as a result, the user’s avatar can not perform optimally (or even adequately) in the tennis application. This suboptimal performance can severely degrade the user experience and hinder the user from expressing the desired virtual motion.

[0016] Moreover, many VR systems that employ direct coupling to provide sensors can fail to track one or more parts of a user's body due to lack of sensors, occlusions, or noise. For example, a VR system that only tracks a user's head and hand positions (via headphones and hand controllers) can map the positions of the head and hands to the user's avatar. However, many other body positions for the avatar (e.g., leg positions, foot positions, etc.) can be unknown. Thus, poses for the untracked body parts are often interpolated by the VR system using, for example, inverse kinematics, animation that adapts to the application context, or pre-captured data. However, such techniques can limit the user's perception of user avatar control, implementation, and / or agency.

[0017] To address such challenges in VR systems, the present disclosure describes systems and methods for providing real-time motion patterns in VR, AR, and / or MR environments. In aspects, a user can access a virtual reality (VR), augmented reality (AR), or mixed reality (MR) system. The system can include a head-mounted display (HMD), zero or more than zero controller devices, and / or zero or more than zero additional sensor components, or remote sensing of user poses from RGB cameras, depth cameras, or any other sensing such as acoustic or magnetic sensing. The HMD can be used to present an environment to the user that includes two-dimensional (2D) and / or three-dimensional (3D) content. The controller(s) or user gestures and speech can be used to manipulate content in the environment. The content can include, for example, 2D and 3D objects such as avatars, video objects, image objects, control objects, audio objects, haptic rendering devices, etc.

[0018] In aspects, input data (e.g., motion data, audio data, speech, textual data, etc.) corresponding to a user’s interaction with a system can be received by one or more sensors (wearable or remote sensing) of the system. The input data can be organized into one or more segments or frames of data. For each segment / frame of data, a feature dataset can be generated. The feature dataset can include, for example, object type information, velocity data, acceleration data, position data, sound data (such as speech data, footfalls, or other motion noise data), pressure data, and / or torque data. In at least some examples, the feature dataset can be transformed to a particular coordinate system / space, or can be defined with respect to a particular coordinate system / space. The feature dataset(s) can be used to scale output motion to be applied to a user avatar. As used herein, scaling refers to interpolating motion from the input data. Examples of scaling techniques include one-to-one mapping (e.g., user motion is directly reflected by the user avatar), position-based scaling, trajectory-based scaling, or combinations thereof. In examples, scaling enables a user to perform motion in a physical environment with a smaller footprint while observing fully expressed avatar motion. Thus, scaling enables a user to perform larger scale or more extensive motion, or to perform more rigorous or more extensive motion, in a spatially limited environment without becoming fatigued.

[0019] In aspects, feature data can also be compared to motion capture datasets. Motion capture data can include motion data corresponding to various actions, activities, and / or motions performed by one or more subjects. Alternatively (or additionally), motion data can be synthetically generated using, for example, animation or physical simulation. In examples, each subject (or grouping of subjects) can exhibit different or specific motion for performing an action, activity, or movement. The specific motion performed by each subject (or grouping of subjects) can represent a “style” of the subject(s). In some aspects, feature data associated with motion data can be stored with the motion data. The feature data can be normalized to one or more body parts or coordinate systems. For example, the feature data can represent head relative motion. Alternatively, the feature data can be associated with depth video, sound motion data, user brain waves (or other biological signals), etc.

[0020] In aspects, feature data for input data can be compared to feature information for motion capture data. Based on the comparison, zero, one, or more candidate motions in the motion capture data can be identified. For each candidate motion, a match score representing a similarity between the candidate motion and the input data can be computed. Based on the match scores, a motion corresponding to the user interaction can be synthesized. The synthesized motion is a blend of the matching motions and the user one-to-one mapping data weighted by the relative distance between the feature vectors of the matching motions. In some examples, the one-to-one data can also be weighted. The synthesized motion can be used to animate portions of the user avatar corresponding to movements detected by sensors of the system. In some aspects, the synthesized motion can further be used to interpolate full body motion. For example, while the input data can only correspond to user head and hand motions, one or more synthesized motions can be generated for user parts that have not been detected by sensors of the system. Synthesizing motion data for untracked parts of the user can include generating separate mappings for each body part, and fusing the separate mappings (each limb) to produce full body motion. The synthesized full body motion can then be applied to the user's avatar. As a result of the synthesis process, the user's original input motions can be transformed in real time into graceful and / or agile full body animations that are performed in one or more subject styles.

[0021] Accordingly, the present disclosure provides a variety of technical advantages, including but not limited to: adaptively decoupling user input motions from virtual avatar rigging, stylizing user input to match motions performed by experts or other users, generating full body motions from a subset of user input, scaling motions as part of motion stylization, preserving hand-eye coordination, macrosensation, and proprioception when rigging virtual avatar motions, using data sets of motion data to stylize motion trajectories, evaluating a set of candidate motions using feature data of user input and stored motion data, converting body motions into a user-centric coordinate system / space, evaluating multiple body segments individually over multiple data frames, using one or more algorithms to rank candidate motions / positions, scaling user input to achieve low motion footprint for users in spatially constrained environments, stylizing motions for users with physical limitations, and enabling users to reduce fatigue and injury risk when using VR, AR, and / or MR systems, among others.

[0022] Figure 1An overview of an example system for providing real-time motion style in VR, AR, and / or MR environments described herein is illustrated. The presented example system 100 is a combination of interdependent components that interact to form an integrated whole for rendering 2D and / or 3D content in a virtual environment or an environment that includes virtual content. The components of the system can be hardware components or software implemented on and / or executed by hardware components of the system. In an example, the system 100 can include any of hardware components (e.g., for executing / running an operating system (OS)) and software components (e.g., applications, application program interfaces (APIs), modules, virtual machines, runtime libraries, etc.) running on the hardware. In one example, the system 100 can provide a runtime environment for software components, adhere to constraints set for operations, and utilize resources or facilities of the system 100, where the components can be software (e.g., applications, programs, modules, etc.) running on one or more processing devices. For example, software (e.g., applications, operating instructions, modules, etc.) can run on a processing device, such as a personal computer (PC), a mobile device (e.g., a smartphone / phone, a tablet, a laptop, a personal digital assistant (PDA), etc.), and / or any other electronic computing device. As an example of a processing device operating environment, reference is made to the example operating environment depicted in Figures 4-7 In other examples, the system components disclosed herein can be distributed across multiple devices. For example, information can be input on a client device and can be processed or accessed from other devices in a network, such as one or more server devices.

[0023] As one example, the system 100 includes a computing device 102 and a computing device 104, a virtual environment service 106, a distributed network 108, and motion capture data store 110. Those skilled in the art will appreciate that the scale of a system such as the system 100 can vary and can include more or fewer components than those shown in Figure 1more or less components than those depicted in the figures. In aspects, computing device 102 and computing device 104 can be any of a variety of computing devices, including but not limited to the processing devices described above. Computing device 102 and computing device 104 can be configured to use one or more input devices for interacting with virtual environment service 106, such as a head-mounted display device or alternative virtual environment visualization system, zero, one, or more controller devices (e.g., joysticks, control sticks, force balls, trackballs, etc.), remote sensing of user motion, including microphones, video cameras, depth cameras, radar, magnetic or acoustic sensing, data gloves, body suits, treadmills or motion platforms, keyboards, microphones, one or more haptic devices, etc. Such input devices can include one or more sensor components, such as accelerometers, magnetometers, gyroscopes, etc. In examples, the input devices can be used to interact with and / or manipulate content presented using virtual environment service 106.

[0024] Virtual environment service 106 can be configured to provide virtual environments and / or apply virtual content to physical or virtual environments. For example, virtual environment service 106 can provide a VR environment, or can provide virtual content and interactions displayed in AR and MR environments. In aspects, virtual environment service 106 can be provided, for example, as part of an interactive productivity or gaming platform. It should be understood that while virtual environment service 106 is illustrated as being separate from computing device 102 and computing device 104, virtual environment service 102 (or one or more components or instances thereof) can be provided by computing device 102 or computing device 104, alone or in combination. As a particular example, computing devices 102 and 104 can each provide a separate instance of virtual environment service 106. In such examples, the instance of virtual environment service 106 can be accessed locally on computing device 102 using a stored executable; however, computing device 104 can access the instance of virtual environment service 106 over a network, such as distributed network 108. In aspects, virtual environment service 106 can provide virtual content to one or more users of computing device 102 and computing device 104. Virtual content can include interactive and non-interactive elements and content. As one example, virtual environment service 106 can provide an interactive virtual avatar of a user. The user can interact with the virtual avatar using the input devices described above. When user motion is detected from one or more input devices, virtual environment service 106 can search motion data store 110 for motion data that matches the user motion. Virtual environment service 106 can apply the matching motion data to the virtual avatar, causing the virtual avatar to perform the motion data.

[0025] Motion capture data store 110 can be configured to store motion capture data and additional data associated with the motion data, such as audio data, pressure data, torque data, depth point clouds, facial expression data, etc. The motion capture data can include motion data associated with various actions, activities, and / or movements. The motion capture data can include full body motion data and / or partial body motion data (e.g., only head and hand motion data). In some examples, the motion data can be organized into one or more categories, such as human interaction data, action data, sports activity data, non-human motion data, etc. The motion data can be collected from multiple subjects using one or more motion capture systems / devices, such as inertial motion capture sensors, mechanical motion capture sensors, magnetic motion capture system sensors, optical motion capture sensors, RGBD cameras, radar, recovery from existing video, animation, physics simulation, non-human actors, etc. The subjects can represent or be categorized into various subject categories, such as experts or professionals, successful or famous subjects, medium or low experience subjects, etc. In aspects, the motion data can be associated with a feature data set. The feature data can include information related to various poses performed during a particular motion. Such information can include, for example, velocity, acceleration, and / or position data for one or more segments of a user’s body during a motion. Such information can additionally include depth point cloud data, audio data, pressure data, torque data, brain wave recordings, etc. For example, a particular motion can be segmented into 360 individual data / pose frames within 3 seconds.

[0026] Figure 2 An overview of an example input processing system 200 for providing real-time motion patterns in VR, AR, and / or MR environments described herein is illustrated. The motion pattern technology implemented by input processing system 200 can include Figure 1 the technology and data described in the systems of

[0027] In aspects, input processing system 200 can generate or provide access to one or more virtual (e.g., VR, AR, or MR) environments. The virtual environments can be viewed using an HMD or similar display technology (not shown) and can include virtual content and / or physical (e.g., real world) content. The virtual content can be manipulated (or otherwise interacted with) using one or more input devices (not shown). In Figure 2In particular embodiments, input processing system 200 includes input detection engine 202, feature analysis engine 204, mapping component 205, style analysis engine 206, motion scaling engine 208, and stylized motion execution engine 210. Those skilled in the art will appreciate that the scale of system 200 can vary, and can include more or fewer components than those described in Figure 1

[0028] Input detection engine 202 can be configured to detect user input. User input can be provided by one or more input devices operated by one or more users, as well as possibly external sensing of the user(s). In examples, input detection engine 202 can include one or more sensor components, such as accelerometers, magnetometers, gyroscopes, etc. Alternatively or additionally, such sensor components can be implemented into the input devices operated by the user. In either case, input detection engine 202 can detect or receive input data corresponding to user interaction with a virtual environment or virtual (or physical) content thereof. Input data can include motion data, audio data, textual data, eye tracking data, object or menu interaction data, etc. In many aspects, input data can be collected as one or more data files or data sessions, and / or segmented into one or more data files. Input detection engine 202 can store input data in one or more data storage locations, and / or provide input data to one or more components of input processing system 200, such as feature analysis engine 204.

[0029] ​The feature analysis engine 204 can be configured to generate a feature dataset from the received input data. In aspects, the feature analysis engine 204 can receive (or otherwise access) the received input data. Upon accessing the input data, the feature analysis engine 204 can perform one or more processing steps on the input data. As one example, the feature analysis engine 204 can segment the input data into a set of data frames. For example, each data frame can represent an "N" millisecond block of the input data. For each set of frames, position data for each input device for which input has been received / detected can be identified. For example, position data for a user's head and hands can be identified. In some aspects, the position data can be translated into a head-centered coordinate system. That is, the position data can be normalized to a user head space. After converting / normalizing the position data, feature data can be generated for each set of frames. The feature data can include acceleration, velocity, and position information for each input device. The feature data can be used to generate one or more feature vectors. As used herein, a feature vector can refer to an n-dimensional vector that represents numerical features of one or more data points or objects. In at least one aspect, the feature analysis engine 204 can generate features for more than one user. The features generated for multiple users can be used by the feature analysis engine 204 to adapt a motion based on input from multiple users.

[0030] In aspects, the feature analysis engine 204 can be further configured to generate a comparative feature dataset from stored motion data. In aspects, the feature analysis engine 204 can access a data repository that includes motion capture data for various actions, activities, events, and / or movements. The motion capture data can be organized and / or stored in data sessions. Each data session can represent a particular action, activity, or event performed by a particular subject (or subjects). For each data session, the feature analysis engine 204 can translate the motion in the data session into a head-centered coordinate system, as previously described, or into an alternative projection of the data. This translation can be invariant to general rotations and translations, or varying for rotations and position (e.g., global position). The data session can then be decimated to match a frame rate of one or more input devices. As a specific example, a data session can be decimated to match a 90 Hz frame rate of a VR system controller. Each data session can then be divided into time blocks that include position data for each input device as a function of time. For example, each data session can be divided into 100 millisecond blocks using a sliding window with 50 millisecond overlap. A representative gesture window (e.g., a sequence of gestures) can be selected from various gesture windows in the data session. The representative gesture window can generally represent the particular action, activity, or event performed in the data session.

[0031] In some aspects, to make the subsequent matching computations invariant to rotation about the Y-axis, each pose window in the motion capture data can be transformed along the Y-axis into the coordinate space of the representative pose window. This transformation includes finding the rotation between two 3D time point clouds where the correspondence is known. In at least one example, the following equation can be used:

[0032]

[0033] In the above equation, w i represents the weight assigned to each point in the time window, n represents the number of frames in the data session, q represents the pose window to be rotated, s represents the representative pose window, and x and z represent the x and z coordinates of the subject’s left and right hands, respectively. In this equation, the weights are used under the assumption that the position vectors at the beginning of the pose window should have a greater weight than the position vectors at the end of the pose window. After rotation normalization for each data session, feature data (e.g., comparison feature data) can be produced for the data session. A feature vector can be produced from the feature data and stored with or for the data session.

[0034] The mapping component 205 can be configured to map the feature data to a sensed motion. In aspects, the mapping component 205 can map the feature data to one or more motions. For example, the mapping component 205 can map the feature data to head and hand movements corresponding to a tennis serve motion. The mapping can include, for example, non-uniform scaling of user motions to full body motions, correcting missing or different limb motions, mapping from human motions to non-human motions, mapping from non-human motions to human motions, and / or other transformations required by the VR system / application. As one specific example, the mapping component 205 can map feature data corresponding to a walking motion to a walking motion of an avatar represented as an eight-legged spider.

[0035] The style analysis engine 206 can be configured to compare the received input data to the stored motion data. In aspects, the style analysis engine 206 can access a feature vector for the input data ("input feature vector"). After accessing the input feature vector, the style analysis engine 206 can iterate through data sessions in the stored motion data. In some examples, upon accessing the stored motion data, the style analysis engine 206 can generate (or cause the feature analysis engine 206 to generate) a comparison feature dataset for the stored motion data as described above. Once the input feature vector and the feature vectors of the data sessions ("comparison feature vectors") are in the same coordinate space (e.g., a head-centered coordinate system), a comparison is performed. The comparison can include computing a matching distance between the input feature vector and each comparison feature vector using, for example, a k-nearest neighbor algorithm. Based on the comparison, the top k values within distance D can be identified. In examples, the comparison can be performed separately for different limbs. For example, the top k t candidate matches for the user's left hand can be identified, and the top k r candidate matches for the user's right hand can be identified.

[0036] In aspects, the various candidate matches for each input device can be used to synthesize a final stylized motion. The synthesized stylized motion can be limited to those portions of the avatar for which motion was detected using one or more input devices. For example, if the user motion was detected only by the HMD and two hand controllers, the corresponding stylized motion can be generated only for the head and hands of the avatar. Alternatively, the synthesized stylized motion can incorporate extrapolated data for one or more portions of the avatar for which motion was not detected. For example, if the user motion was detected only by the HMD and two hand controllers, the corresponding stylized motion can incorporate full body (e.g., torso, leg joints, feet, etc.) motion that matches the detected motion. As a result, the final stylized motion can retain the characteristics of the user input motion while stylizing motion to the entire body of the avatar.

[0037] In aspects, the final stylized motion can be synthesized using various techniques. As one example, the distribution of features in the stored motion data can be modeled as a mixture of Gaussians. A given candidate match pose can be modeled as a weighted linear combination of the means used in the mixture model, and the two matches can be combined to maximize the likelihood of the resulting interpolated match. As another example, the distance metric (D) computed during the comparison described above can be used to accomplish the synthesis. For example, the distances of the k matches, denoted as d k , can be used to construct weights that are inversely proportional to d k . For the left and right hands, the following equations can be used:

[0038] and

[0039] In the above equation, W l represents the weight for the left hand, W r represents the weight for the right hand, d max represents the maximum distance threshold for a valid match, and W is between 0 and 1. Based on the above equation, the output pose O and each joint probability J can be calculated using the following equations:

[0040] and

[0041] To maintain spatial consistency, the output pose O can be adapted based on the joint probabilities with a human skeleton motion model. To maintain temporal consistency, the following exponentially weighted moving average can be used:

[0042] O t = JO t + (1 - J)O t-1

[0043] By rotating O t backwards and adding the head translation vector to O t , the resulting output pose O t can be converted back to the user input space.

[0044] The motion post-processing engine 208 can be configured to transform the output motions for the input data. Example methods of transformation include motion scaling (e.g., one-to-one scaling, position-based scaling, trajectory-based scaling, etc.), applying motion filtering (e.g., cartoon animation filtering, motion blur filtering, motion distortion filtering, etc.), applying motion smoothing techniques (e.g., using a smoothing convolution kernel, a Gaussian smoothing processor, etc.), and the like. Such transformation methods can enable a user to maintain a lower motion footprint when performing motions in a spatially limited environment, or enable a user to avoid fatigue when performing larger or greater amounts of motion. In aspects, the motion post-processing engine 208 can map the input feature vectors and / or position information of the input data to one or more virtual objects in a virtual environment. This mapping can include one or more motion transformation techniques. As one example, the motion post-processing engine 208 can implement a one-to-one mapping, where the motions detected by the user are strongly coupled into the rendered motions of the user’s avatar. That is, the user’s motions are directly reflected in the virtual environment as performed. As another example, the motion post-processing engine 208 can implement position-based scaling, where a factor can be used to elongate or shrink the distance of a virtual object from a particular point. As a particular example, a constant factor can be used to increase or decrease the reachable range of an avatar’s hands relative to the shoulder position of the avatar. This can be mathematically represented as:

[0045]

[0046] In the above equation, s is a scale factor. As yet another example, motion post-processing engine 208 can implement a trajectory-based scaling in which a factor can be used to scale the velocity of the virtual object. As a particular example, rather than scaling the position of the hand of the avatar, a constant factor is used to scale the velocity of the hand of the avatar. This can be mathematically represented as:

[0047] In the above equation, s is a scale factor. As yet another example, motion post-processing engine 208 can implement an adaptive scaling technique using different scaled regions based on the distance of the object to the user’s body. For example, motions near the user’s body can be strongly coupled to maintain proprioception; however, more scaling can be applied when the object (e.g., hand) is far from the user’s body. In this approach, when the virtual object (e.g., virtual hand) is extended to the maximum extent of the avatar, scaling can be disabled to prevent the appearance of unnatural extension of the virtual object.

[0048] Style motion execution engine 210 can be configured to render a stylized motion. In aspects, style motion execution engine 210 can apply the motion stylization determined by style analysis engine 206 and the motion transformations (e.g., scaling, applying filters, etc.) determined by motion post-processing 208 to the user avatar or to a substitute virtual object. Applying the motion stylization and scaling can result in the avatar or virtual object performing one or more stylized motions. The motion stylization and scaling can be applied so that the user is able to maintain full agency over the avatar or virtual object throughout the user’s period of movement.

[0049] Having described various systems that can be employed by aspects disclosed herein, the present disclosure will now describe one or more methods that can be performed by various aspects of the present disclosure. In aspects, method 300 can be performed by an execution environment or system, such as system 100 of Figure 1 or system 200 of Figure 2 However, method 300 is not limited to such examples. In other aspects, method 300 can be performed on an application or service that provides a virtual environment. In at least one aspect, method 300 can be performed (e.g., computer-implemented operations) by one or more components of a distributed network, such as a web service / distributed network service (e.g., a cloud service).

[0050] Figure 3An example method 300 for providing real-time motion stylization in VR environments described herein is illustrated. Although method 300 is described herein with respect to VR environments, it should be understood that the techniques of method 300 can also be applied to other virtual or 3D environments, such as AR and MR environments. In aspects, a virtual application, service, or system, such as virtual environment service 106, can provide a virtual environment that includes various 2D and / or 3D objects. The virtual application / service / system (hereinafter “virtual system”) can be associated with a display device for viewing the virtual environment and one or more input devices for interacting with the virtual environment. The virtual system can utilize a rendering component, such as stylized motion execution engine 210. The rendering component can be used to render various virtual content and objects. In an example, at least one of the rendered objects can be a virtual avatar representing a user interacting with the virtual environment.

[0051] Example method 300 begins at operation 302, where input data corresponding to a user’s interaction with a virtual system can be received. In aspects, one or more input devices including sensor components can be used to provide input data to the virtual system. The input devices (or their sensor components) can detect input data, such as motion data, audio data, text data, eye tracking data, etc., in response to various user inputs. As a specific example, a user of a VR application of a virtual system can interact with the VR application using an HMD and two hand controllers. The HMD and hand controllers can each include one or more sensor components (e.g., accelerometers, magnetometers, gyroscopes, etc.). While using the HMD and hand controllers, the user can simulate an action, such as a tennis serve. The sensors of the HMD and hand controllers can collect and / or record motion data of the simulated action. For example, the motion data can include velocity, acceleration, and / or position information for the user’s head and hands over the duration of the simulated action. In some aspects, the motion data can be collected and / or stored as one or more data frames, as technically capable by the virtual system. For example, the input data can be stored in data frames at a 90 Hz frame rate of the virtual system’s hand controllers. The collected motion data can be provided to the VR application in real-time or at set intervals (e.g., every second). A data collection component of the VR application program, such as input detection engine 202, can receive and / or store the input data.

[0052] At operation 304, feature data can be generated using the received input data. In aspects, the received input data can be provided to an input data analysis component of the VR application or virtual system, such as the feature analysis engine 204. The input data analysis component can perform one or more processing steps on the input data. One processing step can include segmenting the input data into a set of data frames, each data frame representing a time period in the input data. For example, the input data can be segmented into 100 millisecond data frames. Alternatively, the input data analysis component can simply identify the segments in the set of data frames. Another processing step can include identifying motion data in each data frame. For example, in each data frame, position data can be identified for each input device used to interact with the virtual system, regardless of whether the input device has detected input data. Alternatively, position data can be identified only for those input devices that have detected input data. In some aspects, the position data identified for each data frame can be converted into a particular coordinate system. For example, the position data can be normalized to a particular head space or head position of the user or the user's HMD. Although normalization with respect to the user's head space / position has been specifically mentioned herein, alternative approaches and normalization points are also contemplated. Yet another processing step can include generating feature data for the input data. For example, for each data frame, feature data can be generated for each input device represented in the data frame. The feature data can include acceleration, velocity, and / or position information for each input device. The feature data can be used to produce one or more feature vectors representing the feature data for one or more data frames. In at least one example, the feature vector can be a concatenation of acceleration, velocity, and position, encapsulating the spatio-temporal information for one or more data frames.

[0053] At operation 306, the feature data can be compared to a motion dataset. In aspects, the generated feature data can be provided to (or otherwise accessed by) a feature comparison component of the VR application or virtual system, such as the style analysis engine 204. The feature comparison component can access a data store of motion capture data and / or video data, such as the Carnegie Mellon University Motion Capture Dataset. The data store can include motion data corresponding to various actions or activities performed by one or more motion capture subjects. As a specific example, the data store can include motion data files (or other data structures) corresponding to the activities of tennis, boxing, basketball, running, and swimming. The motion data for each activity can be grouped into sub-activities of the activity. For example, the tennis motion data files can be divided into the following categories: serve motion data files, forehand motion data files, backhand motion data files, volley motion data files, etc. For each activity (and / or sub-activity), motion data for multiple motion capture subjects can be stored. For example, the tennis motion data can include motion data for tennis professionals, such as Serena Williams, Roger Federer, Stefan Graf, and Jimmy Connors. Because each motion capture subject can use different or unique motions to perform each activity (and / or sub-activity), the motion data for a particular motion capture subject can represent the subject’s motion “style.” In some aspects, the motion data of the data store can include corresponding feature data or can be stored with corresponding feature information. For example, each motion data file can include a corresponding feature vector.

[0054] In aspects, feature data for input data ("input feature data") can be matched to motion capture data indexed by the features ("comparison feature data"). The evaluation can include comparing one or more feature vectors associated with the input feature data to one or more feature vectors associated with the comparison feature data. The comparison can include iterating through each motion data file or data structure. Alternatively, the comparison can include searching a data store using one or more search utilities. As a specific example, a feature vector associated with input feature data for a simulated tennis serve can be provided to a motion identification model or utility. The motion identification model / utility can analyze the feature vector to determine a category of activity (e.g., tennis) associated with the feature vector. Based on the determination, the motion identification model / utility can identify or retrieve a motion data set classified as tennis (or related to tennis). In some aspects, the comparison can further include using one or more classification algorithms or techniques, such as k-nearest neighbors, logistic regression, Naive Bayes classifier, support vector machines, random forests, neural networks, etc. The classification algorithm / technique can be used to determine a set of candidate motions that approximate match to the input data. For example, a k-nearest neighbors algorithm can be used to determine the top k candidates that match the values of the feature vector associated with the input feature data. The determination can include analysis of motion data received from multiple input devices. For example, the top k candidates that match a user's left hand can be identified, and the top k candidates that match the user's right hand can be identified. t r

[0055] ​​At operation 308, a motion can be synthesized from the set of matching candidates. In aspects, a component of the feature motion synthesis VR application or virtual system, such as the style analysis engine 204, can determine one or more top / best candidates from the list of motion candidates. In a specific example, the determination can include using a k-nearest neighbor algorithm to identify the highest value(s) of Euclidean distance D with one or more features in the feature data. In such an example, if no value is determined to be within distance D, then no candidate can be selected from the list of motion candidates. The determined top candidate motion can be used to synthesize a style motion or select a motion style from the set of motion data. Synthesizing a style motion can include using one or more synthesis techniques. As one example, the distribution of features in the stored motion data can be modeled as a mixture of Gaussians. A given candidate matching pose can be modeled as a weighted linear combination of the means used in the mixture model, and two matches can be combined to maximize the likelihood of the resulting interpolated match. As another example, the distance metric (D) computed during the comparison can be used to construct inverse-proportional weights applied to the motion data of the various input devices. In this example, the external poses can be fitted to a human body kinematics model based on the joint probabilities. In some aspects, synthesizing a style motion can further include interpolating a complete (or partial) body motion / pose from the feature data. For example, continuing from the above example, a user can simulate a tennis serve while using an HMD and hand controllers of a VR system. Based on the input data from the HMD and hand controllers, a motion style can be selected from a store of motion data. The motion data can include head and hand motion data as well as torso and lower body motion data for a tennis serve. Thus, a full body style motion for a tennis serve can be selected for the input data.

[0056] At operation 310, post-processing transformations (e.g., scaling) can be applied to generate avatar motion. In aspects, generated motion (interpolated from the closest motions from the motion dataset) can undergo a transformation to the VR space. Examples of transformations can include one or more motion scaling techniques, such as one-to-one mapping, position-based scaling, trajectory-based scaling, region-based scaling, or some combination thereof. In examples, scaling can enable a user to maintain a low motion footprint when performing motions in a spatially constrained environment, or to avoid fatigue when performing large-scale or extensive motions. For example, continuing the example above, a user of a VR application can simulate a tennis serve. Due to spatial constraints of the user’s physical environment (such as a low ceiling), the user can not be able to fully stretch her hand upwards during the simulation of the tennis serve. Upon receiving input data for the simulation of the tennis serve, a motion scaling component can apply one or more adaptive motion scaling techniques to the received input data. Based on the scaling techniques, position information in the input data can be scaled such that the user’s physical constraints are not applied to the user’s avatar. That is, the user’s avatar’s hand can fully stretch upwards when the avatar simulates the tennis serve.

[0057] At operation 312, stylized motion can be assigned to a virtual avatar by the virtual reality system. In aspects, stylized motion can be applied to one or more objects (components 210) in the VR application. For example, stylized motion can be applied to a user’s avatar. As a result, the avatar can move in the new stylized motion in a manner that fully expresses the avatar’s motion. That is, in a manner that does not indicate physical limitations of the user or the user’s physical environment.

[0058] Figures 4-7 The descriptions of the aspects herein provide illustrations and discussions of various operational environments in which aspects of the disclosure can be practiced. However, the devices and systems illustrated and discussed are for purposes of example and illustration and are not limiting of a vast number of computing device configurations that can be utilized for practicing aspects of the disclosure described herein. Figures 4-7 The devices and systems illustrated and discussed are for purposes of example and illustration and are not limiting of a vast number of computing device configurations that can be utilized for practicing aspects of the disclosure described herein.

[0059] Figure 4 is a block diagram illustrating physical components (e.g., hardware) of a computing device 400 with which aspects of the disclosure can be practiced. The computing device components described below can be suitable for the computing devices described above, including the computing device 102 and the computing device 104, as well as the virtual environment service 106. In a basic configuration, the computing device 400 can include at least one processing unit 402 and a system memory 404. Depending on the configuration and type of computing device, the system memory 404 can comprise, but is not limited to, volatile (e.g. random access memory (RAM)), non-volatile (e.g. read-only memory (ROM)), flash memory, or any combination. The system memory 404 can include an operating system 405 and one or more program modules 406, which can include the virtual environment service 106 and the virtual environment client 108.

[0060] System memory 404 can include operating system 405 and one or more program modules 406 suitable for running software applications 420, such as one or more components supported by the systems described herein. By way of example, system memory 404 can be virtual environment application 424 and text web component 426. The operating system 405, for example, can be suitable for controlling the operation of the computing device 400.

[0061] Furthermore, embodiments of the subject disclosure can be practiced in conjunction with a graphics library, other operating systems, or any other application program and are not limited to any particular Figure 4 application or system. This basic configuration is illustrated in Figure 4 by those components within dashed line 408. The computing device 400 can have additional features or functionality. For example, the computing device 400 can also include additional data storage devices (removable and / or non-removable) such as, for example, magnetic disks, optical disks, or tape. Such additional storage is illustrated in

[0062] As stated above, a number of program modules and data files can be stored in the system memory 404. While executing on the processing unit 402, the program modules 406 (e.g., an application 420) can perform processes including, for example, one or more of the aspects described herein. Other program modules that can be used in accordance with aspects of the present disclosure can include electronic mail and contacts applications, word processing applications, spreadsheet applications, database applications, slide presentation

[0063] application programs, etc. The system memory 404 can also store data that can be operated on by the one or more program modules 406, such as a virtual environment application 424 and a text web component 426. Figure 4 application functions, all of which are integrated (or "burned") onto the chip substrate as a single integrated circuit. When operating via the SOC, the functionality described herein with respect to the capabilities of the client switching protocol can be operated via application-specific logic integrated with other components of the computing device 400 on the single integrated circuit (chip). Embodiments of the subject disclosure can also be practiced using other technologies capable of performing logical operations such as, for example, AND, OR, and NOT, including but not limited to mechanical, optical, fluidic, and quantum technologies. In addition, embodiments of the subject disclosure can be practiced within a general purpose computer or in any other circuits or systems.

[0064] The computing device 400 can also have one or more input device(s) 412 such as a keyboard, a mouse, a pen, a sound or voice input device, a touch or swipe input device, etc. Output device(s) 414 such as a display, speakers, a printer, etc. can also be included. The aforementioned devices are examples and others can be used. The computing device 400 can include one or more communication connections 416 allowing communications with other computing devices 450. Examples of suitable communication connections 416 include, but are not limited to, radio frequency (RF) transmitter, receiver, and / or transceiver circuitry; universal serial bus (USB), parallel, and / or serial ports.

[0065] The term computer readable media as used herein can include computer storage media. Computer storage media can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, or program modules. The system memory 404, the removable storage device 409, and the non-removable storage device 410 are all computer storage media examples (e.g., memory storage). Computer storage media can include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article of manufacture which can be used to store information and which can be accessed by the computing device 400. Any such computer storage media can be part of the computing device 400. Computer storage media does not include a carrier wave or other propagated or modulated data signal.

[0066] Communication media can be embodied by computer readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. According to various embodiments, the term "modulated data signal" can describe a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0067] Figure 5A and 5B Fig. illustrates a mobile computing device 500, for example, a mobile telephone, a smart phone, a wearable computer (such as a smart watch), a tablet computer, a laptop computer, etc., with which embodiments of the present disclosure can be practiced. In some aspects, a client can be a mobile computing device. Reference is made to Fig. 1 for further details of a mobile computing device. Figure 5AFIG. 13 is a block diagram illustrating the architecture of one aspect of a mobile computing device. That is, FIG. 13 illustrates one aspect of a mobile computing device 500 that can implement aspects. In a basic configuration, mobile computing device 500 is a handheld computer that includes both input elements and output elements. The mobile computing device 500 typically includes a display 505 and one or more input buttons 510 that allow the user to enter information into the mobile computing device 500. The display 505 of the mobile computing device 500 can also function as an input device (e.g., a touch screen display).

[0068] Optional side input elements 515 (if included) allow further user input. The side input elements 515 can be a rotary switch, a button, or any other type of manual input element. In alternative aspects, mobile computing device 500 can incorporate more or less of each of the above- described input elements. For example, in some embodiments, the display 505 can not be a touch screen.

[0069] In yet another alternative embodiment, the mobile computing device 500 is a portable phone system, such as a cellular phone. The mobile computing device 500 can also include an optional keypad 535. Optional keypad 535 can be a physical keypad or a "soft" keypad generated on the touch screen display.

[0070] In various embodiments, output elements include the display 505 for showing a graphical user interface (GUI), a visual indicator 520 (e.g., a light emitting diode), and / or an audio transducer 525 (e.g., a speaker). In some aspects, the mobile computing device 500 incorporates input and / or output ports, such as an audio input (e.g., a microphone jack), an audio output (e.g., a headphone jack), and a video output (e.g., a HDMI port) for sending signals to or receiving signals from an external device.

[0071] Figure 5B is a block diagram illustrating the architecture of one aspect of a mobile computing device. That is, the mobile computing device 500 can incorporate a system (e.g., an architecture) 502 to implement some aspects. In one embodiment, the system 502 is implemented as a "smart phone" capable of running one or more applications (e.g., browser, e-mail, calendars, contact managers, messaging clients, games, and media clients / playback). In some aspects, the system 502 is integrated as a computing device, such as an integrated personal digital assistant (PDA) and wireless telephone.

[0072] One or more application programs 566 can be loaded into the memory 562 and run on or in association with the operating system 564. Examples of the application programs include phone dialer programs, e-mail programs, personal information management (PIM) programs, word processing programs, spreadsheet programs, Internet browser programs, messaging programs, and so forth. The system 502 also includes a non-volatile memory area 568 within the memory 562. The non-volatile memory area 568 can be used for storing persistent information that should not be lost if the system 502 powers off. The application programs 566 can use and store information in the non-volatile memory area 568, such as e-mail or other messages used by an e-mail application, and so forth. A synchronization application (not shown) also resides on the system 502 and is programmed to interact with a corresponding synchronization application resident on a host computer to keep the information stored in the non-volatile memory area 568 synchronized with corresponding information stored at the host computer. It should be appreciated that other applications can be loaded into the memory 562 and run on the mobile computing device 500 described herein (e.g., a search engine, an extractor module, a relevance ranking module, an answer scoring module, and so forth).

[0073] The system 502 has a power supply 570, which can be implemented as one or more batteries. The power supply 570 might further include an external power source, such as an AC adapter or a powered docking cradle that supplements or recharges the batteries.

[0074] The system 502 might also include a radio interface layer 572 that works with the radio interface 554 to facilitate wireless communication with a network and any devices coupled with the network. The radio interface layer 572 provides the actual transmit and receive functions for the transmitting and receiving of radio frequency signals to and from the radio interface 554. Transmissions to and from the radio interface layer 572 are conducted under the control of the operating system 564. In other words, communications received by the radio interface layer 572 can be disseminated to the application programs 566 via the operating system 564, and vice versa.

[0075] The visual indicator 520 can be used to provide visual notifications, and / or an audio interface 574 can be used for producing audible notifications via the audio transducer 525. In the illustrated embodiment, the visual indicator 520 is a light emitting diode (LED) and the audio transducer 525 is a speaker. These devices can be directly coupled to the power supply 570 so that when activated, they remain on for a duration dictated by the notification mechanism even though the processor 560 and other components might shut down for conserving battery power. The LED can be programmed to remain on indefinitely until the user takes action to indicate that the device is powered on. The audio interface 574 is used to provide audible signals to and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 525, the audio interface 574 can also be coupled to a microphone to receive audible input, such as to facilitate a telephone conversation. In accordance with embodiments of the present disclosure, the microphone can also serve as an audio sensor to facilitate control of notifications, as will be described below. The system 502 can further include a video interface 576 that enables an operation of a camera 530 to record still images, video stream, and the like.

[0076] A mobile computing device 500 implementing the system 502 can have additional features or functionality. For example, the mobile computing device 500 can also include additional data storage devices (removable and / or non-removable) such as, magnetic disks, optical disks, or tape. Such additional storage is illustrated in Figure 5B by the non-volatile storage area 568.

[0077] Data / information generated or captured by the mobile computing device 500 and stored via the system 502 can be stored locally on the mobile computing device 100, as described above, or the data can be stored on any number of storage media that is accessible by the device via the radio interface layer 572 or via a wired connection between the mobile computing device 500 and a separate computing device associated with the mobile computing device 500, for example, a server computer in a distributed computing network, such as the Internet. As

[0078] Figure 6One aspect of a system architecture for processing data received at a computing system from a remote source, such as a personal computer 604, a tablet computing device 606, or a mobile computing device 608, is illustrated, as described above. Content displayed at the server device 602 can be stored in different communication channels or other storage types. For example, a directory service 622, a web portal 624, a mailbox service 626, an instant messaging store 628, or a social networking site 630 can be used to store various documents.

[0079] A virtual environment application 620 can be employed by a client in communication with the server device 602, and / or a virtual environment data store 621 can be employed by the server device 602. The server device 602 can provide data to and receive data from client computing devices, such as the personal computer 604, the tablet computing device 606, and / or the mobile computing device 608 (e.g., a smart phone) over the network 615. By way of example, the computer system described above can be embodied in the personal computer 604, the tablet computing device 606, and / or the mobile computing device 608 (e.g., a smart phone). Any of these embodiments of the computing device can obtain content from the storage 616, in addition to receiving graphics data that can be used for pre-processing at a graphics initiation system or post-processing at a receiving computing system.

[0080] Figure 7 An exemplary tablet computing device 700 that can perform one or more of the aspects disclosed herein is illustrated. Moreover, the aspects and functionalities described herein can operate over distributed systems (e.g., cloud-based computing systems), where various devices and modules are coupled via a distributed network, such as the Internet or an intranet. Various types of user interfaces and information are presented via on-board computing devices, such as a tablet computing device display, or via remote display units associated with one or more computing devices. For instance, user interfaces and information can be presented via a wall surface display, a projector, a television, a computer screen, a head-mounted display (e.g., a virtual reality display), an augmented reality display, an electronic ink display, an electronic paper display, a wearable device display, a remote projection hub, or a mobile device display. Also, user interfaces and information can be presented via a remote computing device, such as a server computer, a tablet computing device, a mobile computing device, or a wearable computing device. Accordingly, one or more aspects of the present disclosure can be embodied in various types of user interfaces and information, and in various types of software implementations, including those accessible via a web browser, a mobile application, a desktop application, or a stand-alone software package.

[0081] For example, aspects of the disclosure are described above with reference to flowchart illustrations and / or operational illustrations of methods, systems, and computer program products according to aspects of the disclosure. As it will be appreciated by those skilled in the art, the order of the blocks in the flowcharts, as well as the order of the blocks in the operational illustrations, can be changed. For example, two blocks shown in succession can be executed substantially concurrently or the blocks can be executed in the reverse order, depending upon the functionality / acts involved.

[0082] The description and drawings of one or more aspects provided herein are not intended to limit or restrict the scope of the disclosure to such precise system configurations. The aspects, examples, and details provided herein are considered sufficient to convey possession of the general inventive concepts. The application should not be construed as limited to any aspect, example, or detail provided herein. Various features (structural and methodological) are intended to be selectively included or omitted to produce an embodiment within the scope of the claimed disclosure. Having described the application in terms of embodiments, reference is made to the claims for determining the scope of the application.

Claims

1. A system comprising: One or more processors; as well as A memory coupled to at least one of the one or more processors, the memory including computer-executable instructions that, when executed by the at least one processor, perform a method comprising: Receive real-time input data corresponding to user movements in the virtual reality system; Extract feature data from the input data; The feature data is compared with stored motion data, wherein the comparison includes identifying a set of motions in the stored motion data, wherein the set of motions includes at least a first motion corresponding to a first body part and a second motion corresponding to a second body part, the second body part being different from the first body part; The motion sets are blended to generate stylized motion; A transformation is applied to the stylized motion, wherein the transformation maps the stylized motion to a virtual reality space; and The user avatar is manipulated based on the transformed stylized motion.

2. The system of claim 1, wherein the virtual reality system is accessed using a head-mounted display device.

3. The system according to claim 1, wherein the input data is at least one of the following: motion data, audio data, text data, eye-tracking data, 3D point cloud, depth data, or biosignals.

4. The system of claim 1, wherein the input data is collected from two or more input devices, each of the input devices comprising one or more sensor components.

5. The system according to claim 4, wherein the feature data includes at least one of the following: acceleration information, velocity information, or position information of the one or more input devices.

6. The system of claim 1, wherein generating the feature data includes converting the input data to a head-normalized coordinate system.

7. The system of claim 1, wherein the stored motion data includes actions performed by a first motion capture subject and a second motion capture subject, the first motion capture subject performing the actions in a first style, and the second motion capture subject performing the actions in a second style.

8. The system of claim 1, wherein comparing the feature data with the stored motion data comprises: The first feature vector associated with the feature data is compared with one or more feature vectors associated with the stored motion data.

9. The system of claim 1, wherein comparing the feature data with the stored motion data comprises: One or more matching algorithms are used to identify one or more candidate matches for the feature data.

10. The system of claim 9, wherein the one or more matching algorithms include at least one of the following: k-nearest neighbors, logistic regression, Naive Bayes classifier, support vector machine, random forest, or neural network.

11. The system of claim 10, wherein matching the feature data to the stylized motion comprises: Stylized motion is synthesized from one or more candidate matches.

12. The system of claim 11, wherein synthesizing the stylized motion includes constructing weights inversely proportional to the Euclidean distance between a first feature vector of the feature data and a second feature vector of the stored motion data.

13. The system of claim 1, wherein applying the transformation includes applying at least one of the following: motion scaling, motion filtering, or motion smoothing.

14. The system of claim 13, wherein the motion scaling includes using at least one of the following: one-to-one mapping, position-based scaling, trajectory-based scaling, or region-based scaling.

15. The system of claim 1, wherein the stylized motion is applied to the user's virtual avatar such that the stylized motion mimics the user interaction.

16. A method comprising: The virtual environment system receives user motion data corresponding to activities performed by the user, wherein the input data is detected using one or more input devices associated with the virtual environment system; Use the input data to generate feature data; The feature data is compared with stored motion data, wherein the comparison includes identifying a set of motions in the stored motion data, wherein the set of motions includes at least a first motion corresponding to a first body part and a second motion corresponding to a second body part, the second body part being different from the first body part; The motion sets are blended to generate stylized motion; A transformation is applied to the stylized motion, wherein the transformation maps the stylized motion to a virtual reality space; as well as The user avatar is operated in the virtual environment system according to the transformed stylized movement.

17. The method of claim 16, wherein generating feature data includes creating a feature vector representing the feature data, the feature vector including at least one of the following associated with the input data: acceleration information, velocity information, or position information.

18. The method of claim 16, wherein comparing the feature data with the stored motion data comprises: The feature data is classified into activity types; and Search the stored motion data for the activity type.

19. The method of claim 16, wherein the stored motion data includes styled motion data for at least one of: an expert in the activity, a professional in the activity, or a well-known entity.

20. A virtual environment system, comprising: One or more processors; as well as A memory coupled to at least one of the one or more processors, the memory including computer-executable instructions that, when executed by the at least one processor, perform a method comprising: User motion data is received from the user performing the activity, wherein the user motion data is detected using one or more input devices associated with the virtual environment system; The user motion data is used to generate feature data, wherein the feature data includes at least one of the following associated with the activity: acceleration information, velocity information, or position information; The feature data is compared with stored motion data, wherein the comparison includes identifying a set of motions in the stored motion data, wherein the set of motions includes at least a first motion corresponding to a first body part and a second motion corresponding to a second body part, the second body part being different from the first body part; The motion sets are blended to generate stylized motion; A transformation is applied to the stylized motion, wherein the transformation maps the stylized motion to the virtual space of the virtual environment system; and The transformed, stylized motion is performed in the virtual space.

Citation Information

Patent Citations

  • Interacting with user interface through metaphoric body

    CN102221886A

  • Human motion capture device

    CN102323854A