Virtual reality management device and virtual reality management method

The virtual reality management device uses a movement recognition model to differentiate between long-range and short-range movements based on arm and hand gestures, addressing precision and multitasking challenges in VR environments, enhancing immersion and reducing motion sickness.

WO2025238676A1PCT designated stage Publication Date: 2025-11-20HITACHI SYST LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/017593
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-13
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Existing VR technologies face challenges in facilitating precise long-range and short-range movement within virtual environments, particularly in multitasking scenarios, leading to inefficiencies and motion sickness, especially in complex environments like Virtual Training Environments (VTEs).

Method used

A virtual reality management device and method utilizing a movement recognition model to analyze user arm and hand gestures, distinguishing between long-range and short-range movements, enabling precise teleportation and walking/running through arm swinging motions, and allowing multitasking like object manipulation.

Benefits of technology

Enables seamless transitions between long-distance locomotion and short-distance movements, reducing motion sickness and enhancing immersion by integrating natural arm dynamics into virtual environments, supporting complex multitasking actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024017593_20112025_PF_FP_ABST
    Figure JP2024017593_20112025_PF_FP_ABST
Patent Text Reader

Abstract

Aspects relate to a virtual reality management technique for facilitating precise long-range and short-range movement in a virtual environment while enabling complex multi-task actions. The virtual reality management technique includes determining, by analyzing the set of video data using a movement recognition model, a short-range movement action for an avatar in the virtual reality environment in a case that a set of movement factors detected for both a first arm and a second arm of the user satisfy a predetermined movement threshold and executing the short-range movement action, determining, by analyzing the set of video data using the movement recognition model, a long-range movement action for the avatar in the virtual reality environment in a case that a first hand of the user corresponds to a pointing gesture and a confirmation action is received from the user, and executing the long-range movement action.
Need to check novelty before this filing date? Find Prior Art

Description

VIRTUAL REALITY MANAGEMENT DEVICE AND VIRTUAL REALITY MANAGEMENT METHOD

[0001] The present disclosure relates to a virtual reality management device and a virtual reality management method.

[0002] In recent years, virtual reality (VR) technology, in which virtual-reality headsets or multi-projected environments are used to simulate a user’s physical presence in a virtual environment, are growing in popularity. Generally, VR technology is implemented using a virtual-reality headset having a head-mounted display comprising a small screen positioned in front of the eyes. Various modalities exist for allowing users to interact with the VR environment. For example, some systems can employ one or more of gaze controls, hand-held hardware devices, gesture controls, wearable devices (e.g., wrist bands), voice controls, or the like.

[0003] In some cases, the physical spaces in which users reside (e.g., rooms inside users’ homes) may be smaller than the virtual spaces presented via VR technology. This poses challenges for achieving realistic locomotion (e.g., walking, running) of users within VR environments.

[0004] Conventionally, techniques for handling the movement of user avatars within VR environments have been considered. As an example, US Patent 11,461,973 (Patent Literature 1) discloses “Embodiments described herein disclose methods and systems directed to locomotion in virtual reality (VR) based on hand gestures of a user. In some implementations, the user can navigate between locations using hand gestures that trigger teleportation. In other implementations, the user can separately control forward / backward movement and the direction orientation the user is facing. The separate control can either be with one hand when making different gestures or by using different hands to control movement and orientation. In some implementations, a dragging gesture by the user is interpreted by the VR system to trigger movement. In this way, the user can turn in place without moving forward, move forward without turning, or can move forward and turn, while controlling speed, all with a single gesture.”

[0005] US Patent 11,461,973

[0006] Patent Literature 1 discloses a technique for gesture-based locomotion control in a VR environment. For example, Patent Literature 1 describes a user making a first hand gesture to set a first origin point and a second hand gesture to set a destination point. Upon releasing the second hand gesture, the user is transported to the destination with an orientation direction determined by the second hand gesture.

[0007] While Patent Literature 1 discloses a technique for allowing users to change orientation and teleport to a desired location in the VR environment using hand gestures, it does not take into consideration the navigational challenges associated with multitasking (e.g., holding virtual objects while performing complex hand gestures for movement) and precision for accurate displacements. As a result, the efficacy of the technique disclosed in Patent Literature 1 is limited with respect to virtual environments that require more complex object manipulation and precise displacement, such as Virtual Training Environments (VTEs). Similarly, other conventional techniques for controlling locomotion in virtual environments, such as controller-based movement or additional wearable sensors, face challenges such as motion sickness in the case of controller-based approaches or increased cost in the case of additional wearable sensors.

[0008] Accordingly, it is an object of the present disclosure to provide a virtual reality management technique for facilitating precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions.

[0009] One representative example of the present disclosure relates to a virtual reality management device comprising: a set of cameras for collecting a set of video data of a real-world environment around a user; a display for displaying a virtual reality environment including an avatar representing the user; a processor; a memory; and a storage unit, wherein: the storage unit includes: a movement recognition model trained to identify body movements of the user in the set of video data; the memory includes a set of computer-readable instructions that cause the processor to: determine, by analyzing the set of video data using the movement recognition model, a short-range movement action for the avatar in the virtual reality environment in a case that a set of movement factors detected for both a first arm and a second arm of the user satisfy a predetermined movement threshold; execute the short-range movement action by moving the avatar in the virtual reality environment at a speed calculated based on the set of movement factors; determine, by analyzing the set of video data using the movement recognition model, a long-range movement action for the avatar in the virtual reality environment in a case that a first hand of the user corresponds to a pointing gesture and a confirmation action is received from the user; and execute the long-range movement action by moving the avatar in the virtual reality environment to a destination location indicated by the pointing gesture.

[0010] According to the present disclosure it is possible to provide a virtual reality management technique for facilitating precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions.

[0011] Problems, configurations, and effects other than those described above will be made clear by the following description in the embodiments for carrying out the invention.

[0012] FIG. 1 illustrates an example computing architecture for executing the embodiments of the present disclosure.FIG. 2 is a diagram illustrating an example hardware configuration of a virtual reality management system according to the embodiments of the present disclosure.FIG. 3 is a flowchart illustrating a movement action management method according to the embodiments of the present disclosure.FIG. 4 is a diagram illustrating an example of training and using the movement recognition model according to the embodiments of the present disclosure.FIG. 5 is a diagram illustrating an example of a first phase of the hybrid grasp / long-range movement action according to the embodiments of the present disclosure.FIG. 6 is a diagram illustrating an example of a second phase of the hybrid grasp / long-range movement action according to the embodiments of the present disclosure.FIG. 7 is a diagram illustrating an example of a third phase of the hybrid grasp / long-range movement action according to the embodiments of the present disclosure.FIG. 8 is a diagram illustrating an example of capturing video data for movement recognition model training, according to embodiments of the present disclosure.FIG. 9 is a diagram illustrating an example of the flow of training the movement recognition model for short-range movement gesture detection.FIG. 10 is a diagram illustrating an example of the movement factors used for detecting the short-range movement gesture according to the embodiments of the present disclosure.FIG. 11 is a diagram illustrating an example of determining displacement in the virtual environment when performing the short-range movement action, according to the embodiments of the present disclosure.FIG. 12 is a diagram illustrating an example of the hybrid grasp / short-range movement action according to the embodiments of the present disclosure.FIG. 13 is a flowchart illustrating a training and use method of the virtual reality management device according to the embodiments of the present disclosure.FIG. 14 is a flowchart illustrating an example of a training process of the movement recognition model of the virtual reality management device according to the embodiments of the present disclosure.Description of Embodiment(s)

[0013] Herein, embodiments of the present invention will be described with reference to the Figures. It should be noted that the embodiments described herein are not intended to limit the invention according to the claims, and it is to be understood that each of the elements and combinations thereof described with respect to the embodiments are not strictly necessary to implement the aspects of the present invention.

[0014] Various aspects are disclosed in the following description and related drawings. Alternate aspects may be devised without departing from the scope of the disclosure. Additionally, well-known elements of the disclosure will not be described in detail or will be omitted so as not to obscure the relevant details of the disclosure.

[0015] The words “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term “aspects of the disclosure” does not require that all aspects of the disclosure include the discussed feature, advantage or mode of operation.

[0016] Further, many aspects are described in terms of sequences of actions to be performed by, for example, elements of a computing device. It will be recognized that various actions described herein can be performed by specific circuits (e.g., an application specific integrated circuit (ASIC)), by program instructions being executed by one or more processors, or by a combination of both. Additionally, the sequence of actions described herein can be considered to be embodied entirely within any form of computer readable storage medium having stored therein a corresponding set of computer instructions that upon execution would cause an associated processor to perform the functionality described herein. Thus, the various aspects of the disclosure may be embodied in a number of different forms, all of which have been contemplated to be within the scope of the claimed subject matter.

[0017] Hereinafter, a detailed description of the embodiments of the present disclosure will be described with reference to the Figures.

[0018] (Summary of the Invention) Aspects of the present disclosure relate to facilitating navigation in a virtual environment using long-range (e.g., teleportation) and short-range (e.g., running / walking) movement modes based on hand-tracked gestures. A movement recognition model distinguishes between gestures for long-range and short-range movement, and refines avatar movements for enhanced immersion and reduced motion sickness. Users can perform complex multi-tasking actions like grasping virtual objects during navigation without using a physical controller.

[0019] More particularly, by utilizing a movement recognition model for hand and arm movement recognition, it is possible to enable smooth transitions between long-distance locomotion methods, such as teleportation, and short-distance techniques, like walking / running. Further, utilizing two-hand gestures for long-distance teleportation facilitates differentiation between teleportation and other VR actions. Additionally, aspects of the disclosure relate to supporting multi-tasking to allow users to teleport while holding virtual objects in one hand, for example. Precise short-range displacements can be achieved based on recognition of natural arm swinging (arm-walking) motions. In the present disclosure, “arm walking” refers to a user moving their arms to simulate naturalistic arm swinging motions while remaining in place.

[0020] One objective of the present disclosure relates to enhancing immersion and practicality in virtual reality by utilizing pre-recorded arm movement in relation to the user's body that is collected by a front-facing camera during a pre-training phase. Virtual avatar displacement can be performed based on the arm swing frequency, arm joint angles, and the relative arm positions of the user.

[0021] To refine long-range and short-range movement mode gesture detection, a pre-trained AI module may be utilized to leverage a diverse dataset of other users’ walking patterns. This data may then correlated with the present users’ recorded walking pattern and speed, enabling the system of the present disclosure to intelligently adapt its virtual displacement calculations realistically. This collaborative learning approach enhances the system's accuracy and versatility, accommodating different users' natural walking behaviors.

[0022] Furthermore, users can adjust their walking speed in the virtual environment with an intuitive UI presented in the virtual space. As a result, users can experience a more intuitive and personalized VR interaction, as the system seamlessly integrates their arm dynamics and walking behaviors into the virtual environment, fostering a realistic and engaging VR experience.

[0023] Turning now to the Figures, FIG. 1 depicts a high-level block diagram of a computer system 100 for implementing various embodiments of the present disclosure, according to embodiments. The mechanisms and apparatus of the various embodiments disclosed herein apply equally to any appropriate computing system. The major components of the computer system 100 include one or more processors 102, a memory 104, a terminal interface 112, a storage interface 113, an I / O (Input / Output) device interface 114, and a network interface 115, all of which are communicatively coupled, directly or indirectly, for inter-component communication via a memory bus 106, an I / O bus 108, bus interface unit 109, and an I / O bus interface unit 110.

[0024] The computer system 100 may contain one or more general-purpose programmable central processing units (CPUs) 102A and 102B, herein generically referred to as the processor 102. In embodiments, the computer system 100 may contain multiple processors; however, in certain embodiments, the computer system 100 may alternatively be a single CPU system. Each processor 102 executes instructions stored in the memory 104 and may include one or more levels of on-board cache.

[0025] In embodiments, the memory 104 may include a random-access semiconductor memory, storage device, or storage medium (either volatile or non-volatile) for storing or encoding data and programs. In certain embodiments, the memory 104 represents the entire virtual memory of the computer system 100, and may also include the virtual memory of other computer systems coupled to the computer system 100 or connected via a network. The memory 104 can be conceptually viewed as a single monolithic entity, but in other embodiments the memory 104 is a more complex arrangement, such as a hierarchy of caches and other memory devices. For example, memory may exist in multiple levels of caches, and these caches may be further divided by function, so that one cache holds instructions while another holds non-instruction data, which is used by the processor or processors. Memory may be further distributed and associated with different CPUs or sets of CPUs, as is known in any of various so-called non-uniform memory access (NUMA) computer architectures.

[0026] The memory 104 may store all or a portion of the various programs, modules and data structures for processing data transfers as discussed herein. For instance, the memory 104 can store a virtual reality management application 150. In embodiments, the virtual reality management application 150 may include instructions or statements that execute on the processor 102 or instructions or statements that are interpreted by instructions or statements that execute on the processor 102 to carry out the functions as further described below. In certain embodiments, the virtual reality management application 150 is implemented in hardware via semiconductor devices, chips, logical gates, circuits, circuit cards, and / or other physical hardware devices in lieu of, or in addition to, a processor-based system. In embodiments, the virtual reality management application 150 may include data in addition to instructions or statements. In certain embodiments, a camera, sensor, or other data input device (not shown) may be provided in direct communication with the bus interface unit 109, the processor 102, or other hardware of the computer system 100. In such a configuration, the need for the processor 102 to access the memory 104 and the virtual reality management application 150 may be reduced.

[0027] The computer system 100 may include a bus interface unit 109 to handle communications among the processor 102, the memory 104, a display system 124, and the I / O bus interface unit 110. The I / O bus interface unit 110 may be coupled with the I / O bus 108 for transferring data to and from the various I / O units. The I / O bus interface unit 110 communicates with multiple I / O interface units 112, 113, 114, and 115, which are also known as I / O processors (IOPs) or I / O adapters (IOAs), through the I / O bus 108. The display system 124 may include a display controller, a display memory, or both. The display controller may provide video, audio, or both types of data to a display device 126. Further, the computer system 100 may include one or more sensors or other devices configured to collect and provide data to the processor 102. As examples, the computer system 100 may include biometric sensors (e.g., to collect heart rate data, stress level data), environmental sensors (e.g., to collect humidity data, temperature data, pressure data), motion sensors (e.g., to collect acceleration data, movement data), or the like. Other types of sensors are also possible. The display memory may be a dedicated memory for buffering video data. The display system 124 may be coupled with a display device 126, such as a standalone display screen, computer monitor, television, or a tablet or handheld device display. In one embodiment, the display device 126 may include one or more speakers for rendering audio. Alternatively, one or more speakers for rendering audio may be coupled with an I / O interface unit. In alternate embodiments, one or more of the functions provided by the display system 124 may be on board an integrated circuit that also includes the processor 102. In addition, one or more of the functions provided by the bus interface unit 109 may be on board an integrated circuit that also includes the processor 102.

[0028] The I / O interface units support communication with a variety of storage and I / O devices. For example, the terminal interface unit 112 supports the attachment of one or more user I / O devices 116, which may include user output devices (such as a video display device, speaker, and / or television set) and user input devices (such as a keyboard, mouse, keypad, touchpad, trackball, buttons, light pen, or other pointing device). A user may manipulate the user input devices using a user interface in order to provide input data and commands to the user I / O device 116 and the computer system 100, and may receive output data via the user output devices. For example, a user interface may be presented via the user I / O device 116, such as displayed on a display device, played via a speaker, or printed via a printer.

[0029] The storage interface 113 supports the attachment of one or more disk drives or direct access storage devices 117 (which are typically rotating magnetic disk drive storage devices, although they could alternatively be other storage devices, including arrays of disk drives configured to appear as a single large storage device to a host computer, or solid-state drives, such as flash memory). In some embodiments, the storage device 117 may be implemented via any type of secondary storage device. The contents of the memory 104, or any portion thereof, may be stored to and retrieved from the storage device 117 as needed. The I / O device interface 114 provides an interface to any of various other I / O devices or devices of other types, such as printers or fax machines. The network interface 115 provides one or more communication paths from the computer system 100 to other digital devices and computer systems; these communication paths may include, for example, one or more networks 130.

[0030] Although the computer system 100 shown in FIG. 1 illustrates a particular bus structure providing a direct communication path among the processors 102, the memory 104, the bus interface 109, the display system 124, and the I / O bus interface unit 110, in alternative embodiments the computer system 100 may include different buses or communication paths, which may be arranged in any of various forms, such as point-to-point links in hierarchical, star or web configurations, multiple hierarchical buses, parallel and redundant paths, or any other appropriate type of configuration. Furthermore, while the I / O bus interface unit 110 and the I / O bus 108 are shown as single respective units, the computer system 100 may, in fact, contain multiple I / O bus interface units 110 and / or multiple I / O buses 108. While multiple I / O interface units are shown which separate the I / O bus 108 from various communications paths running to the various I / O devices, in other embodiments, some or all of the I / O devices are connected directly to one or more system I / O buses.

[0031] In various embodiments, the computer system 100 is a multi-user mainframe computer system, a single-user system, or a server computer or similar device that has little or no direct user interface, but receives requests from other computer systems (clients). In other embodiments, the computer system 100 may be implemented as a desktop computer, portable computer, laptop or notebook computer, tablet computer, pocket computer, telephone, smart phone, or any other suitable type of electronic device.

[0032] Next, with reference to FIG. 2, an example hardware configuration of a virtual reality management system according to the embodiments of the present disclosure will be described.

[0033] FIG. 2 is a diagram illustrating an example hardware configuration of a virtual reality management system 200 according to the embodiments of the present disclosure. The virtual reality management system 200 relates to a system configured to facilitate precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions.

[0034] As illustrated in FIG. 2, the virtual reality management system 200 according to the embodiments of the present disclosure includes a virtual reality management device 210, a communication network 250, and a user terminal 260. The virtual reality management device 210 and the user terminal 260 may be communicatively connected via the communication network 250. Here, the communication network 250 may include a Local Area Network (LAN) connection, the Internet, a Wide Area Network (WAN) connection, a Metropolitan Area Network (MAN) connection, a Bluetooth (registered trademark) connection, or the like.

[0035] The virtual reality management device 210 is an apparatus for generating, providing, and managing VR content with respect to a user. In embodiments, the virtual reality management device 210 may be a headset configured to be worn on the head of the user and provide generated VR content via input / output units 246 comprising displays, speakers, and the like. As illustrated in FIG. 2, the virtual reality management device 210 primarily includes a memory 220, a storage unit 230, a processor 244, input / output units 246, and sensor units 248. In embodiments, aspects of the virtual reality management device 210 may be implemented using a computer system such as the computer system 100 illustrated in FIG. 1.

[0036] The memory 220 is a memory for storing a virtual reality management application 150 for implementing the functions of the virtual reality management technique according to the embodiments of the present disclosure. As illustrated in FIG. 2, the virtual reality management application 150 may include a VR generation unit 222, a movement recognition unit 224, a movement management unit 226, and a model training unit 228. The VR generation unit 222, the movement recognition unit 224, the movement management unit 226, and the model training unit 228 may be implemented as software modules included in the virtual reality management application 150.

[0037] The VR generation unit 222 is a functional unit for generating VR content for provision to the user of the virtual reality management device 210. The VR generation unit 222 may generate a virtual environment including images, sounds, and haptic feedback for output to the user via the input / output units 246 of the virtual reality management device 210. In embodiments, the VR generation unit 222 may generate a user avatar for representing the body of the user of the virtual reality management device 210 in the virtual environment, and the user may view the virtual environment from the first-person perspective of the avatar. As an example, the VR generation unit 222 may be configured to generate a virtual training environment (VTE) for facilitating occupational training for a particular task (e.g., machine maintenance work). The VR generation unit 222 is not particularly limited herein, and may be implemented using known VR content generation techniques.

[0038] The movement recognition unit 224 is a functional unit for recognizing user movements performed by the user of the virtual reality management device 210 in the real world environment. In the present disclosure, user movements refer to gestures, actions, or any type of body movement performed by the user. In embodiments, the movement recognition unit may be configured to make use of a movement recognition model (e.g., the movement recognition model 231) trained to identify various user movements such as open and closed hand gestures, pointing gestures, object grasping gestures, arm-swinging, extended arm postures, standing arm postures, walking postures and the like within captured video data. The user movements recognized by the movement recognition unit 224 may be used to facilitate implementation of short-range and long-range movement actions in the virtual environment.

[0039] The movement management unit 226 is a functional unit for implementing short-range and long-range movement actions as well as virtual-object grasping actions in the virtual environment based on the user movements detected by the movement recognition unit 224. In embodiments, the long-range movement action may include teleportation from an origin location to a destination location in the virtual environment, and may be performed in response to determining that a first hand of the user corresponds to a pointing gesture and receiving a confirmation action from the user. Additionally, in embodiments, the short-range movement action may include walking / running of the avatar representing the user in the virtual environment, and may be performed in response to detecting that a set of movement factors detected for both a first arm and a second arm of the user satisfy a predetermined movement threshold (e.g., the user is performing arm-walking). As the detailed functionality of the movement management unit 226 will be described later, it will be omitted here.

[0040] The model training unit 228 is a functional unit for training the movement recognition model 231 used to identify various user movements in captured video data. In embodiments, the model training unit 228 may be configured to collect a set of movement factors of the user of the virtual reality management device 210, generate a movement training dataset including the collected movement factors, and train the movement recognition model based on the movement training dataset. In this way, it is possible to train a movement recognition model 231 that is specifically customized to the movement factors of the user of the virtual reality management device 210. As the detailed functionality of the model training unit 228 will be described later, it will be omitted here.

[0041] The storage unit 230 is a unit for storing various data and information used in implementing the aspects of the present disclosure. The storage unit 230 may include a collection of hard disc drives, solid state drives, flash memory, cloud, storage, or the like. As illustrated in FIG. 2, the storage unit 230 may include a movement recognition model 231 and a wearer movement database 232.

[0042] The movement recognition model 231 is a machine learning model trained to identify various user movements such as open and closed hand gestures, pointing gestures, grabbing gestures, arm-swinging, extended arm postures, standing arm postures, walking postures and the like within captured video data. In embodiments, the movement recognition model 231 may be pre-trained on a dataset including movement factors collected from a number of users, and subsequently retrained by the model training unit 228 based on the movement factors of the user of the virtual reality management device 210. In this way, it is possible to train a movement recognition model 231 that is specifically customized to the movement characteristics of the user of the virtual reality management device 210.

[0043] The wearer movement database 232 is a database configured to store a set of wearer movement data. In embodiments, the set of wearer movement data may include the user movement factors collected by the model training unit 228 for training the movement recognition model 231. As described herein, the stored user movement factors (e.g., arm swinging frequency) may be used to calculate a walking speed for the avatar representing the user in the virtual reality environment.

[0044] The processor 244 is a processing unit for executing processing instructions for the various software modules and functional units included in the virtual reality management application 150 stored in the memory 220, and may substantially correspond to the processor 102 of the computer system 100 illustrated in FIG. 1.

[0045] The input / output units 246 are a collection of devices for facilitating communication between the virtual reality management device 210 and external sources, such as the user of the virtual reality management device 210 as well as one or more user terminals 260 to which the virtual reality management device 210 is communicably connected. In embodiments, the input / output units 246 may include a built-in display unit, speakers, haptic devices, and the like for providing virtual content to the user of the virtual reality management device 210. Further, the input / output units 246 may include microphones, touch screens, and eye-tracking cameras for receiving input from the user of the virtual reality management device 210.

[0046] The set of sensor units 248 are a collection of sensors for collecting information about the user of the virtual reality management device 210 as well as the real-world environment occupied by the user. In embodiments, the set of sensor units 248 may include a set of cameras for collecting a set of video data of the real-world environment around the user. The set of cameras may include front-facing cameras configured to capture video data for the real-world environment in front of the user as well as down-facing cameras configured to capture video data for the body (torso, arms, hands, legs) of the user. Further, the set of sensor units 248 may include motion sensors (e.g., to collect acceleration data, movement data), biometric sensors (e.g., to collect heart rate data, stress level data), environmental sensors (e.g., to collect humidity data, temperature data, pressure data), or the like. Other types of sensors are also possible.

[0047] The user terminal 260 is a device that can be used by a user (e.g., a client) of the virtual reality management device 210. In embodiments, the user terminal 260 may be used to configure the virtual reality management device 210, provide additional computing resources for generation of virtual reality content by the VR generation unit 222, facilitate training of the movement recognition model 231, manage and store video data captured by the cameras of the virtual reality management device 210, or the like. As examples, the user terminal 260 may be implemented using a personal computer, tablet computer, smart phone, or other computing device.

[0048] According to the virtual reality management system 200 illustrated in FIG. 2, it is possible to provide a virtual reality management technique for facilitating precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions.

[0049] Next, with reference to FIG. 3, a movement action management method according to the embodiments of the present disclosure will be described.

[0050] FIG. 3 is a flowchart illustrating a movement action management method 300 according to the embodiments of the present disclosure. As described herein, aspects of the present disclosure relate to providing a virtual reality management technique for facilitating precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions. Accordingly, the movement action management method 300 is a method for determining and executing long-range and short-range movement actions in the virtual environment based on detected user movements.

[0051] First, at Step S302, the movement recognition unit 224 uses the trained movement recognition model 231 to analyze a set of video data collected by the cameras of the sensor units 248 to identify user movements performed by the user of the virtual reality management device 210. For instance, the movement recognition unit 224 may use the trained movement recognition model 231 to identify whether the user is performing movements corresponding to a long-range movement action or a short-range movement action.

[0052] Next, at Step S304, the movement recognition unit 224 uses the trained movement recognition model 231 to identify whether the user movements correspond to a short-range movement action. Here, the short-range movement action may refer to walking or running of the user avatar in the virtual environment, and may be identified in response to detecting that the user is performing an arm-walking gesture in the real world environment. Here, arm-walking refers to the user moving their arms in a swinging motion (e.g., as the user would when walking) while remaining stationary. More particularly, as described herein, the movement recognition unit 224 may identify that the user movements correspond to arm-walking in the case that a set of movement factors (e.g., a first set of arm joint angles between a torso of the user and the first arm of the user, a second set of arm joint angles between the torso of the user and the second arm of the user, an arm swing frequency, a walking cadence of the user) detected for both a first arm and a second arm of the user satisfy a predetermined movement threshold. In the case that the movement recognition unit 224 identifies user movements corresponding to the short-range movement action, the method proceeds to Step S316. In the case that the movement recognition unit 224 does not identify user movements corresponding to the short-range movement action, the method proceeds to Step S306. In the present disclosure, user movements detected by the movement recognition unit as corresponding to the short-range movement action may be referred to as short-range movement gestures, and user movements detected by the movement recognition unit as corresponding to the long-range movement action may be referred to as long-range movement gestures.

[0053] Next, at Step S306, the movement recognition unit 224 uses the trained movement recognition model 231 to identify whether the user movements correspond to a long-range movement action. Here, the long-range movement action may refer to teleportation of the user avatar in the virtual environment from a source location to a destination location, and may be identified in response to detecting that the user is performing a pointing gesture (e.g., extending the arm and one or more fingers of the hand toward a desired destination location) in the real world environment. Additionally, the user may perform a confirmation action in addition to the pointing gesture to indicate intent to perform the long-range movement action. Here, the confirmation action may include a gesture with another hand (e.g., a fist gesture, an open-hand gesture), a voice command, two rapid eye blinks performed within a certain time frame (e.g., 1 second, 2 seconds), a wink (e.g., a single eye blink) or the like. In the case that the movement recognition unit 224 identifies user movements corresponding to the long-range movement action, the method proceeds to Step S308. In the case that the movement recognition unit 224 does not identify user movements corresponding to the long-range movement action, the method returns to Step S302 to continue video analysis.

[0054] Next, at Step S308, in the case that the movement recognition unit 224 identifies user movements corresponding to the long-range movement action, the movement management unit 226 determines to perform the long-range movement action with respect to the user avatar in the virtual environment.

[0055] Next, at Step S310, the movement recognition unit 224 uses the trained movement recognition model 231 to identify whether the user movements correspond to an object grasping action. Here, the object-grasping action refers to the act of picking-up, holding, or gripping a virtual object in the virtual environment, and may be identified in response to detecting that the user is performing an object grasping gesture (e.g., curling the thumb and one or more fingers of the hand) in the real-world environment. In the case that the movement recognition unit 224 does not identify user movements corresponding to the object-grasping action, the method proceeds to Step S311. In the case that the movement recognition unit 224 identifies user movements corresponding to the object-grasping action, the method proceeds to Step S312.

[0056] At Step S311, in the case that the movement recognition unit 224 does not identify user movements corresponding to the object-grasping action, the movement management unit 226 performs the long-range movement action determined at Step S308. For example, the movement management unit 226 may perform the long-range movement action by teleporting the avatar of the user to the location in the virtual environment indicated by the pointing gesture of the user.

[0057] At Step S312, in the case that the movement recognition unit 224 identifies user movements corresponding to the object-grasping action at Step S310, the movement management unit 226 performs a hybrid grasp / long-range movement action with respect to the avatar of the user in the virtual environment. Here, the hybrid grasp / long-range movement action refers to an action for moving the user avatar in the virtual reality environment to a destination location while a virtual object is attached to (e.g., in the hand of) the avatar. More particularly, the movement management unit 226 may move the avatar of the user to the location in the virtual environment indicated by the pointing gesture of the user while the virtual object targeted by the object grasping gesture detected in Step S310 remains attached to the avatar.

[0058] Next, at Step S314, the movement recognition unit 224 determines whether long-range movement is completed. More particularly, the movement recognition unit 224 may use the trained movement recognition model 231 to identify whether the user movements corresponding to the long-range movement action are being performed by the user (e.g., the user is maintaining a pointing-gesture to perform another teleportation action). In the case that long-range movement is completed, the method ends. In the case that long-range movement is not completed (e.g., the user wishes to perform another long-range movement action), the method returns to Step S308.

[0059] At Step S316, in the case that the movement recognition unit 224 identifies user movements corresponding to the short-range movement action, the movement management unit 226 determines to perform the short-range movement action with respect to the user avatar in the virtual environment.

[0060] Next, at Step S318, the movement recognition unit 224 uses the trained movement recognition model 231 to identify whether the user movements correspond to an object grasping action. Here, the object-grasping action refers to the act of picking-up, holding, or gripping a virtual object in the virtual environment, and may be identified in response to detecting that the user is performing an object grasping gesture (e.g., curling the thumb and one or more fingers of the hand) in the real-world environment. In the case that the movement recognition unit 224 does not identify user movements corresponding to the object-grasping action, the method proceeds to Step S319. In the case that the movement recognition unit 224 identifies user movements corresponding to the object-grasping action, the method proceeds to Step S320.

[0061] At Step S319, in the case that the movement recognition unit 224 does not identify user movements corresponding to the object-grasping action, the movement management unit 226 performs the short-range movement action determined at Step S316. For example, the movement management unit 226 may perform the short-range movement action by moving the avatar of the user in a specified direction (e.g., the direction in which the avatar is facing) in the virtual environment at a speed calculated based on the set of movement factors (e.g., the arm swinging frequency of the user) stored in the wearer movement database 232 as well as a movement speed multiplier adjustable by the user.

[0062] At Step S320, in the case that the movement recognition unit 224 identifies user movements corresponding to the object-grasping action at Step S318, the movement management unit 226 performs a hybrid grasp / short-range movement action with respect to the avatar of the user in the virtual environment. Here, the hybrid grasp / short-range movement action refers to an action for moving the user avatar in the virtual reality environment in a specified direction at a speed calculated based on the set of movement factors (e.g., the arm swinging frequency of the user) stored in the wearer movement database 232 as well as a movement speed scaling multiplier adjustable by the user while a virtual object remains attached to the avatar. More particularly, the movement management unit 226 may move the avatar of the user in the specified direction (e.g., the direction in which the avatar is facing) at the calculated speed while the virtual object targeted by the object grasping gesture detected in Step S318 remains attached to the avatar.

[0063] Next, at Step S322, the movement recognition unit 224 determines whether short-range movement is completed. More particularly, the movement recognition unit 224 may use the trained movement recognition model 231 to identify whether the user movements corresponding to the short-range movement action are being performed by the user (e.g., the user is continuing to perform arm-walking to continue short-range movement). In the case that short-range movement is completed, the method ends. In the case that long-range movement is not completed (e.g., the user wishes to perform another long-range movement action), the method returns to Step S316.

[0064] According to the movement action management method 300 described above with reference to FIG. 3, by distinguishing between user movements for short-range movement actions, long-range movement actions, as well as object-grasping actions, it becomes possible to provide a virtual reality management technique for facilitating precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions.

[0065] Next, with reference to FIG. 4, an example of training and using the movement recognition model according to the embodiments of the present disclosure will be described.

[0066] FIG. 4 is a diagram illustrating an example of training and using the movement recognition model 231 according to the embodiments of the present disclosure. As illustrated in FIG. 4, the movement recognition model 231 may be trained based on a collected set of movement training data, and then used to recognize user movements in subsequently captured video data.

[0067] First, during a training phase 410, the model training unit 228 (illustrated in FIG. 2) may be configured to use the sensor units 248 to collect a set of movement training data for the user 415 of the virtual reality management device 210. In embodiments, collecting the set of movement training data may include prompting the user 415 to perform a variety of hand gestures 412 and arm movements 414, and using the sensor units 248 to capture video data of the users’ movements. As examples, the hand gestures 412 may include open and closed hand gestures, pointing gestures, object grasping gestures, and the like. More particularly, the objecting-grasping gestures may be grasping gestures specific to holding certain objects, such as a wrenches, pliers, hammers, drills or the like. The grasping gestures to be performed by the user 415 may be selected based on the content of the virtual reality environment generated by the VR generation unit 222. Similarly, the arm movements 414 may include a standing arm posture, arm-walking movements, an extended arm posture, and a pointing arm gesture.

[0068] Next, the model training unit 228 uses the captured set of movement training data to train, as the movement recognition model 231, a set of convolutional neural network (CNN) models to identify the captured user movements. In this way, the model training unit 228 can be trained to identify various user movements such as open and closed hand gestures, pointing gestures, object grasping gestures, arm-swinging, extended arm postures, standing arm postures, arm-walking, and the like within captured image data.

[0069] Next, during a use phase 420, the trained movement recognition model 231 may be used to identify user movements within subsequently captured video data, and the movement management unit 226 may implement one or more virtual environment management actions based on the detected user movements. Here, “virtual environment management actions” is used to refer to any action performed in the virtual environment based on detected user movements, and includes object grasping actions, long-range movement actions, short-range movement actions, and combinations thereof. For example, as illustrated in FIG. 4, in response to identifying an object grasping gesture, the movement management unit 226 may execute an object grasping action to cause an avatar 423 representing the user 415 to grasp a virtual object 424 in the virtual environment VE. Further, it should be noted that by training a movement recognition model 231 including separate CNN models trained to identify separate movements, it is possible to facilitate recognition of multiple simultaneously input movements (e.g., object grasping gestures and pointing gesture). As a detailed description of the user movements and corresponding virtual environment actions will be described later herein, a description thereof will be omitted here.

[0070] In this way, by training the movement recognition model 231 to detect simultaneously input user movements, such as movements for locomotion as well as movements for object grasping in the virtual environment, it is possible to facilitate precise long-range and short-range movement in a virtual environment while enabling complex multi-task actions.

[0071] Next, with reference to FIGs. 5-7, an example of the hybrid grasp / long-range movement action according to the embodiments of the present disclosure will be described. It should be noted that, in FIGs. 5-12, the upper portion (above the thick black line) illustrates a real world environment RE and the lower portion (below the thick black line) illustrates a virtual environment VE.

[0072] As described herein, aspects of the disclosure relate to using the movement recognition model 231 to recognize multiple user movements to facilitate precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions. In the present disclosure, actions performed based on multiple recognized user movements are referred to as hybrid actions. The hybrid actions referred to herein may include a hybrid grasp / long-range movement action for moving a user avatar in the virtual reality environment to a destination location while a virtual object remains in the hand of the avatar. FIG. 5 is a diagram illustrating an example of a first phase of the hybrid grasp / long-range movement action according to the embodiments of the present disclosure.

[0073] First, as illustrated in FIG. 5, the movement recognition unit 224 may detect, by analyzing a set of video data collected by the sensor units 248 of the virtual reality management device 210 using the movement recognition model 231, that a first hand of the user 415 corresponds to an object grasping gesture 510 (e.g., thumb and one or more fingers curled to simulate gripping an object) for grasping a virtual object in the virtual environment VE. In response to detecting the object grasping gesture 510, the movement management unit 226 executes an object grasping action in the virtual environment VE to cause the avatar 423 to grasp a virtual object 424.

[0074] Next, the movement recognition unit 224 may detect, by analyzing the set of video data collected by the sensor units 248 of the virtual reality management device 210 using the movement recognition model 231, that the first hand of the user 415 also corresponds to a pointing gesture 512 for initiating a long-range movement action in the virtual environment VE. Here, it should be noted that the object grasping gesture 510 and the pointing gesture 512 may be performed by the same hand of the user 415. For example, the user 415 may extend their arm and index finger forward to indicate the pointing gesture 512 while simultaneously using their other fingers and thumb of the same hand to create the object grasping gesture 510.

[0075] In this case, the movement recognition model 231 determines that the hand of the user 415 corresponds to both the object grasping gesture 510 with respect to a virtual object 424 in the virtual reality environment VE and the pointing gesture 512, and the movement management unit 226 determines to execute a hybrid grasp / long-range movement action for the avatar 423 of the user in the virtual reality environment VE. More particularly, the movement management unit 226 may present a virtual pointer 522 represented by a ray extending from the avatar 423 in the virtual environment VE for use in selection of a destination location 524 of the hybrid grasp / long-range movement action.

[0076] FIG. 6 is a diagram illustrating an example of a second phase of the hybrid grasp / long-range movement action according to the embodiments of the present disclosure.

[0077] As illustrated in FIG. 6, in embodiments, execution of the hybrid grasp / long-range movement action in the VE may require input of a confirmation action 610 from the user 415. Here, the confirmation action 610 refers to an action performed by the user 415 to verify their intent to perform a particular virtual environment management action such as the hybrid grasp / long-range movement action. In embodiments, the confirmation action 610 may include a particular gesture, such as an open-hand gesture, a closed-hand (fist) gesture, a hand wave, a voice command, two rapid eye blinks performed within a certain time frame (e.g., 1 second, 2 seconds), a wink (e.g., a single eye blink). In the case that the confirmation action 610 is a hand gesture, the confirmation action 610 may be performed by a second hand of the user (e.g., the hand opposite to the first hand used to perform the object grasping gesture 510 and the pointing gesture 512 illustrated in FIG. 5). Here, for ease of explanation, an example will be described in which the confirmation action 610 is performed by a hand gesture, but it should be noted that the confirmation action 610 is not particularly limited, and any predetermined user input that indicates their intent to perform the hybrid grasp / long-range movement action may be detected as the confirmation action 610.

[0078] The movement recognition unit 224 may detect, by analyzing a set of video data collected by the sensor units 248 of the virtual reality management device 210 using the movement recognition model 231, that the second hand of the user 415 is indicating the confirmation action 610 while the first hand of the user 415 continues to indicate both the object grasping gesture 510 and the pointing gesture 512 illustrated in FIG. 5. Subsequently, in response to detection of the confirmation action 610, the movement management unit 22 may present a visual verification element 615 in the virtual environment VE for indicating a verification time threshold. The visual confirmation element 615 may, for example, be represented using a circular icon that changes color (e.g., fills up) as long as the confirmation action 610 is maintained by the user 415, such that when the verification time threshold elapses and the icon becomes full, the intent of the user to perform the hybrid grasp / long-range movement action can be verified.

[0079] FIG. 7 is a diagram illustrating an example of a third phase of the hybrid grasp / long-range movement action according to the embodiments of the present disclosure.

[0080] As illustrated in FIG. 7, in response to the visual confirmation element 615 becoming full (e.g., the verification time threshold elapsing), the movement management unit 226 may execute the hybrid grasp / long-range movement action by moving the avatar 423 in the virtual reality environment VE to the destination location 524 designated by the virtual pointer 522 while the virtual object 424 is attached to the avatar 423.

[0081] As illustrated in FIGs. 5-7, by training the movement recognition model 231 to identify multiple simultaneously input user movements such as object grasping gestures and pointing gestures, it is possible to input hybrid actions such as a hybrid grasp / long-range movement action for moving the avatar of the user to a desired location in the virtual environment while retaining possession of a virtual object. By means of the hybrid grasp / long-range movement action described above, users can travel large distances in the virtual environment while retaining virtual objects, facilitating tool-based task execution in large virtual environments.

[0082] Next, with reference to FIGs. 8-9, an example of training the movement recognition model for short-range movement gesture detection will be described. As described herein, the movement recognition model 231 may be trained based on a collected set of movement training data, and then used to recognize user movements (e.g., arm walking) as short-range movement gestures in subsequently captured video data. FIG. 8 is a diagram illustrating an example of capturing video data for movement recognition model training, according to embodiments of the present disclosure.

[0083] First, as illustrated in FIG. 8, the virtual reality management device 210 uses front-facing cameras of the sensor units 248 to perform a scan 810 of the real-world environment. Here, performing the scan of the real-world environment may include using the front-facing cameras of the virtual reality management device 210 to capture video data of the surrounding environment, and using one or more image analysis techniques and / or object recognition techniques to ascertain the initial location of the user (e.g., location relative to other objects in the room, absolute location based on GPS data), the size of the environment (e.g., available walking distance), objects in the real-world environment, and other factors that characterize the real-world environment.

[0084] FIG. 9 is a diagram illustrating an example of the flow of training the movement recognition model for short-range movement gesture detection. As illustrated in FIG. 9, the virtual reality management device 210 uses the display, speaker, or other output device of the input / output units 246 to provide instructions 910 to the user 415 to begin walking in the real-world environment. At this time, the virtual reality management device 210 may operate in a pass-through mode, such that the real-world environment can be displayed to the user via the display of the virtual reality management device 210. While the user 415 walks, the model training unit 228 collects a set of video data 915 using cameras of the set of sensor units 248 that shows the hand and arm movements of the user. Upon walking for a predetermined distance, the virtual reality management device 210 may provide instructions 920 to the user to stop walking in the real-world environment.

[0085] Subsequently, the model training unit 228 may analyze the collected set of video data 915 to track user hand and arm movements (e.g., arm-swinging) and real world displacement in the real-world environment, and ascertain a set of movement factors for the user 415. The set of movement factors may include arm joint angles between the torso and arms of the user 415, an arm swing frequency, a walking cadence of the user 415, and the like. More particularly, the walking cadence may include data characterizing the walking pattern of the user 415, and may include the walking speed of the user 415 in terms of steps per minute. The set of user movement factors may be stored in the wearer movement database 232 illustrated in FIG. 2. As the details of using the set of movement factors to facilitate detection of the short-range movement gesture will be described later herein, a description thereof will be omitted here.

[0086] The model training unit 228 may generate a set of movement training data based on the set of movement factors 925 characterizing the hand / arm movement and walking speed of the user as well as additional datasets 930 for a number of other users. By training the movement recognition model 231 based on this set of movement training data including both movement factors of the user 415 as well as other users, it is possible to generate a robust movement recognition model that is specifically tailored to identify the unique movements of the user 415.

[0087] Next, with reference to FIG. 10, an example of the movement factors used for detecting the short-range movement gesture will be described.

[0088] FIG. 10 is a diagram illustrating an example of the movement factors used for detecting the short-range movement gesture according to the embodiments of the present disclosure. In embodiments, the movement recognition unit 224 may use the movement recognition model 231 to analyze a set of video data collected by the cameras of the input / output units 246 of the virtual reality management device 210, and ascertain a first set of arm joint angles between a torso of the user and a first arm of the user, a second set of arm joint angles between the torso of the user and the second arm of the user, an arm swing frequency, and a walking cadence of the user.

[0089] As illustrated in FIG. 10, the first set of arm joint angles may include an arm joint angle A measured between the torso of the user 415 and the first arm (e.g., the right arm) when the user 415 swings their arms while walking. For instance, the arm joint angle A may be the maximum angle between the torso of the user 415 and the first arm when they extend their arm to its farthest point while swinging their arms when walking. Additionally, the first set of arm joint angles may include an arm joint angle B measured between the torso of the user 415 and the upper portion of the first arm (e.g., the elbow) when the user 415 returns their arms to a position by their sides while swinging their arms when walking. For instance, the arm joint angle B may be the minimum angle between the torso of the user 415 and the upper portion of the first arm when they return the first arm to their side while swinging their arms when walking. Further, the first set of arm joint angles may include an arm joint angle C measured between the torso of the user 415 and the lower portion of the first arm (e.g., the wrist) when the user 415 returns the first arm to their side while walking. For instance, the arm joint angle C may be the minimum angle between the torso of the user 415 and the lower portion of the first arm when they return the first arm to their side while swinging their arms when walking. It should be noted that, although an example of the first set of arm joint angles was described herein, as the second set of arm joint angles substantially correspond to the first set of arm joint angles, a redundant description thereof will be omitted here.

[0090] Further, in addition to the first and second set of arm joint angles, the movement recognition unit 224 may use the movement recognition model 231 to analyze a set of video data collected by the cameras of the input / output units 246 of the virtual reality management device 210 to ascertain an arm swing frequency (e.g., the rate at which the user 415 swings their arms while walking) and a walking cadence (e.g., the number of steps per unit time) of the user 415.

[0091] In embodiments, the movement recognition unit 224 may determine that the set of movement factors detected for both the first arm and a second arm of the user satisfy a predetermined movement threshold in a case that that the first set of arm joint angles satisfy a first arm joint angle threshold (e.g., the arm joint angle A is 25 degrees or more, the arm joint angle B is between 0 degrees and 15 degrees, the arm joint angle C is between 0 degrees and 20 degrees), the second set of arm joint angles satisfy a second arm joint angle threshold (e.g., substantially corresponding to the above described thresholds for the first set of arm joint angles), the arm swing frequency satisfies an arm swing frequency threshold (e.g., 0.5:1 or greater with respect to leg movements), and the walking cadence satisfies a walking cadence threshold (less than 10 steps per minute). In response to determining that the set of movement factors detected for both the first arm and a second arm of the user satisfy the predetermined movement threshold, the movement management unit 226 may detect that the short-range movement gesture indicating user intent to execute the short-range movement action in the virtual environment has been performed. Here, it should be noted that the thresholds for the first and second set of arm joint angles, the arm swing frequency, and the walking cadence may be set based on the individual movement characteristics of the user 415. In this way, it becomes possible to detect short-range arm gestures for a particular user with a high degree of accuracy.

[0092] Next, with reference to FIG. 11, an example of determining displacement in the virtual environment when performing the short-range movement action will be described.

[0093] FIG. 11 is a diagram illustrating an example of determining displacement in the virtual environment when performing the short-range movement action, according to the embodiments of the present disclosure.

[0094] In embodiments, the movement management unit 226 may be configured to calculate the displacement of the user avatar in the virtual environment while performing the short-range movement action based on a scaling factor determined based on a user adjustable speed multiplier, the arm swing frequency of the user, the length of time the user is performing arm walking (e.g., swinging their arms), the arm joint angles of the users arms, and a relative position factor of the user relative to virtual objects in the virtual environment. More particularly, the displacement of the user avatar in the virtual environment while performing the short-range movement action may be determined using the following equation.Equation 1

[0095] DisplacementVR         =Scaling Factor×Arm Swing Frequency×Time×cos(Arm Joint Angle)         ×Relative Position Factor

[0096] Here, the arm joint angles may include a difference between the arm joint angle A and the average of the arm joint angles B and C as described above with reference to FIG. 10. Additionally, the relative position factor may be set based on the size of the virtual environment, the distance between the user and a nearest virtual object in the virtual environment, or the like. Additionally, as illustrated in FIG. 11, the scaling factor may be set based on a speed multiplier adjustable by a user via a speed adjustment interface 1110 provided to the user via the display of the virtual reality management device 210. In this way, the user may increase or decrease the speed of the movement of the avatar in the virtual environment according to their preference. According to the displacement calculation method described with reference to FIG. 11, it is possible to achieve precise control of short-range movements in the virtual environment. For example, by increasing or decreasing the frequency of arm swing movements, the user can increase or decrease their speed in an intuitive manner to avoid overshooting or undershooting the desired location. Further, by setting the scaling factor based on a speed multiplier adjustable by the user, the displacement of the user may be scaled appropriately in consideration of the size of the virtual environment.

[0097] Next, with reference to FIG. 12, an example of the hybrid grasp / short-range movement action according to the embodiments of the present disclosure will be described.

[0098] As described herein, aspects of the disclosure relate to using the movement recognition model 231 to recognize multiple user movements to facilitate precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions. The hybrid actions referred to herein may include a hybrid grasp / short-range movement action for moving a user avatar a short distance in the virtual reality environment while a virtual object remains attached to an avatar representing the user. FIG. 12 is a diagram illustrating an example of the hybrid grasp / short-range movement action according to the embodiments of the present disclosure.

[0099] First, as illustrated in FIG. 12, the movement recognition unit 224 may detect, by analyzing a set of video data collected by the sensor units 248 of the virtual reality management device 210 using the movement recognition model 231, that a first hand of the user corresponds to an object grasping gesture 1210 with respect to a virtual object in the virtual reality environment, and the set of movement factors 1205 detected for both the first arm and the second arm of the user 415 satisfy the predetermined movement threshold, a hybrid grasp / short-range movement action for the avatar 423 in the virtual reality environment VE. As described herein, the set of movement factors 1205 detected for both the first arm and the second arm of the user 415 may be determined to satisfy the predetermined movement threshold in a case that the first set of arm joint angles satisfy a first arm joint angle threshold (e.g., the arm joint angle A is 25 degrees or more, the arm joint angle B is between 0 degrees and 15 degrees, the arm joint angle C is between 0 degrees and 20 degrees), the second set of arm joint angles satisfy a second arm joint angle threshold (e.g., substantially corresponding to the above described thresholds for the first set of arm joint angles), the arm swing frequency satisfies an arm swing frequency threshold (e.g., 0.5:1 or greater with respect to leg movements), and the walking cadence satisfies a walking cadence threshold (less than 10 steps per minute). In this way, arm swinging by the user while keeping the legs stationary (e.g., arm-walking) can be detected as the short-range movement gesture indicating a user’s intent to execute a short-range movement action in the virtual environment VE.

[0100] In the case that the movement recognition unit 224 detects that the first hand of the user corresponds to an object grasping gesture 1210 with respect to a virtual object in the virtual reality environment, and the set of movement factors 1205 detected for both the first arm and the second arm of the user 415 satisfy the predetermined movement threshold, the movement management unit 226 determines to execute a hybrid grasp / short-range movement action for the avatar 423 of the user in the virtual reality environment VE. More particularly, the movement management unit 226 may execute the hybrid grasp / short-range movement action by moving the avatar 423 in the virtual reality environment VE in a specified direction (e.g., the direction that the avatar is facing) at a speed calculated based on the arm swing frequency of the user and user inputs to the speed adjustment UI. Accordingly, the avatar 423 may move from a start position to an end position while the virtual object 424 remains attached to avatar.

[0101] As illustrated in FIG. 12, by training the movement recognition model 231 to identify multiple simultaneously input user movements such as object grasping gestures and arm-walking, it is possible to input hybrid actions such as a hybrid grasp / short-range movement actions for moving the avatar of the user to a desired location in the virtual environment while maintaining the avatar’s grasp on a virtual object. By means of the hybrid grasp / short-range movement action described above, users can make precision movements in the virtual environment while retaining virtual objects, facilitating tool-based task execution in small virtual environments.

[0102] Next, with reference to FIG. 13, a training and use method of the virtual reality management device 210 according to the embodiments of the present disclosure will be described.

[0103] FIG. 13 is a flowchart illustrating a training and use method 1300 of the virtual reality management device 210 according to the embodiments of the present disclosure. The training and use method 1300 of the virtual reality management device 210 is a method for training the movement recognition model 231 based on a set of movement training data collected from the user of the virtual reality management device 210, and subsequently using the trained movement recognition model 231 to facilitate precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions.

[0104] First, at Step S1302, the virtual reality management device 210 uses the front-facing cameras of the sensor units 248 to perform a scan of the real-world environment. Here, performing the scan of the real-world environment may include using the front-facing cameras of the virtual reality management device 210 to capture video data of the surrounding environment, and using one or more image analysis techniques and / or object recognition techniques to ascertain the initial location of the user (e.g., location relative to other objects in the room, absolute location based on GPS data), the size of the environment (e.g., available walking distance), objects in the real-world environment, and other factors that characterize the real-world environment.

[0105] Next, at Step S1304, the virtual reality management device 210 uses the display, speaker, or other output device of the input / output units 246 to provide instructions to the user to begin walking forward in the real-world environment. At this time, the virtual reality management device 210 may operate in a pass-through mode, such that the real-world environment can be displayed to the user via the display of the virtual reality management device 210.

[0106] Next, at Step S1306, the down-facing cameras of the sensor units 248 begin recording to collect a first set of video data that shows the hand and arm movements of the user while walking.

[0107] Next, at Step S1308, the virtual reality management device 210 uses the display, speaker, or other output device of the input / output units 246 to provide instructions to the user to stop walking in the real-world environment. Additionally, the down-facing cameras stop recording the first set of video data.

[0108] Next, at Step S1310, the model training unit 228 may analyze the collected first set of video data to track user hand and arm movements (e.g., arm-swinging) and real world displacement in the real-world environment, and generate a set of movement training data based on movement factors (e.g., arm joint angles between the torso and arms of the user, an arm swing frequency, a walking cadence of the user, and the like) ascertained by analyzing the collected first set of video data. The set of user movement factors may be stored in the wearer movement database 232 illustrated in FIG. 2.

[0109] Next, at Step S1312, the model training unit 228 trains the movement recognition model 231 based on the set of movement training data generated in Step S1310. More particularly, in embodiments, the model training unit 228 may use stochastic gradient descent, learning rate scheduling, data augmentation, dropout, batch normalization, transfer learning, weight initialization, early stopping, gradient clipping, adaptive learning rate methods, or other training techniques to train a set of CNN models to identify separate movements (e.g., pointing gestures, object grasping gestures, arm-walking) based on the set of movement training data. In this way, it is possible to train the movement recognition model to identify a wide range of user movements.

[0110] Next, at Step S1314, the down-facing cameras of the sensor units 248 begin recording to collect a second set of video data showing the hand / arm movements of the user.

[0111] Next, at Step S1316, the movement recognition unit 224 uses the trained movement recognition model 231 to analyze the second set of video data and detect user movements, and the movement management unit 226 performs one or more virtual environment management actions to facilitate locomotion and task completion in the virtual environment. It should be noted that, as the process used to detect user movements and implement virtual environment management actions may substantially correspond to the movement action management method 300 described with reference to FIG. 3, the details thereof will be omitted here.

[0112] Next, at Step S1318, the movement management unit 226 sets a movement speed of the user avatar in the virtual environment to a value based on the arm swing frequency of the user ascertained in Step S1310. More particularly, in embodiments, the movement management unit 226 may set a movement speed of the user avatar in the virtual environment to a value obtained by calculating the product of the arm swing frequency of the user and an adjustable scaling factor. Here, the adjustable scaling factor may be a value determined based on a speed multiplier adjustable by the user of the virtual reality management device 210 to increase their movement speed in the virtual environment.

[0113] Next, at Step S1320, the movement management unit 226 may increase or decrease the speed scaling factor based on user inputs to the speed adjustment UI provided on the display of the virtual reality management unit.

[0114] According to the training and use method 1300 of the virtual reality management device 210 according to the embodiments of the present disclosure, it is possible to train a robust movement recognition model that is specifically tailored to identify the unique movements of the user, and subsequently use this trained movement recognition model to facilitate precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions.

[0115] Next, with reference to FIG. 14, an example of a training process of the movement recognition model of the virtual reality management device according to the embodiments of the present disclosure will be described.

[0116] FIG. 14 is a flowchart illustrating an example of a training process 1400 of the movement recognition model 231 of the virtual reality management device 210 according to the embodiments of the present disclosure. The training process 1400 is a process for training the movement recognition model 231 to detect user movements in captured video data to facilitate locomotion as well as multi-tasking operations such as object grasping in the virtual environment. In embodiments, the training process 1400 may substantially correspond to Steps S1310-S1312 of the training and use method 1300 illustrated in FIG. 13.

[0117] First, at Step S1402, the down-facing cameras of the sensor units 248 may be used to collect a set of video data that shows the hand and arm movements of the user while walking.

[0118] Next, at Step S1404, the model training unit 228 may analyze the collected set of video data to track user hand and arm movements (e.g., arm-swinging) and real world displacement in the real-world environment, and generate a set of movement training data based on movement factors (e.g., arm joint angles A, B, and C between the torso and arms of the user, an arm swing frequency, a walking cadence of the user, and the like) ascertained by analyzing the collected set of video data. The set of user movement factors may be stored in the wearer movement database 232 illustrated in FIG. 2. In certain embodiments, in the case that the movement recognition model 231 has been pretrained based on a dataset characterizing the hand movements, arm movements, and walking cadence of other users, the model training unit 228 may generate the set of movement training data by combining the movement factors ascertained for the user of the virtual reality management device 210 and the additional datasets characterizing the hand movements, arm movements, and walking cadence of other users, and subsequently fine-tune the movement recognition model 231 based on this combined dataset, as described below.

[0119] Next, at Step S1406, the model training unit 228 may extract a set of features from the set of movement training data using a feature extraction technique (e.g., autoencoding, principle component analysis, image processing techniques), and analyze the extracted set of features using techniques such as feature importance scores or recursive feature elimination to identify, from among the set of features, a subset of the set of features that achieves an impact threshold with respect to detection accuracy of the short-range movement action and the long-range movement action. Here, the impact threshold may be a predetermined threshold used for assessing the degree of impact that a particular feature or group of features has on the detection accuracy of the short-range movement action and the long-range movement action. In embodiments, features that have a 5% or greater impact on the detection accuracy of the short-range movement action and the long-range movement action may be identified as the subset of the set of features that achieves the impact threshold. Here, the extracted set of features may be ranked based on the degree of impact they have on the detection accuracy of the short-range movement action and the long-range movement action, such that features that have a greater impact on detection accuracy are given a higher rank, and features that have a lower impact on detection accuracy are given a lower rank. In embodiments, features that have a rank below a predetermined rank (e.g., lowest 25% of ranked features) may be removed from the set of features.

[0120] Next, at Step S1408, the model training unit 228 may adjust the hyperparameters of the movement recognition model 231 to facilitate increased model performance. Here, the hyperparameters refer to parameters that govern the machine learning process of the movement recognition model 231, and may include learning rate, regularization strength, kernel parameters for support-vector machine-based models, or the like. More particularly, the model training unit 228 may use grid search or random search techniques to systematically explore the hyperparameter space to find optimal values. In this way, the movement recognition model 231 can be adapted to the individual characteristics of the movement characteristics of the user.

[0121] Next, at Step S1410, the model training unit 228 may train the movement recognition model 231 using the movement training dataset generated in Step S1404. More particularly, the model training unit 228 may split the movement training dataset into a training set and a validation set, and use stochastic gradient descent, learning rate scheduling, data augmentation, dropout, batch normalization, transfer learning, weight initialization, early stopping, gradient clipping, adaptive learning rate methods, or other training techniques to train a set of CNN models to identify separate movements (e.g., pointing gestures, object grasping gestures, arm-walking) based on the set of training set. Additionally, the model training unit 228 may assign weights to the set of features extracted from the training set, such that greater weights are assigned to the subset of the set of features identified as achieves an impact threshold with respect to detection accuracy of the short-range movement action and the long-range movement action. To avoid drastic changes to existing weights, a smaller learning rate may be utilized during training. Additionally, the model training unit 228 may monitor loss and accuracy during training to assess the convergence and performance of the movement recognition model 231.

[0122] Next, at Step S1412, the model training unit 228 may validate the movement recognition model 231 using the validation set created in Step S1410. More particularly, the model training unit 228 may calculate the accuracy, precision, recall, and F1-score of the movement recognition model 231 with respect to the validation set, and perform cross-validation techniques to reduce false-positives. The parameters of the movement recognition model 231 may be adjusted to increase accuracy, precision, recall, and F1-score of the movement recognition model 231 with respect to the validation set.

[0123] Next, at Step S1414, the model training unit 228 may perform incremental learning to cause the movement recognition model 231 to adapt to variations in the movement characteristics of the user of the virtual reality management device 210. For instance, the model training unit 228 may periodically update the movement recognition model 231 with new data collected from the user, and retrain the movement recognition model 231 using a mix of old and new data.

[0124] Next, at Step S1416, the model training unit 228 may set the predetermined movement threshold to facilitate movement detection. Here, the predetermined movement threshold is a threshold for ascertaining whether the movements of a user identified in video data correspond to a particular gesture indicating intent to perform a specified virtual environment management action such as a short-range movement action, a long-range movement action, an object grasping action, a hybrid grasp / long-range movement action, a hybrid grasp / short-range movement action, or the like. More particularly, the predetermined movement threshold may include a set of individual thresholds, such as a first arm joint angle threshold, a second arm joint angle threshold, an arm swing frequency threshold, and a walking cadence threshold. More particularly, the model training unit 228 may determine the various thresholds included in the predetermined movement threshold based on the rate of false positives and false negatives detected by the movement recognition model 231 during validation with respect to the validation set. In certain embodiments, the model training unit 228 may use a receiver operating characteristic to identify appropriate thresholds for balancing the rate of false positives and false negatives detected by the movement recognition model 231 and facilitating favorable performance.

[0125] Next, at Step S1418, the model training unit 228 may collect feedback data from the user of the virtual reality management device 210, and use this feedback data to further refine the movement recognition model 231. For example, the model training unit 228 may receive, as the feedback data, a notification from the user that a particular gesture was misclassified, and adjust the feature weights and predetermined movement threshold based on this feedback data. In this way, the robustness of the movement recognition model 231 may be increased, and variations in the movement characteristics of the user may be accounted for.

[0126] Next, at Step S1420, the movement recognition unit 224 may use the trained movement recognition model 231 to detect user movements in subsequently collected video data, and the movement management unit 226 may perform one or more virtual environment management actions to facilitate locomotion and task completion in the virtual environment. As described herein, this process may include using the cameras of the input / output units 246 to capture video data to monitor movements of the user, using the movement recognition unit 224 to identify user movements corresponding to the short-range movement action, the long-range movement action, the object grasping action, the hybrid grasp / long-range movement action, the hybrid grasp / short-range movement action, or the like, calculating displacement in the virtual environment as described with reference to FIG. 11, comparing model predictions with ground truth data, and revising the model based on feedback data from the user.

[0127] According to the training process 1400 of the movement recognition model 231 described above, it is possible to train a robust movement recognition model that is specifically tailored to identify the unique movements of the user, and subsequently use this trained movement recognition model to facilitate precise long-range and short-range movement in a virtual environment while enabling complex multi-task actions.

[0128] (Alternative Embodiment) As described herein, aspects of the present disclosure relate to a hybrid grasp / long-range movement action for moving a user avatar in the virtual reality environment to a destination location while a virtual object remains in the hand of the avatar. However, additional aspects of the disclosure relate to the recognition that, in certain cases, it may be desirable to transport multiple objects from one location to another within the virtual environment. Accordingly, in embodiments, aspects of the present disclosure relate to an object transportation action for transporting one or more objects from a first location to a second location within the virtual environment.

[0129] More particularly, in embodiments, the movement recognition may determine, by analyzing a captured set of video data using the movement recognition model, that a first hand of the user corresponds to a first pointing gesture for selecting a first virtual object in the virtual environment and a second hand of the user corresponds to an object grasping gesture, an object transportation action for moving the first virtual object in the virtual environment, and the movement management unit may execute the object transportation action by moving the first virtual object in the virtual environment to a destination location indicated by a second pointing gesture in response to detecting that the object grasping gesture has been released (e.g., the second hand of the user has returned to an open-hand gesture). In embodiments, the movement management unit may present a virtual pointer represented by a ray extending from the avatar in the virtual environment VE for use in selection of the first virtual object and the destination location.

[0130] Further, in embodiments, the user may chain together multiple object transportation actions to transport multiple objects from a first location to a second location within the virtual environment. The intention of the user to chain together multiple object transportation actions to transport multiple objects may be determined in response to detecting a confirmation action performed by the user after selection of the first virtual object. More particularly, after detecting the first pointing gesture for selecting the first virtual object in the virtual environment, the movement recognition may detect a confirmation action from the user (e.g., a double eye blink input within a predetermined time as detected by eye-tracking cameras of the virtual reality management device 210), and subsequently determine, by analyzing the captured set of video data using the movement recognition model, that the first hand of the user corresponds to a second pointing gesture for selecting a second virtual object in the virtual environment and the second hand of the user corresponds to the object grasping gesture, an object transportation action for moving the first virtual object and the second virtual object in the virtual environment, and the movement management unit may execute the object transportation action by moving the first virtual object and the second virtual object in the virtual environment to a destination location indicated by a third pointing gesture in response to detecting that the object grasping gesture has been released (e.g., the second hand of the user has returned to an open-hand gesture). As described above, the movement management unit may present a virtual pointer represented by a ray extending from the avatar in the virtual environment VE for use in selection of the first and second virtual objects and the destination location.

[0131] In this way, by performing a first pointing gesture to select a first object, performing a confirmation action (e.g., double blink) to indicate the intention of the user to select additional virtual objects, performing a second pointing gesture to select a second object, performing the object grasping gesture to initiate execution of the object transportation action, performing a third pointing gesture to select a destination location, and then releasing the object grasping gesture to complete the execution of the object transportation action, users may easily transport multiple objects from a first location to a second location within the virtual environment. Such a technique may be especially useful in virtual training environments where it may be desirable for users to transport multiple virtual objects corresponding to tools or supplies from one location to another in the virtual environment.

[0132] As described herein, aspects of the present disclosure relate to using a movement recognition model to distinguish between gestures for long-range and short-range movement actions in a virtual environment. By training a movement recognition model including separate CNN models trained to identify separate user movements, it is possible to facilitate recognition of multiple simultaneously input movements (e.g., object grasping gestures and pointing gestures), and enable smooth transitions between long-distance locomotion methods, such as teleportation, and short-distance techniques, like walking / running.

[0133] In embodiments, long-range movement actions may be detected when a user performs a pointing gesture and a confirmation action. By using the long-range movement action, the user may easily teleport long distances in the virtual environment. Two-hand gestures (e.g., a pointing gesture and a confirmation gesture) may be utilized for long-range movement action detection. In this way, by utilizing two hand gestures, differentiation between long-range movement gesture detection and other gestures can be facilitated.

[0134] In embodiments, short-range movement actions may be detected when movement factors for both of a user’s arms satisfy a predetermined movement threshold (e.g., the angle, movement frequency, and walking cadence indicate that the user is performing arm walking). The speed of the short-range movement action may be determined based on the movement characteristics (e.g., the arm swing frequency) and a speed multiplier adjustable by the user. In this way, the user may make precision movements in small virtual environments that would be difficult to achieve using teleportation techniques.

[0135] Additionally, aspects of the disclosure relate to supporting multi-tasking to allow users to perform movement actions while holding virtual objects in one hand. In this way, users can carry virtual objects with them while moving throughout the virtual environment, facilitating tool-based task execution.

[0136] In embodiments, the movement recognition model may be trained based on a dataset including training data indicating the movement characteristics of the user together with additional datasets indicating the movement characteristics of other users. By training the movement recognition model in this way, it is possible to generate a robust movement recognition model that is specifically tailored to identify the unique movements of the user.

[0137] Further, aspects of the disclosure relate to an object transportation action for transporting one or more objects from a first location to a second location within the virtual environment. Successive virtual object selections can be chained together based on detected confirmation actions performed by the user. In this way, a user may transport multiple virtual objects corresponding to tools or supplies from one location to another in the virtual environment.

[0138] According to the above-described embodiments, it is possible to provide a virtual reality management technique for facilitating precise long-range and short-range movement in a virtual environment while enabling complex multi-tasking actions.

[0139] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention.

[0140] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0141] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0142] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0143] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0144] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

[0145] While the foregoing is directed to exemplary embodiments, other and further embodiments of the invention may be devised without departing from the basic scope thereof, and the scope thereof is determined by the claims that follow. The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

[0146] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the various embodiments. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. “Set of,” “group of,” “bunch of,” etc. are intended to include one or more. It will be further understood that the terms “includes” and / or “including,” when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. In the previous detailed description of exemplary embodiments of the various embodiments, reference was made to the accompanying drawings (where like numbers represent like elements), which form a part hereof, and in which is shown by way of illustration specific exemplary embodiments in which the various embodiments may be practiced. These embodiments were described in sufficient detail to enable those skilled in the art to practice the embodiments, but other embodiments may be used and logical, mechanical, electrical, and other changes may be made without departing from the scope of the various embodiments. In the previous description, numerous specific details were set forth to provide a thorough understanding the various embodiments. But, the various embodiments may be practiced without these specific details. In other instances, well-known circuits, structures, and techniques have not been shown in detail in order not to obscure embodiments.

[0147] As described herein, the virtual reality management technique according to the present disclosure relates to the following aspects.

[0148] (Aspect 1) A virtual reality management device comprising: a set of cameras for collecting a set of video data of a real-world environment around a user; a display for displaying a virtual reality environment including an avatar representing the user; a processor; a memory; and a storage unit, wherein: the storage unit includes: a movement recognition model trained to identify body movements of the user in the set of video data; the memory includes a set of computer-readable instructions that cause the processor to: determine, by analyzing the set of video data using the movement recognition model, a short-range movement action for the avatar in the virtual reality environment in a case that a set of movement factors detected for both a first arm and a second arm of the user satisfy a predetermined movement threshold; execute the short-range movement action by moving the avatar in the virtual reality environment at a speed calculated based on the set of movement factors; determine, by analyzing the set of video data using the movement recognition model, a long-range movement action for the avatar in the virtual reality environment in a case that a first hand of the user corresponds to a pointing gesture and a confirmation action is received from the user; and execute the long-range movement action by moving the avatar in the virtual reality environment to a destination location indicated by the pointing gesture.

[0149] (Aspect 2) The virtual reality management device according to aspect 1, wherein the memory further includes a set of computer-readable instructions that cause the processor to: ascertain, by analyzing the set of video data using the movement recognition model, a first set of arm joint angles between a torso of the user and the first arm of the user, a second set of arm joint angles between the torso of the user and the second arm of the user, an arm swing frequency, and a walking cadence of the user as the set of movement factors; and determine, in a case that the first set of arm joint angles satisfy a first arm joint angle threshold, the second set of arm joint angles satisfy a second arm joint angle threshold, the arm swing frequency satisfies an arm swing frequency threshold, and the walking cadence satisfies a walking cadence threshold, that the set of movement factors detected for both the first arm and a second arm of the user satisfy the predetermined movement threshold.

[0150] (Aspect 3) The virtual reality management device according to either of aspects 1 or 2, wherein: the set of the memory further includes a set of computer-readable instructions that cause the processor to: detect, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to both an object grasping gesture with respect to a virtual object in the virtual reality environment and the pointing gesture, a hybrid grasp / long-range movement action for the avatar in the virtual reality environment; and execute the hybrid grasp / long-range movement action by moving the avatar in the virtual reality environment to the destination location indicated by the pointing gesture while the virtual object is attached to the avatar.

[0151] (Aspect 4) The virtual reality management device according to any one of aspects 1 to 3, wherein the memory further includes a set of computer-readable instructions that cause the processor to: detect, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to an object grasping gesture with respect to a virtual object in the virtual reality environment and the set of movement factors detected for both the first arm and the second arm of the user satisfy the predetermined movement threshold, a hybrid grasp / short-range movement action for the avatar in the virtual reality environment; and execute the hybrid grasp / short-range movement action by moving the avatar in the virtual reality environment at a speed calculated based on the set of movement factors while the virtual object is attached to the avatar.

[0152] (Aspect 5) The virtual reality management device according to any one of aspects 1 to 4, wherein the movement recognition model is a convolutional neural network trained to identify an open-hand gesture, a pointing gesture, a fist gesture, an object grasping gesture, and an arm swing gesture.

[0153] (Aspect 6) The virtual reality management device according to any one of aspects 1 to 5, wherein the memory further includes a set of computer-readable instructions that cause the processor to: collect, using the set of cameras, a set of movement training data characterizing the set of movement factors for the user; and generate a combined set of training data by combining the set of movement training data with an additional data set characterizing a set of movement factors for a second user; and retrain the movement recognition model based on the combined set of training data.

[0154] (Aspect 7) The virtual reality management device according to aspect 6, wherein, to retrain the movement recognition model based on the set of movement training data, the memory further includes a set of computer-readable instructions that cause the processor to: extract a set of features from the set of movement training data; identify, from among the set of features, a subset of the set of features that achieves an impact threshold with respect to detection accuracy of the short-range movement action and the long-range movement action; assign greater weights to the subset of the set of features; and train the movement recognition model using the weights assigned to the subset of the set of features.

[0155] (Aspect 8) The virtual reality management device according to any one of aspects 1 to 7, wherein the memory further includes a set of computer-readable instructions that cause the processor to: determine, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to a first pointing gesture for selecting a first virtual object in the virtual reality environment and a second hand of the user corresponds to an object grasping gesture, an object transportation action for moving the first virtual object in the virtual reality environment; and execute, in response to detecting that the object grasping gesture has been released, the object transportation action by moving the first virtual object in the virtual reality environment to a destination location indicated by a second pointing gesture.

[0156] (Aspect 9) The virtual reality management device according to any one of aspects 1 to 8, wherein: the set of cameras include an eye-tracking camera for monitoring eye movements of the user; and the memory further includes a set of computer-readable instructions that cause the processor to: detect, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to a first pointing gesture for selecting a first virtual object in the virtual reality environment; detect, using the eye-tracking camera, a double eye blink input within a predetermined time frame as a confirmation action; detect, by analyzing the set of video data using the movement recognition model, that the first hand of the user corresponds to a second pointing gesture for selecting a second virtual object in the virtual reality environment; detect, by analyzing the set of video data using the movement recognition model, that a second hand of the user corresponds to an object grasping gesture; and execute, in response to detecting that the object grasping gesture has been released, an object transportation action for moving the first virtual object and the second virtual object to a destination location in the virtual reality environment indicated by a third pointing gesture.

[0157] (Aspect 10) A virtual reality management method implemented by a virtual reality management device including: a set of cameras for collecting a set of video data of a real-world environment around a user; a display for displaying a virtual reality environment including an avatar representing the user; a processor; a memory; and a storage unit, wherein: the storage unit includes: a movement recognition model trained to identify body movements of the user in the set of video data; the memory includes a set of computer-readable instructions that cause the processor to execute a virtual reality management method comprising steps of: ascertaining, by analyzing the set of video data using the movement recognition model, a first set of arm joint angles between a torso of the user and a first arm of the user, a second set of arm joint angles between the torso of the user and a second arm of the user, an arm swing frequency, and a walking cadence of the user as the set of movement factors; determining, a short-range movement action for the avatar in the virtual reality environment in a case that the first set of arm joint angles satisfy a first arm joint angle threshold, the second set of arm joint angles satisfy a second arm joint angle threshold, the arm swing frequency satisfies an arm swing frequency threshold, and the walking cadence satisfies a walking cadence threshold; executing the short-range movement action by moving the avatar in the virtual reality environment at a speed calculated based on the arm swing frequency; determining, by analyzing the set of video data using the movement recognition model, a long-range movement action for the avatar in the virtual reality environment in a case that a first hand of the user corresponds to a pointing gesture and a confirmation action is received from the user; executing the long-range movement action by moving the avatar in the virtual reality environment to a destination location indicated by the pointing gesture; detecting, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to both an object grasping gesture with respect to a virtual object in the virtual reality environment and the pointing gesture, a hybrid grasp / long-range movement action for the avatar in the virtual reality environment; executing the hybrid grasp / long-range movement action by moving the avatar in the virtual reality environment to the destination location indicated by the pointing gesture while the virtual object is attached to the avatar; detecting, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to an object grasping gesture with respect to a virtual object in the virtual reality environment, the first set of arm joint angles satisfy the first arm joint angle threshold, the second set of arm joint angles satisfy the second arm joint angle threshold, the arm swing frequency satisfies the arm swing frequency threshold, and the walking cadence satisfies the walking cadence threshold, a hybrid grasp / short-range movement action for the avatar in the virtual reality environment; and executing the hybrid grasp / short-range movement action by moving the avatar in the virtual reality environment at a speed calculated based on the arm swing frequency while the virtual object is attached to the avatar.

[0158] (Aspect 11) The virtual reality management method according to aspect 10, wherein the memory further includes a set of computer-readable instructions that cause the processor to execute steps of: determining, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to a first pointing gesture for selecting a first virtual object in the virtual reality environment and a second hand of the user corresponds to an object grasping gesture, an object transportation action for moving the first virtual object in the virtual reality environment; and executing, in response to detecting that the object grasping gesture has been released, the object transportation action by moving the first virtual object in the virtual reality environment to a destination location indicated by a second pointing gesture.

[0159] (Aspect 12) The virtual reality management method according to either aspect 11 or 11, wherein: the set of cameras include an eye-tracking camera for monitoring eye movements of the user; and the memory further includes a set of computer-readable instructions that cause the processor to execute steps of: detecting, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to a first pointing gesture for selecting a first virtual object in the virtual reality environment; detecting, using the eye-tracking camera, a double eye blink input within a predetermined time frame as a confirmation action; detecting, by analyzing the set of video data using the movement recognition model, that the first hand of the user corresponds to a second pointing gesture for selecting a second virtual object in the virtual reality environment; detecting, by analyzing the set of video data using the movement recognition model, that a second hand of the user corresponds to an object grasping gesture; and executing, in response to detecting that the object grasping gesture has been released, an object transportation action for moving the first virtual object and the second virtual object to a destination location in the virtual reality environment indicated by a third pointing gesture.

[0160] 150... Virtual reality management application 200... Virtual reality management system 210... Virtual reality management device 220... Memory 222... VR generation unit 224... Movement recognition unit 226... Movement management unit 228... Model training unit 230… Storage unit 231... Movement recognition model 232... Wearer movement database 244... Processor 246… Input / output units 248… Sensor units 250… Communication network 260... User terminal

Claims

1. A virtual reality management device comprising: a set of cameras for collecting a set of video data of a real-world environment around a user; a display for displaying a virtual reality environment including an avatar representing the user; a processor; a memory; and a storage unit, wherein: the storage unit includes: a movement recognition model trained to identify body movements of the user in the set of video data; the memory includes a set of computer-readable instructions that cause the processor to: determine, by analyzing the set of video data using the movement recognition model, a short-range movement action for the avatar in the virtual reality environment in a case that a set of movement factors detected for both a first arm and a second arm of the user satisfy a predetermined movement threshold; execute the short-range movement action by moving the avatar in the virtual reality environment at a speed calculated based on the set of movement factors; determine, by analyzing the set of video data using the movement recognition model, a long-range movement action for the avatar in the virtual reality environment in a case that a first hand of the user corresponds to a pointing gesture and a confirmation action is received from the user; and execute the long-range movement action by moving the avatar in the virtual reality environment to a destination location indicated by the pointing gesture.

2. The virtual reality management device according to claim 1, wherein the memory further includes a set of computer-readable instructions that cause the processor to: ascertain, by analyzing the set of video data using the movement recognition model, a first set of arm joint angles between a torso of the user and the first arm of the user, a second set of arm joint angles between the torso of the user and the second arm of the user, an arm swing frequency, and a walking cadence of the user as the set of movement factors; and determine, in a case that the first set of arm joint angles satisfy a first arm joint angle threshold, the second set of arm joint angles satisfy a second arm joint angle threshold, the arm swing frequency satisfies an arm swing frequency threshold, and the walking cadence satisfies a walking cadence threshold, that the set of movement factors detected for both the first arm and a second arm of the user satisfy the predetermined movement threshold.

3. The virtual reality management device according to claim 1, wherein: the set of the memory further includes a set of computer-readable instructions that cause the processor to: detect, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to both an object grasping gesture with respect to a virtual object in the virtual reality environment and the pointing gesture, a hybrid grasp / long-range movement action for the avatar in the virtual reality environment; and execute the hybrid grasp / long-range movement action by moving the avatar in the virtual reality environment to the destination location indicated by the pointing gesture while the virtual object is attached to the avatar.

4. The virtual reality management device according to claim 1, wherein the memory further includes a set of computer-readable instructions that cause the processor to: detect, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to an object grasping gesture with respect to a virtual object in the virtual reality environment and the set of movement factors detected for both the first arm and the second arm of the user satisfy the predetermined movement threshold, a hybrid grasp / short-range movement action for the avatar in the virtual reality environment; and execute the hybrid grasp / short-range movement action by moving the avatar in the virtual reality environment at a speed calculated based on the set of movement factors while the virtual object is attached to the avatar.

5. The virtual reality management device according to claim 1, wherein the movement recognition model is a convolutional neural network trained to identify an open-hand gesture, a pointing gesture, a fist gesture, an object grasping gesture, and an arm swing gesture.

6. The virtual reality management device according to claim 1, wherein the memory further includes a set of computer-readable instructions that cause the processor to: collect, using the set of cameras, a set of movement training data characterizing the set of movement factors for the user; and generate a combined set of training data by combining the set of movement training data with an additional data set characterizing a set of movement factors for a second user; and retrain the movement recognition model based on the combined set of training data.

7. The virtual reality management device according to claim 6, wherein, to retrain the movement recognition model based on the set of movement training data, the memory further includes a set of computer-readable instructions that cause the processor to: extract a set of features from the set of movement training data; identify, from among the set of features, a subset of the set of features that achieves an impact threshold with respect to detection accuracy of the short-range movement action and the long-range movement action; assign greater weights to the subset of the set of features; and train the movement recognition model using the weights assigned to the subset of the set of features.

8. The virtual reality management device according to claim 1, wherein the memory further includes a set of computer-readable instructions that cause the processor to: determine, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to a first pointing gesture for selecting a first virtual object in the virtual reality environment and a second hand of the user corresponds to an object grasping gesture, an object transportation action for moving the first virtual object in the virtual reality environment; and execute, in response to detecting that the object grasping gesture has been released, the object transportation action by moving the first virtual object in the virtual reality environment to a destination location indicated by a second pointing gesture.

9. The virtual reality management device according to claim 1, wherein: the set of cameras include an eye-tracking camera for monitoring eye movements of the user; and the memory further includes a set of computer-readable instructions that cause the processor to: detect, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to a first pointing gesture for selecting a first virtual object in the virtual reality environment; detect, using the eye-tracking camera, a double eye blink input within a predetermined time frame as a confirmation action; detect, by analyzing the set of video data using the movement recognition model, that the first hand of the user corresponds to a second pointing gesture for selecting a second virtual object in the virtual reality environment; detect, by analyzing the set of video data using the movement recognition model, that a second hand of the user corresponds to an object grasping gesture; and execute, in response to detecting that the object grasping gesture has been released, an object transportation action for moving the first virtual object and the second virtual object to a destination location in the virtual reality environment indicated by a third pointing gesture.

10. A virtual reality management method implemented by a virtual reality management device including: a set of cameras for collecting a set of video data of a real-world environment around a user; a display for displaying a virtual reality environment including an avatar representing the user; a processor; a memory; and a storage unit, wherein: the storage unit includes: a movement recognition model trained to identify body movements of the user in the set of video data; the memory includes a set of computer-readable instructions that cause the processor to execute a virtual reality management method comprising steps of: ascertaining, by analyzing the set of video data using the movement recognition model, a first set of arm joint angles between a torso of the user and a first arm of the user, a second set of arm joint angles between the torso of the user and a second arm of the user, an arm swing frequency, and a walking cadence of the user as the set of movement factors; determining, a short-range movement action for the avatar in the virtual reality environment in a case that the first set of arm joint angles satisfy a first arm joint angle threshold, the second set of arm joint angles satisfy a second arm joint angle threshold, the arm swing frequency satisfies an arm swing frequency threshold, and the walking cadence satisfies a walking cadence threshold; executing the short-range movement action by moving the avatar in the virtual reality environment at a speed calculated based on the arm swing frequency; determining, by analyzing the set of video data using the movement recognition model, a long-range movement action for the avatar in the virtual reality environment in a case that a first hand of the user corresponds to a pointing gesture and a confirmation action is received from the user; executing the long-range movement action by moving the avatar in the virtual reality environment to a destination location indicated by the pointing gesture; detecting, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to both an object grasping gesture with respect to a virtual object in the virtual reality environment and the pointing gesture, a hybrid grasp / long-range movement action for the avatar in the virtual reality environment; executing the hybrid grasp / long-range movement action by moving the avatar in the virtual reality environment to the destination location indicated by the pointing gesture while the virtual object is attached to the avatar; detecting, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to an object grasping gesture with respect to a virtual object in the virtual reality environment, the first set of arm joint angles satisfy the first arm joint angle threshold, the second set of arm joint angles satisfy the second arm joint angle threshold, the arm swing frequency satisfies the arm swing frequency threshold, and the walking cadence satisfies the walking cadence threshold, a hybrid grasp / short-range movement action for the avatar in the virtual reality environment; and executing the hybrid grasp / short-range movement action by moving the avatar in the virtual reality environment at a speed calculated based on the arm swing frequency while the virtual object is attached to the avatar.

11. The virtual reality management method according to claim 10, wherein the memory further includes a set of computer-readable instructions that cause the processor to execute steps of: determining, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to a first pointing gesture for selecting a first virtual object in the virtual reality environment and a second hand of the user corresponds to an object grasping gesture, an object transportation action for moving the first virtual object in the virtual reality environment; and executing, in response to detecting that the object grasping gesture has been released, the object transportation action by moving the first virtual object in the virtual reality environment to a destination location indicated by a second pointing gesture.

12. The virtual reality management method according to claim 10, wherein: the set of cameras include an eye-tracking camera for monitoring eye movements of the user; and the memory further includes a set of computer-readable instructions that cause the processor to execute steps of: detecting, by analyzing the set of video data using the movement recognition model, that a first hand of the user corresponds to a first pointing gesture for selecting a first virtual object in the virtual reality environment; detecting, using the eye-tracking camera, a double eye blink input within a predetermined time frame as a confirmation action; detecting, by analyzing the set of video data using the movement recognition model, that the first hand of the user corresponds to a second pointing gesture for selecting a second virtual object in the virtual reality environment; detecting, by analyzing the set of video data using the movement recognition model, that a second hand of the user corresponds to an object grasping gesture; and executing, in response to detecting that the object grasping gesture has been released, an object transportation action for moving the first virtual object and the second virtual object to a destination location in the virtual reality environment indicated by a third pointing gesture.

Citation Information

Patent Citations

  • Methods and systems for controlling a device using hand gestures in multi-user environment

    US11112875B1

  • Gesture personalization and profile roaming

    US20110093820A1

  • Virtual object operating method and virtual object operating system

    US20210365104A1

  • VR device, method, program and recording medium

    WO2020017440A1