Space interaction method, head mounted display system, readable medium, and program product

By acquiring angular velocity information through the inertial sensor of a handheld device, recognizing gestures, and controlling the operation of a head-mounted display device, the problem of poor device mobility and universality in existing technologies is solved, and lightweight gesture interaction is achieved.

CN120994060APending Publication Date: 2025-11-21HANGZHOU LINGBAN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511087065.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing spatial interaction technologies rely on optical tracking and multi-sensor solutions, resulting in poor device mobility and versatility.

Method used

By acquiring angular velocity information sequences through the inertial sensors of handheld devices, generating target angular velocity information, recognizing hand gestures, and controlling the head-mounted display device to perform operations, the use of external devices and additional sensors is avoided.

Benefits of technology

It improves the mobility and versatility of the device, makes the device lighter, and enables effective gesture interaction without the need for external devices and multiple sensors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994060A_ABST
    Figure CN120994060A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a space interaction method, a head-mounted display system, a readable medium and a program product. A specific embodiment of the method comprises the following steps: acquiring an angular velocity information sequence corresponding to the handheld equipment through an inertial sensor of the handheld equipment; generating target angular velocity information according to the angular velocity information sequence; according to the generated at least one piece of target angular velocity information, gesture action information corresponding to the handheld device is recognized, and the at least one piece of target angular velocity information has a time sequence; and in response to determining that the gesture action information meets a preset space interaction triggering condition, controlling a head-mounted display device in communication connection with the handheld device to execute a space operation triggered by the handheld device. According to the embodiment, external equipment and mark points do not need to be arranged, multiple sensors do not need to be used either, the equipment can be lighter, and the mobility and universality of the equipment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of computer technology, and particularly to a spatial interaction method, a head-mounted display system, a readable medium and a program product. BACKGROUND

[0002] Spatial interaction technology is a technology in which a user interacts with content displayed in an extended reality space through gestures and other actions. At present, spatial interaction usually relies on the following methods: an optical scheme based on optical tracking, which requires capturing marker points on a user device or a limb through a camera, an infrared sensor or a laser radar, calculating the spatial position and posture thereof, or a hybrid scheme based on an IMU combined with other sensors (optical, ultrasonic, UWB), which fuses to calculate the posture of a device.

[0003] However, when the above methods are used, the following technical problems often exist: the optical scheme requires external devices and marker points, and the hybrid scheme requires multiple sensors, resulting in poor mobility and universality of the device.

[0004] The above information disclosed in this BACKGROUND section is only for the purpose of enhancing the understanding of the background of the present inventive concepts, and therefore, it can contain information that does not form the prior art that is already known in this country to those skilled in the art. SUMMARY

[0005] The summary section of the present disclosure is used to introduce the concepts in a brief manner, which will be described in detail in the following detailed description section. The summary section of the present disclosure is not intended to identify key or essential features of the claimed technology, nor is it intended to be used to limit the scope of the claimed technology.

[0006] Some embodiments of the present disclosure propose a spatial interaction method, a head-mounted display system, a computer readable medium and a computer program product to solve one or more of the technical problems mentioned in the background section.

[0007] In a first aspect, some embodiments of the present disclosure provide a spatial interaction method, which comprises: acquiring a sequence of angular velocity information corresponding to a handheld device through an inertial sensor of the handheld device; generating target angular velocity information according to the sequence of angular velocity information; identifying gesture action information corresponding to the handheld device according to at least one generated target angular velocity information, wherein the at least one target angular velocity information has a time sequence; and in response to determining that the gesture action information satisfies a preset spatial interaction triggering condition, controlling a head-mounted display device communicatively connected to the handheld device to perform a spatial operation triggered by the handheld device.

[0008] Optionally, the generating the target angular velocity information according to the sequence of angular velocity information comprises: performing smoothing processing on the sequence of angular velocity information to obtain smoothed angular velocity information as the target angular velocity information.

[0009] Optionally, the smoothing processing on the sequence of angular velocity information to obtain smoothed angular velocity information as the target angular velocity information comprises: determining average angular velocity information of each angular velocity information in the sequence of angular velocity information as the smoothed angular velocity information.

[0010] Optionally, the target angular velocity information comprises angular velocity corresponding to each coordinate axis; and the identifying the gesture action information corresponding to the handheld device according to the generated at least one target angular velocity information comprises: in response to determining that the generated target angular velocity information comprises angular velocity corresponding to any coordinate axis greater than a first preset threshold for a duration greater than a preset time length, and the at least one target angular velocity information comprises maximum angular velocity corresponding to the any coordinate axis greater than a second preset threshold, determining a gesture type corresponding to a shaking action of the any coordinate axis as the gesture action information corresponding to the handheld device.

[0011] Optionally, the method further comprises: in response to detecting that the generated one target angular velocity information comprises angular velocity corresponding to any coordinate axis greater than the first preset threshold, modifying an action flag bit to an in-process state; and in response to determining that the generated target angular velocity information comprises angular velocity corresponding to the any coordinate axis greater than the first preset threshold for a duration greater than the preset time length, and the at least one target angular velocity information comprises maximum angular velocity corresponding to the any coordinate axis greater than the second preset threshold, resetting the action flag bit to a stop state.

[0012] Optionally, the method further comprises: in response to determining that an interactive action supported by a current display window in the head-mounted display device comprises a shaking action corresponding to any coordinate axis, determining that the gesture action information satisfies a preset spatial interaction triggering condition; and determining a preset spatial operation of the coordinate axis included in the gesture action information as a spatial operation triggered by the handheld device corresponding to the shaking action of the current display window.

[0013] Optionally, the method further comprises: in response to determining that the handheld device changes from a stationary state to a moving state, and the gesture action information satisfies the preset spatial interaction triggering condition, determining that the handheld device satisfies a calibration triggering condition; and in response to determining that the handheld device satisfies the calibration triggering condition, calibrating a ray direction of the handheld device according to camera direction data of the head-mounted display device.

[0014] In a second aspect, some embodiments of the present disclosure provide a head-mounted display system, comprising: a handheld device configured to interact with a head-mounted display device; the head-mounted display device is communicatively connected with the handheld device and configured to image in front of a user's eyes; the handheld device and / or the head-mounted display device comprises: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.

[0015] In a third aspect, some embodiments of the present disclosure provide a computer readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0016] In a fourth aspect, some embodiments of the present disclosure provide a computer program product comprising a computer program, which, when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0017] The above various embodiments of the present disclosure have the following beneficial effects: through the space interaction method of some embodiments of the present disclosure, without setting external devices and marker points, and without using multiple sensors, the device can be made more lightweight, and the mobility and universality of the device are improved. Specifically, the reason why the mobility and universality of the device are poor is that the optical scheme needs to set external devices and marker points, and the hybrid scheme needs multiple sensors, resulting in poor mobility and universality of the device. Based on this, the space interaction method of some embodiments of the present disclosure first acquires a sequence of angular velocity information corresponding to the handheld device through the inertial sensor of the handheld device. Thus, the time-series angular velocity information can be collected through the inertial sensor of the handheld device. Then, target angular velocity information is generated according to the sequence of angular velocity information. Thus, the time-series angular velocity information collected in real time can be dynamically calibrated. Then, gesture action information corresponding to the handheld device is identified according to the generated at least one target angular velocity information, wherein the at least one target angular velocity information has a time sequence. Thus, gesture recognition and action mapping can be performed using the dynamically calibrated angular velocity information. Secondly, in response to determining that the gesture action information satisfies a preset space interaction triggering condition, the head-mounted display device communicatively connected with the handheld device is controlled to perform a space operation triggered by the handheld device. Thus, when the gesture action of the user triggers the space interaction logic of the head-mounted display device, the head-mounted display device can perform the triggered space operation. Also, since only angular velocity information is needed when performing gesture recognition and action mapping, without external device recognition of marker points and without additional other sensors to collect data, the device can be made more lightweight, and the mobility and universality of the device are improved. BRIEF DESCRIPTION OF DRAWINGS

[0018] The above and other features, aspects and advantages of the present disclosure will become more apparent with reference to the following detailed description when taken in conjunction with the accompanying drawings. Throughout the drawings, similar or same reference numerals are used to denote similar or same elements. It is to be understood that the drawings are schematic, and elements and features are not necessarily to scale.

[0019] Figure 1 is a flowchart of some embodiments of a spatial interaction method according to the present disclosure;

[0020] Figure 2 is a schematic diagram of one application scenario of a spatial interaction method according to some embodiments of the present disclosure;

[0021] Figure 3 is a structural schematic diagram of a head-mounted display system according to some embodiments of the present disclosure;

[0022] Figure 4 is a structural schematic diagram of an electronic device suitable for use to implement some embodiments of the present disclosure. DETAILED DESCRIPTION

[0023] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are only for illustrative purposes and should not be construed as limiting the scope of protection of the present disclosure.

[0024] It should also be noted that, for the sake of brevity, only the portions of the drawings that are relevant to the present disclosure are shown. The embodiments in the present disclosure and the features in the embodiments can be combined with each other where there is no conflict.

[0025] It should be noted that the terms “first”, “second”, and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.

[0026] It should be noted that the terms “one”, “multiple” in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as “one or more”.

[0027] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0028] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0029] Figure 1 A flow 100 of some embodiments of the spatial interaction method according to the present disclosure is shown. The spatial interaction method comprises the following steps:

[0030] In step 101, a sequence of angular velocity information corresponding to a handheld device is acquired by an inertial sensor of the handheld device.

[0031] In some embodiments, the execution subject of the spatial interaction method (for example, a handheld device included in a head-mounted display system or a head-mounted display device) can acquire the above-mentioned sequence of angular velocity information corresponding to the handheld device by an inertial sensor of the handheld device. Wherein, the handheld device can be in communication connection with the head-mounted display device through a wired connection mode or a wireless connection mode. The handheld device can be at least a device used for interactive control of the head-mounted display device. The handheld device can also be used to power, charge, and provide computing power for the head-mounted display device. The handheld device can include but is not limited to a mobile phone, a tablet computer, and a mobile host. The above-mentioned head-mounted display device can be a device worn on the head and imaging in front of the user's eyes. The above-mentioned head-mounted display device can be but is not limited to AR glasses, MR glasses, and VR glasses. The above-mentioned inertial sensor can be an IMU sensor of the handheld device, which can be used to acquire IMU data of the handheld device. The IMU data can include angular velocity information. The angular velocity information can include the angular velocity of three coordinate axes. In practice, the above-mentioned execution subject can acquire each continuous angular velocity information of the above-mentioned handheld device in the last preset frame number as the sequence of angular velocity information. For example, the preset frame number can be 20 frames. It should be pointed out that the above-mentioned wireless connection mode can include but is not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other now known or future developed wireless connection modes.

[0032] In step 102, target angular velocity information is generated according to the sequence of angular velocity information.

[0033] In some embodiments, the above-mentioned execution subject can generate target angular velocity information according to the above-mentioned sequence of angular velocity information. In practice, a low-pass filter can be used to filter the above-mentioned sequence of angular velocity information to obtain a smoothed sequence of angular velocity information. Then, the last angular velocity information in the smoothed sequence of angular velocity information can be determined as the target angular velocity information.

[0034] In some optional implementations of some embodiments, the above-mentioned execution subject can smooth the above-mentioned sequence of angular velocity information to obtain smoothed angular velocity information as the target angular velocity information.

[0035] In some optional implementations of some embodiments, the subject performing the above can determine average angular velocity information of each angular velocity information in the above angular velocity information sequence as the smooth angular velocity information. The average angular velocity information can include the average of the angular velocity corresponding to each coordinate axis.

[0036] Optionally, the above angular velocity information sequence is collected in a moving state of the handheld device.

[0037] In some optional implementations of some embodiments, the subject performing the above can generate target angular velocity information according to the above angular velocity information sequence by the following steps:

[0038] First, display stationary prompt information corresponding to the handheld device in the above head-mounted display device. The stationary prompt information can be prompt content for prompting the user to keep the handheld device stationary. For example, the stationary prompt information can be "Please place the handheld device on a flat surface and keep it stationary".

[0039] Second, in response to detecting stationary confirmation information corresponding to the handheld device, obtain a stationary angular velocity information sequence corresponding to the handheld device within a preset time period through the inertial sensor of the handheld device. The stationary confirmation information can be information confirming that the user has kept the handheld device stationary. For example, the user can confirm that the handheld device has been kept stationary through a button on the head-mounted display device. The specific setting of the preset time period is not limited. The stationary angular velocity information sequence can be an angular velocity information sequence collected in a stationary state of the handheld device.

[0040] Third, generate offset angular velocity information according to the above stationary angular velocity information sequence. In practice, the average stationary angular velocity information corresponding to each stationary angular velocity information in the above stationary angular velocity information sequence can be determined as the offset angular velocity information. Specifically, the average stationary angular velocity information can be composed by averaging each stationary angular velocity corresponding to each coordinate axis. The average stationary angular velocity information can include the average stationary angular velocity corresponding to each coordinate axis.

[0041] Fourth, determine the sampling interval corresponding to the above angular velocity information sequence and the historical angle information. The historical angle information can be the angle at the previous time point. The historical angle information can include the angle corresponding to each coordinate axis. The sampling interval can be the sampling time interval of the angular velocity information sequence.

[0042] Fifth, for each angular velocity information in the above angular velocity information sequence, based on the above sampling interval and the historical angle information, the following correction steps are performed:

[0043] A first sub-step is to perform offset calibration on the angular velocity information according to the offset angular velocity information to obtain calibrated angular velocity information. In practice, for each coordinate axis, a difference between an angular velocity corresponding to the coordinate axis included in the angular velocity information and an angular velocity corresponding to the coordinate axis included in the offset angular velocity information can be determined as a calibrated angular velocity corresponding to the coordinate axis. Then, the calibrated angular velocities corresponding to the coordinate axes obtained can be determined as the calibrated angular velocity information.

[0044] A second sub-step is to generate change angle information according to the calibrated angular velocity information and the sampling interval. In practice, a product of a calibrated angular velocity corresponding to each coordinate axis included in the calibrated angular velocity information and the sampling interval can be determined as each change angle. Then, the change angles can be determined as the change angle information.

[0045] A third sub-step is to generate initial correction angle information according to the historical angle information and the change angle information. In practice, for each coordinate axis, a sum of a historical angle corresponding to the coordinate axis included in the historical angle information and a change angle corresponding to the coordinate axis included in the change angle information can be determined as an initial correction angle. Then, the initial correction angles corresponding to the coordinate axes obtained can be determined as the initial correction angle information.

[0046] A fourth sub-step is to determine acceleration information corresponding to the angular velocity information. In practice, the acceleration information can correspond to the sampling time same as that of the angular velocity information. The acceleration information can include accelerations corresponding to the coordinate axes. The acceleration information can be collected by an accelerometer in an inertial sensor of the handheld device.

[0047] A fifth sub-step is to generate reference angle information according to the acceleration information. In practice, a reference angle corresponding to the X axis can be generated according to an acceleration corresponding to the Y axis and an acceleration corresponding to the Z axis included in the acceleration information. For example, the reference angle corresponding to the X axis can be generated by the following formula: AccelAngleX = arctan2(a y , a z ). Wherein, AccelAngleX represents the reference angle corresponding to the X axis. a y represents the acceleration corresponding to the Y axis. a z represents the acceleration corresponding to the Y axis. arctan2 represents a four-quadrant arctangent function. Then, a reference angle corresponding to the Y axis can be generated according to an acceleration corresponding to the X axis, an acceleration corresponding to the Y axis, and an acceleration corresponding to the Z axis included in the acceleration information. For example, the reference angle corresponding to the Y axis can be generated by the following formula: Wherein, AccelAngleY represents the reference angle corresponding to the Y axis. a xrepresents the acceleration corresponding to the X-axis. Then, the reference angle corresponding to the X-axis and the reference angle corresponding to the Y-axis can be determined as the reference angle information.

[0048] In the sixth sub-step, the initial correction angle information and the reference angle information are fused to obtain correction angle information. In practice, for each coordinate axis except the Z-axis, a product of a preset weight and an initial correction angle corresponding to the coordinate axis in the initial correction angle information can be determined as a first angle, a product of a difference between 1 and the preset weight and a reference angle corresponding to the coordinate axis in the reference angle information can be determined as a second angle, and a sum of the first angle and the second angle can be determined as a correction angle corresponding to the coordinate axis. For the Z-axis, the initial correction angle corresponding to the Z-axis in the initial correction angle information can be directly determined as the correction angle. Then, the correction angles corresponding to the respective coordinate axes can be determined as the correction angle information.

[0049] In the seventh sub-step, correction angular velocity information corresponding to the angular velocity information is generated according to the correction angle information. In practice, for each coordinate axis, a derivative of a correction angle corresponding to the coordinate axis in the correction angle information can be obtained to obtain a correction angular velocity corresponding to the coordinate axis. For example, a difference between the correction angle and a correction angle at a previous time point corresponding to the coordinate axis can be determined, and then a ratio of the difference to the sampling interval can be determined as the correction angular velocity corresponding to the coordinate axis. Then, the obtained correction angular velocities corresponding to the respective coordinate axes can be determined as the correction angular velocity information.

[0050] In the eighth sub-step, the correction angle information is determined as a historical angle corresponding to the next angular velocity information to update the historical angle.

[0051] In the sixth step, the obtained average angular velocity information corresponding to each correction angular velocity information is determined as the target angular velocity information. Thus, the first step to the sixth step can serve as some of the inventive points of the embodiments of the present disclosure, and solve the technical problem that the gyroscope may output a non-zero value (bias) in a stationary state, resulting in continuous drift of the integrated angle; and the angle obtained by only integrating the gyroscope will drift over time, and the error will continuously accumulate, especially when running for a long time. To solve the technical problem, the present disclosure can estimate the bias value by collecting the angular velocity in the stationary state and averaging, so as to subtract the bias in the subsequent data, thereby significantly improving the accuracy of the angular velocity. In addition, the reference angle (such as the pitch angle and the roll angle) can be calculated by the accelerometer, and the angle value can be periodically “pulled back” to suppress the drift. Thus, the influence of the angular velocity bias and the angle drift can be effectively removed, thereby improving the accuracy of the angular velocity.

[0052] At step 103, gesture action information corresponding to the handheld device is identified according to the generated at least one target angular velocity information.

[0053] In some embodiments, the execution subject can identify gesture action information corresponding to the handheld device according to the generated at least one target angular velocity information. The at least one target angular velocity information has a time sequence. The at least one target angular velocity information can be each target angular velocity information generated continuously when the handheld device starts moving from a stationary state. It should be noted that the identification of gesture action information is performed in real time. In practice, the execution subject can first generate time domain feature information corresponding to the at least one target angular velocity information. The time domain feature information can include but is not limited to angular velocity mean information, angular velocity standard deviation information, angular velocity peak value number information, angular velocity peak value amplitude information, and angular velocity zero-crossing rate information. The angular velocity mean information can include average angular velocity corresponding to the X-axis, average angular velocity corresponding to the Y-axis, and average angular velocity corresponding to the Z-axis. The angular velocity standard deviation information can include angular velocity standard deviation corresponding to the X-axis, angular velocity standard deviation corresponding to the Y-axis, and angular velocity standard deviation corresponding to the Z-axis. The angular velocity peak value number information can include the number of angular velocity peaks corresponding to the X-axis, the number of angular velocity peaks corresponding to the Y-axis, and the number of angular velocity peaks corresponding to the Z-axis. The angular velocity peak value amplitude information can include the maximum angular velocity corresponding to the X-axis, the maximum angular velocity corresponding to the Y-axis, and the maximum angular velocity corresponding to the Z-axis. The angular velocity zero-crossing rate information can include the zero-crossing rate of angular velocity corresponding to the X-axis, the zero-crossing rate of angular velocity corresponding to the Y-axis, and the zero-crossing rate of angular velocity corresponding to the Z-axis. Then, frequency domain feature information corresponding to the at least one target angular velocity information can be generated. Specifically, for each coordinate axis, fast Fourier transform can be performed on the angular velocity corresponding to the coordinate axis in the at least one target angular velocity information to extract the main frequency corresponding to the coordinate axis as the frequency domain feature information. The frequency domain feature information can include but is not limited to the main frequency, the spectrum energy, and the frequency band energy ratio. The main frequency can be the frequency with the strongest energy in the angular velocity change. The spectrum energy can include energy values corresponding to each frequency band. Each frequency band can include a pre-set low frequency band, a middle frequency band, and a high frequency band. The energy value corresponding to a frequency band can be the sum of the squares of the amplitudes in the frequency band. The frequency band energy ratio can include the proportion of the energy value of each frequency band to the total energy value. The total energy value can be the sum of the energy values of each frequency band. Next, time sequence dynamic feature information corresponding to the at least one target angular velocity information can be generated. The time sequence dynamic feature information can include but is not limited to angular velocity change direction number information, angular velocity change slope information, and angular velocity modulus length change rate information. The angular velocity change direction number information can include the number of angular velocity change directions corresponding to each coordinate axis. This feature is helpful for identifying the directional change in the gesture. The angular velocity change direction number can be the number of times the angular velocity direction changes. The angular velocity change slope information can include the average angular velocity slope corresponding to each coordinate axis. This feature can be used to identify the speed change of the gesture, such as acceleration swinging or deceleration stopping. The average angular velocity slope can be the mean of the angular velocity slopes at each time point.The angular velocity module length change rate information can include a significant change quantity of the angular velocity module length change rate of each coordinate axis, which is particularly suitable for identifying "explosive" actions in the gesture, such as sudden acceleration or sudden stop. The angular velocity module length change rate can be the change rate of the angular velocity module length of each time point relative to the last time point. The significant change quantity can be the number of angular velocity module length change rates greater than a preset threshold. Then, for each coordinate axis, each feature value in the above time domain feature information, the above frequency domain feature information, and the above time sequence dynamic feature information corresponding to the coordinate axis can be combined into a feature vector corresponding to the coordinate axis. Secondly, each feature vector can be normalized to obtain feature information. Then, the feature information can be input into a pre-trained gesture action recognition network to obtain gesture action information. The gesture action recognition network can be a pre-trained neural network model that takes feature information (refer to the above feature information) as input and corresponding gesture action information as output. For example, the gesture action recognition network can include a feature encoder, a pattern library, and a similarity matching module. The feature encoder can use a fully connected network (MLP) to compress the original features into a low-dimensional space to obtain a low-dimensional feature embedding vector to extract common expressions. The pattern library can pre-store "pattern vectors" constructed for each gesture action, that is, the mean of the feature embedding vectors of all samples corresponding to the gesture action. The similarity matching module can select the gesture action corresponding to the pattern vector with the highest similarity from the pattern library as the gesture action information by taking the feature embedding vector output by the feature encoder as input. In this way, time domain, frequency domain, and time sequence features can be combined to achieve multi-dimensional feature extraction, and the gesture action recognition network used is relatively lightweight, suitable for deployment in mobile or wearable devices, and can support the addition of new gesture samples without the need to retrain the model, and has strong scalability.

[0054] In some optional implementations of some embodiments, the execution subject can determine, in response to determining that the generated target angular velocity information includes an angular velocity corresponding to any coordinate axis greater than a first preset threshold for a duration greater than a preset time length, and that the maximum angular velocity included in the at least one target angular velocity information corresponding to the any coordinate axis is greater than a second preset threshold, that the any coordinate axis and the gesture type corresponding to the shaking action are the gesture action information corresponding to the handheld device. For example, the first preset threshold can be 5 degrees per second. The preset time length can be 0.3 seconds. The second preset threshold can be 20 degrees per second. In this way, when the angular velocity of any coordinate axis is greater than the first preset threshold, it is determined that the action has started, when the action has lasted for the preset time length, it is determined that the action meets the time requirement, and when the maximum angular velocity amplitude during the action process is greater than the second preset threshold, it is determined that the intensity requirement is met, so that it can be determined that the shaking action on the coordinate axis is completed, that is, the handheld device has performed a shaking action on the coordinate axis.

[0055] Optionally, the execution subject can also modify the action flag bit to an in-process state in response to detecting that the generated target angular velocity information includes an angular velocity corresponding to any coordinate axis that is greater than the first preset threshold. The action flag bit can represent whether there is a currently in-process shaking action. In an initial state, the action flag bit can be in a stop state, indicating that there is no currently in-process shaking action, for example, can be represented by "false". The in-process state can indicate that there is a currently in-process shaking action, for example, can be represented by "true". Then, the action flag bit can be reset to the stop state in response to determining that the duration for which the generated target angular velocity information includes an angular velocity corresponding to any coordinate axis that is greater than the first preset threshold is greater than a preset time length, and the maximum angular velocity included in the at least one target angular velocity information corresponding to any coordinate axis is greater than a second preset threshold. Thus, after being reset to the stop state, the next shaking action detection can be allowed to proceed.

[0056] In step 104, in response to determining that the gesture action information satisfies a preset spatial interaction trigger condition, the head-mounted display device in communication connection with the handheld device is controlled to perform a spatial operation triggered by the handheld device.

[0057] In some embodiments, the execution subject can control the head-mounted display device in communication connection with the handheld device to perform a spatial operation triggered by the handheld device in response to determining that the gesture action information satisfies a preset spatial interaction trigger condition. The preset spatial interaction trigger condition can be that an interaction action supported by a current display window in the head-mounted display device includes a gesture action represented by the gesture action information. The interaction action supported by the current display window can be pre-configured by a developer, and the supported interaction action can correspond to a set of triggerable preset spatial operations. For example, when a photo is displayed in the current display window and the interaction action is a fast flick, the triggerable preset spatial operation can be photo page turning (e.g., displaying the next photo).

[0058] As an example, referring to Figure 2 , the user can hold the handheld device and shake it to the left, and this interaction action of the handheld device can be identified.

[0059] Optionally, the execution subject can also determine that the gesture action information satisfies a preset spatial interaction trigger condition in response to determining that an interaction action supported by a current display window in the head-mounted display device includes a shaking action corresponding to any coordinate axis. Then, the shaking action corresponding to the current display window and the preset spatial operation of the coordinate axis included in the gesture action information can be determined as a spatial operation triggered by the handheld device. Thus, a single-axis shaking action can be identified and a single-axis shaking action triggered spatial operation can be performed.

[0060] Optionally, the execution subject can further determine that the handheld device satisfies a calibration trigger condition in response to determining that the handheld device changes from a stationary state to a moving state and the gesture action information satisfies a preset spatial interaction trigger condition. Next, the ray direction of the handheld device is calibrated according to camera direction data of the head-mounted display device in response to determining that the handheld device satisfies the calibration trigger condition. In practice, the current camera direction data of the head-mounted display device can be determined as an initial direction of the ray direction of the handheld device to calibrate the ray direction. The ray direction can be a spatial direction of a ray used to indicate an interaction direction of the handheld device in the head-mounted display device. Both the camera direction data and the ray direction can include directions of respective coordinate axes. Thus, dynamic calibration can be triggered when the user picks up the handheld device and a shaking action is detected, and the direction offset of the handheld device is corrected using the camera data.

[0061] The above various embodiments of the present disclosure have the following beneficial effects: The spatial interaction method of some embodiments of the present disclosure can make the device more lightweight and improve the mobility and universality of the device without setting external devices and marker points and without using multiple sensors. Specifically, the reason why the mobility and universality of the device are poor is that the optical scheme needs to set external devices and marker points, and the hybrid scheme needs multiple sensors, resulting in poor mobility and universality of the device. Based on this, the spatial interaction method of some embodiments of the present disclosure first acquires a sequence of angular velocity information corresponding to the handheld device through an inertial sensor of the handheld device. Thus, the time-series angular velocity information can be collected through the inertial sensor of the handheld device. Then, target angular velocity information is generated according to the sequence of angular velocity information. Thus, the time-series angular velocity information collected in real time can be dynamically calibrated. Then, gesture action information corresponding to the handheld device is identified according to the generated at least one target angular velocity information, wherein the at least one target angular velocity information has a time sequence. Thus, gesture recognition and action mapping can be performed using the dynamically calibrated angular velocity information. Next, the head-mounted display device in communication connection with the handheld device is controlled to perform a spatial operation triggered by the handheld device in response to determining that the gesture action information satisfies a preset spatial interaction trigger condition. Thus, the triggered spatial operation can be performed by the head-mounted display device when the gesture action of the user triggers the spatial interaction logic of the head-mounted display device. Also, since only angular velocity information is needed when performing gesture recognition and action mapping, without external device recognition of marker points and without additional other sensors to collect data, the device can be made more lightweight, and the mobility and universality of the device are improved.

[0062] Further reference is made to Figure 3As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a head-mounted display system, which are similar to... Figure 1 Corresponding to the method embodiments shown, the system can be specifically applied to various electronic devices.

[0063] like Figure 3 As shown, some embodiments of the head-mounted display system 300 include a handheld device 301 and a head-mounted display device 302. The handheld device 301 is configured to interact with the head-mounted display device. The head-mounted display device 302 is communicatively connected to the handheld device and is configured to project an image in front of the user's eyes.

[0064] In some embodiments, the handheld device 301 and / or the head-mounted display device 302 may include one or more processors and a storage device. The storage device stores one or more programs, which, when executed by the one or more processors, cause the one or more processors to perform the following functions: Figure 1 The methods described in the corresponding embodiments. It should be noted that the handheld device can be used as an external device for interactive control of the head-mounted display device, and the head-mounted display device can also be used as an external device for display of the handheld device.

[0065] It is understandable that the methods and references implemented by the processor described in the head-mounted display system 300 are similar. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the head-mounted display system 300, and will not be repeated here.

[0066] The following is for reference. Figure 4 It illustrates an electronic device 400 suitable for implementing some embodiments of the present disclosure (e.g., Figure 1 A schematic diagram of the structure of a handheld device or head-mounted display device. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0067] like Figure 4As shown, the electronic device 400 can include a processing device 401 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 402 or loaded into a random access memory (RAM) 403 from a storage device 408. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0068] Generally, the following devices can be connected to the I / O interface 405: input devices 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; and communication devices 409. The communication devices 409 can allow the electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present. Figure 4 Each block shown in the flowcharts can represent a device or multiple devices as needed.

[0069] In particular, processes described above with reference to the flowcharts can be implemented as a computer software program according to some embodiments of the present disclosure. For example, some embodiments of the present disclosure include a computer program product including a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In some such embodiments, the computer program can be downloaded and installed from a network through the communication devices 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-described functions defined in the methods of some embodiments of the present disclosure are performed.

[0070] Note that the computer readable medium in some embodiments of the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In some embodiments of the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used by an instruction execution system, apparatus or device, or that can be used by or in connection with an instruction execution system, apparatus or device. In some embodiments of the present disclosure, the computer readable signal medium can include a computer readable program code propagated in or on a carrier medium, in which the computer readable program code is embodied. Such propagated computer readable program code can take many forms, including but not limited to, an electromagnetic signal, an optical signal or any suitable combination of the foregoing. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. Program code embodied on a computer readable medium can be transmitted using any suitable medium, including but not limited to, wire, cable, wireless, RF, infrared or any suitable combination of the foregoing.

[0071] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0072] The computer readable medium can be included in the electronic device; or can exist separately from the electronic device. The computer readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: acquire a sequence of angular velocity information corresponding to the handheld device through an inertial sensor of the handheld device; generate target angular velocity information according to the sequence of angular velocity information; identify gesture action information corresponding to the handheld device according to at least one generated target angular velocity information, wherein the at least one target angular velocity information has a time sequence; and in response to determining that the gesture action information satisfies a preset spatial interaction triggering condition, control a head-mounted display device in communication connection with the handheld device to perform a spatial operation triggered by the handheld device.

[0073] Computer program code for carrying out operations of some embodiments of the present disclosure can be written in any of one or more programming languages, including object oriented programming languages, such as Java, Smalltalk, C++, or conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0074] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the block can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0075] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, non-limiting examples of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), etc.

[0076] Some embodiments of the present disclosure also provide a computer program product comprising a computer program which, when executed by a processor, implements any of the above-mentioned spatial interaction methods.

[0077] The above description is merely exemplary of some preferred embodiments of the present disclosure and of the application principles underlying the present disclosure. It is to be understood that the scope of the present disclosure is not limited to the specific embodiments described above, and that the application scope of the embodiments involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, but also covers other technical solutions formed by the combinations of the above technical features or equivalent features thereof without deviating from the above inventive concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the embodiments of the present disclosure (but not limited to) having similar functions.

Claims

1. A spatial interaction method, comprising: The angular velocity information sequence corresponding to the handheld device is obtained through the inertial sensor of the handheld device; Target angular velocity information is generated based on the angular velocity information sequence; Based on at least one generated target angular velocity information, gesture information corresponding to the handheld device is identified, wherein the at least one target angular velocity information has a temporal order; In response to determining that the gesture information meets the preset spatial interaction triggering conditions, the head-mounted display device that is communicatively connected to the handheld device is controlled to execute the spatial operation triggered by the handheld device.

2. The method according to claim 1, wherein, The step of generating target angular velocity information based on the angular velocity information sequence includes: The angular velocity information sequence is smoothed to obtain smoothed angular velocity information, which is then used as the target angular velocity information.

3. The method according to claim 2, wherein, The smoothing process of the angular velocity information sequence to obtain smoothed angular velocity information as target angular velocity information includes: The average angular velocity information of each angular velocity information in the angular velocity information sequence is determined as smooth angular velocity information.

4. The method according to claim 1, wherein, The target angular velocity information includes the angular velocity corresponding to each coordinate axis; And the step of identifying the gesture information corresponding to the handheld device based on at least one generated target angular velocity information includes: In response to determining that the duration for which the angular velocity of any coordinate axis corresponding to the generated target angular velocity information is greater than a first preset threshold is greater than a preset duration, and that the maximum angular velocity of any coordinate axis corresponding to the at least one target angular velocity information is greater than a second preset threshold, the arbitrary coordinate axis and the gesture type corresponding to the shaking action are determined as the gesture action information corresponding to the handheld device.

5. The method according to claim 4, wherein, The method further includes: In response to detecting that the angular velocity of any coordinate axis in a generated target angular velocity information is greater than the first preset threshold, the action flag is modified to a state of in progress. In response to determining that the duration for which the generated target angular velocity information includes an angular velocity corresponding to any coordinate axis greater than the first preset threshold is greater than a preset duration, and that the maximum angular velocity of the at least one target angular velocity information including an angular velocity corresponding to any coordinate axis is greater than a second preset threshold, the action flag is reset to a stopped state.

6. The method according to claim 4, wherein, The method further includes: In response to determining that the interactive actions supported by the current display window in the head-mounted display device include shaking actions corresponding to any coordinate axis, the gesture action information is determined to meet the preset spatial interaction triggering conditions; The preset spatial operation of the coordinate axis included in the shaking action and the gesture action information corresponding to the current display window is determined as the spatial operation triggered by the handheld device.

7. The method according to any one of claims 1-6, wherein, The method further includes: In response to determining that the handheld device changes from a stationary state to a moving state, and that the gesture information satisfies a preset spatial interaction trigger condition, it is determined that the handheld device meets the calibration trigger condition; In response to determining that the handheld device meets the calibration trigger condition, the ray direction of the handheld device is calibrated based on the camera orientation data of the head-mounted display device.

8. A head-mounted display system, comprising: The handheld device is configured to interact with the head-mounted display device; A head-mounted display device, communicatively connected to the handheld device, is configured to project an image in front of the user's eyes; The handheld device and / or the head-mounted display device includes: one or more processors and a storage device, wherein the storage device stores one or more programs that, when executed by the one or more processors, cause the one or more processors to implement the method as described in any one of claims 1-7.

9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.

10. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Gesture recognition method and system based on gyroscope sensor and mobile terminal

    CN103513788A

  • Mobile control method and device of intelligent equipment, equipment and medium

    CN115278205A

  • Navigation method in walking scene, intelligent terminal and storage medium

    CN118882683A

  • Intelligent interaction method, terminal, and storage medium

    WO2022193257A1