Method for mapping three-dimensional motion gesture into key combination

Through multi-stage processing of inertial measurement unit data and interactive context matching, the modification key and arrow key combination commands are generated, which solves the problem of difficult to efficiently generate keyboard combination keys in the prior art, and achieves efficient and accurate device control.

CN120295481APending Publication Date: 2025-07-11SHENZHEN HULE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510464912.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing interaction methods based on inertial measurement units are difficult to efficiently generate keyboard combination key commands required by complex devices, especially the combination of modifier keys and arrow keys, and face problems such as large individual differences, serious environmental interference, and difficulty in gesture recognition.

Method used

By performing multi-level signal processing on the three-dimensional motion data collected by the inertial measurement unit, basic gesture recognition and sequence matching are performed, and the command intention of modifier keys and arrow keys is generated, and state control and verification are performed, and platform event conversion and multi-modal feedback signal generation are performed.

Benefits of technology

It realizes efficient and accurate generation of keyboard combination key commands required by complex devices, reduces user operation burden, improves interaction fluency and accuracy, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295481A_ABST
    Figure CN120295481A_ABST
Patent Text Reader

Abstract

The invention discloses a method, device and equipment for mapping a three-dimensional motion gesture into a key combination and a computer readable storage medium. The method comprises the following steps: generating a standardized motion feature flow according to original three-dimensional motion data collected by an inertial measurement unit; obtaining a basic gesture identifier time sequence according to the standardized motion feature flow; generating a combined command intention of a modification key and a direction key according to the basic gesture identifier time sequence; performing modification key state control and combined key event triggering processing on the intention to generate a target key event sequence; verifying the target key event sequence to obtain an event sequence passing verification; and converting the target key event sequence into a target platform standard input event and injecting the target platform standard input event, and generating a multi-mode feedback signal corresponding to the command intention or the execution state at the same time. The method for mapping the three-dimensional motion gesture into the key combination has the advantages of high efficiency, high accuracy and high adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and in particular, to a method, device, equipment and computer-readable storage medium for mapping three-dimensional motion gestures to key combinations. Background Art

[0002] In the field of human-computer interaction, to meet the needs of different user groups and diverse application scenarios, researchers have been continuously exploring input methods to replace traditional keyboards and mice. Using natural human movements for device control, especially achieving interaction by capturing limb movements through wearable sensors, has become an important research direction. The Inertial Measurement Unit (IMU) has been widely used in the fields of gesture recognition, attitude tracking, and assistive technology control due to its advantages such as small size, low power consumption, relatively controllable cost, and the ability to capture attitude and motion information in three-dimensional space. In the prior art, IMU-based interaction methods usually involve directly mapping specific and relatively independent static postures or dynamic gestures to a single device command or function call, such as simulating mouse pointer movement, triggering simple on / off instructions, or recognizing specific symbolic gestures. These technologies provide effective interaction means in certain scenarios, especially in mobile environments or when the user's hands are occupied.

[0003] However, the existing IMU-based gesture interaction technologies face significant challenges in dealing with complex device control requirements. The operations of many modern computing devices and assistive technologies (such as computer operating systems, professional application software, multi-degree-of-freedom robotic arms) require a large number of different commands, especially frequently involving keyboard combination keys, such as the combination of modifier keys (Ctrl, Shift, Alt) and arrow keys (up, down, left, right), which are crucial in text editing, window navigation, file management, and shortcut operations. Currently, most IMU-based interfaces can usually only provide a limited number of clearly distinguishable independent gesture inputs, and their "input dimension" is much lower than the "control dimension" required by complex devices. To bridge this gap, existing solutions often adopt the method of mode switching or hierarchical menus, but this significantly increases the user's operation burden and time cost, and reduces the fluency and intuitiveness of the interaction. In addition, IMU-based gesture recognition itself also faces technical problems such as large differences in individual user actions, susceptibility to environmental motion interference, difficulty in distinguishing similar gestures, and high requirements for real-time processing algorithms, which limit its application in scenarios where precise and efficient input combination keys are required.

[0004] Therefore, there is an urgent need for a new interaction method that can effectively utilize the limited basic three-dimensional motion gestures captured by the IMU and that users can easily perform to reliably and efficiently generate a large number of combination key commands, especially those containing modifier keys and arrow keys. Summary of the Invention

[0005] Embodiments of the present application provide a method for mapping three-dimensional motion gestures to key combinations, aiming to provide an efficient, accurate, robust, personalized, and user-friendly human-computer interaction solution.

[0006] To achieve the above object, embodiments of the present application provide a method for mapping three-dimensional motion gestures to key combinations, including:

[0007] Performing multi-level signal processing on the original three-dimensional motion data collected from an inertial measurement unit to obtain a standardized motion feature stream;

[0008] Performing basic gesture recognition on the standardized motion feature stream to obtain a sequence of basic gesture recognition identifiers;

[0009] Interpreting the sequence of basic gesture recognition identifiers by applying a preset sequence matching rule based on the current interaction context to obtain a modifier key and direction key combination command intention corresponding to a confirmation sequence;

[0010] Performing modifier key state control and combined key event triggering processing on the modifier key and direction key combination command intention and combining direction key information to obtain a target key event sequence;

[0011] Performing command verification processing on the target key event sequence by applying a verification rule including intention confirmation, context rationality verification, and repeated command suppression to obtain a verified target key event sequence;

[0012] Performing target platform event conversion and injection processing on the verified target key event sequence, and simultaneously performing multi-modal feedback signal generation processing corresponding to the modifier key and direction key combination command intention or execution state to obtain a standard input event injected into the target platform and the simultaneously generated multi-modal feedback signal.

[0013] In one embodiment, performing multi-level signal processing on the original three-dimensional motion data collected from an inertial measurement unit includes:

[0014] Applying a digital low-pass filter to the original three-dimensional motion data for filtering processing to obtain filtered three-dimensional motion data;

[0015] Applying a Kalman filter to the filtered three-dimensional motion data for sensor fusion processing to generate fusion motion data including real-time attitude information;

[0016] Apply a motion state detection algorithm to the fused motion data for processing. By calculating the short-time energy, signal amplitude variance, and main frequency components of the fused motion data within a sliding time window, and comparing them with a motion activation threshold dynamically adjusted based on the user's recent activity level, identify and filter out potential gesture signal segments representing the user's intention to perform gestures;

[0017] Map the data of the potential gesture motion segment to a preset range to obtain the standardized motion feature stream.

[0018] In one embodiment, perform basic gesture recognition on the standardized motion feature stream, including:

[0019] Apply a dynamic segmentation algorithm based on signal variance and derivative threshold determination to the standardized motion feature stream for processing, and real-time define the start and end time points of a single basic gesture action to obtain gesture time segments;

[0020] Perform feature extraction processing on the fused motion data within the gesture time segment, and calculate a multi-dimensional feature set including time-domain statistics, frequency-domain distribution characteristics, and autoregressive model coefficients;

[0021] Apply the principal component analysis method to the multi-dimensional feature set for dimensionality reduction processing to obtain a dimensionality-reduced gesture feature vector;

[0022] Input the dimensionality-reduced gesture feature vector into a pre-trained linear kernel support vector machine classification model for classification processing to determine the basic gesture category corresponding to the gesture time segment;

[0023] Organize the continuously determined basic gesture categories in chronological order to obtain the chronological sequence of basic gesture identifiers.

[0024] In one embodiment, interpret the chronological sequence of basic gesture identifiers by applying a preset sequence matching rule based on the current interaction context, including:

[0025] Obtain the current active application information by querying the operating system application programming interface to determine the current interaction context;

[0026] According to the current interaction context, select and load an activated sequence matching rule library from multiple predefined context-specific sequence rule libraries, and the rule library contains multiple target gesture sequences and corresponding command intents;

[0027] Append the latest identifier in the chronological sequence of basic gesture identifiers to the currently constructed user input sequence;

[0028] Perform real-time prefix matching processing on the currently constructed user input sequence and all target gesture sequences in the activated sequence matching rule library, and dynamically maintain a candidate sequence set;

[0029] Detect whether the candidate sequence set is reduced to a uniquely matched target gesture sequence, or detect whether a preset confirmation waiting time has elapsed after the last identifier of the timing sequence is input, to determine the final confirmation sequence;

[0030] Retrieve and obtain the corresponding modifier key and direction key combination command intention from the activated sequence matching rule library according to the confirmation sequence.

[0031] In one embodiment, perform modifier key state control and combined key event triggering processing on the modifier key and direction key combination command intention, and combine direction key information to obtain a target key event sequence, including:

[0032] Parse the command intention to identify the specific modifier key and specific direction key contained therein;

[0033] Determine the triggering mode of the modifier key according to the command intention:

[0034] If the command intention requires continuously holding down the specific modifier key, activate the continuous triggering mode, and control the pressing and releasing states of the specific modifier key by continuously monitoring and maintaining the stability of the basic gesture signal of the specific modifier key and applying a hysteresis threshold determination;

[0035] If the command intention requires instantaneous triggering of the combination, activate the instantaneous triggering mode;

[0036] If it is detected that a specific sequence is used to switch the locking state, activate the locking triggering mode to switch and maintain the locked or unlocked state of the modifier key;

[0037] Generate the target key event sequence including keyboard scan codes and actions according to the activated triggering mode and the parsed direction key information, in the order and timing required by the target operating system.

[0038] In one embodiment, perform command verification processing on the target key event sequence by applying verification rules including intention confirmation, context rationality verification, and duplicate command suppression, including:

[0039] Evaluate the classification confidence of the basic gesture based on which the target key event sequence is generated and the matching confidence of the confirmation sequence for intention confirmation verification;

[0040] Compare the operation corresponding to the target key event sequence with the behavior rules preset based on the current interaction context for context rationality verification;

[0041] Detect whether the target key event sequence is exactly the same as the event sequence generated within a preset short time window to perform duplicate command suppression verification;

[0042] Based on the verification results of the intention confirmation, context rationality, and duplicate command suppression, obtain the target key event sequence that passes the verification.

[0043] In one embodiment, for the target key event sequence that passes the verification, perform target platform event conversion and injection processing, and synchronously perform multi-modal feedback signal generation processing corresponding to the modifier key and arrow key combination command intention or execution status, including:

[0044] Perform target platform event conversion processing on the target key event sequence that passes the verification to obtain a standard input event that conforms to the input specification of the target operating system;

[0045] Call the underlying input simulation interface function of the target operating system to inject the standard input event data structure into the system event queue according to the parsing time sequence to obtain the standard input event injected into the target platform;

[0046] According to the modifier key and arrow key combination command intention or the execution status of the target key event sequence, query a preset feedback mapping table to generate a feedback instruction including a specified feedback channel, feedback mode, and feedback intensity level, where the feedback channel includes at least one of a tactile channel, an auditory channel, and a visual channel, and the feedback mode includes at least one of a preset vibration waveform, audio sample, and visual element style;

[0047] Apply adaptive adjustment processing to the feedback instruction, adjust the intensity level or channel priority in the feedback instruction according to the real-time environmental parameters obtained from the environmental sensor or the preference parameters set by the user through the configuration interface, and then output through the corresponding feedback actuator to obtain the synchronously generated multi-modal feedback signal.

[0048] To achieve the above object, an embodiment of the present application also proposes a device for mapping three-dimensional motion gestures to key combinations, including:

[0049] An inertial measurement unit interface for receiving the original three-dimensional motion data collected from the inertial measurement unit;

[0050] A signal processing module, connected to the inertial measurement unit interface, configured to perform multi-level signal processing on the original three-dimensional motion data to obtain a standardized motion feature stream;

[0051] A basic gesture recognition module, connected to the signal processing module, is configured to perform basic gesture recognition on the standardized motion feature stream to obtain a sequence of basic gesture recognition identifiers;

[0052] A sequence interpretation module, connected to the basic gesture recognition module, is configured to interpret the sequence of basic gesture recognition identifiers by applying a preset sequence matching rule based on the current interaction context to obtain a modifier key and direction key combination command intention corresponding to the confirmation sequence;

[0053] A state trigger module, connected to the sequence interpretation module, is configured to perform modifier key state control and combined key event triggering processing on the modifier key and direction key combination command intention and combine direction key information to obtain a target key event sequence;

[0054] A command verification module, connected to the state trigger module, is configured to perform command verification processing on the target key event sequence by applying a verification rule including intention confirmation, context rationality verification, and duplicate command suppression to obtain a verified target key event sequence;

[0055] An event injection and feedback module, connected to the command verification module, is configured to perform target platform event conversion and injection processing on the verified target key event sequence, and synchronously perform multimodal feedback signal generation processing corresponding to the modifier key and direction key combination command intention or execution state, so as to output a standard input event injected into the target platform and the synchronously generated multimodal feedback signal.

[0056] To achieve the above object, an embodiment of the present application further provides a device for mapping three-dimensional motion gestures to key combinations, including a memory, a processor, and a program for mapping three-dimensional motion gestures to key combinations stored on the memory and executable on the processor. When the processor executes the program for mapping three-dimensional motion gestures to key combinations, the method for mapping three-dimensional motion gestures to key combinations as described in any one of the above is implemented.

[0057] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, on which a program for mapping three-dimensional motion gestures to key combinations is stored. When the program for mapping three-dimensional motion gestures to key combinations is executed by a processor, the method for mapping three-dimensional motion gestures to key combinations as described in any one of the above is implemented.

[0058] The method for mapping three-dimensional motion gestures to modifier key and direction key combinations in the present application has the following beneficial effects:

[0059] First, by interpreting the time series of basic gesture identifiers using a preset sequence matching rule based on the current interaction context, the present application can generate a large number of complex combination commands including modifier keys and arrow keys through different time series combinations using a limited number of basic gestures that are easy for users to learn and execute. This effectively overcomes the mismatch problem between the low-dimensional gesture input channel and the high-dimensional device control requirements, avoids the cumbersome mode switching or hierarchical menu navigation in traditional applications, greatly expands the scope of the user's ability to perform fine control through simple actions, and enhances the expressiveness and efficiency of input.

[0060] Second, the modifier key state control and combined key event trigger processing included in the present application, especially its ability to distinguish between continuous trigger, instantaneous trigger, and locked trigger modes, can accurately simulate the behaviors of pressing, holding, releasing, and locking modifier keys on a physical keyboard. This fine management of the modifier key state enables users to complete complex operations that require long-term holding of modifier keys or quick operations that require precise combinations through gestures, improving the fidelity and practicality of gesture input for simulating keyboard behavior and meeting the real needs of combined key control in practical applications.

[0061] Furthermore, the present application integrates multiple technical features for improving robustness in the signal processing and recognition process. The multi-level signal processing including motion state recognition can effectively filter out environmental noise and interference from the user's unconscious body movements at the preprocessing stage, and only transmit potential gesture signals with clear intentions to the subsequent recognition module, reducing the computational load and false trigger rate of subsequent processing. The SVM classifier and incremental learning mechanism based on small sample training used in basic gesture recognition enable the system to quickly establish a personalized model for new users and have good adaptability to changes in the user's operation habits. The multi-level verification rules included in the command verification process, including intention confirmation, context rationality verification, and duplicate command suppression, can intercept potential recognition errors, unreasonable operations, or accidental repetitions before the command is executed, further ensuring the accuracy and security of the interaction.

[0062] Finally, the present application synchronously generates multi-modal feedback signals corresponding to the command intention or execution status, and optimizes the feedback effect through an adaptive adjustment mechanism. This clear and timely feedback (tactile, auditory, visual) enables users to clearly understand the gestures currently recognized by the system, the sequences being constructed, the confirmed command intentions, and the states of the modifier keys, reducing the user's cognitive load and operational uncertainty, and enhancing the learnability of the system and the user experience. Description of the Drawings

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0064] Figure 1 FIG. 4 is a module structure diagram of an embodiment of a device for mapping three-dimensional motion gestures to key combinations according to the present invention;

[0065] Figure 2 FIG. 8 is a flowchart of an embodiment of a method for mapping three-dimensional motion gestures to key combinations according to the present invention.

[0066] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0067] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0068] To better understand the above technical solutions, the exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0069] It should be noted that in the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" in the text does not exclude the presence of components or steps not listed in the claims. The quantifier "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware including several different components and by means of a computer appropriately programmed. In the unit claims listing several devices, several of these devices can be embodied by the same hardware item. The use of "first", "second", and "third" etc. does not denote any order and these words can be interpreted as names.

[0070] As Figure 1 shown, Figure 1 FIG. 29 is a schematic structural diagram of a server 1 (also called a device for mapping three-dimensional motion gestures to key combinations) in a hardware operating environment related to the embodiment solution of the present invention.

[0071] The server according to the embodiment of the present invention, such as "Internet of Things devices", intelligent air conditioners with networking functions, intelligent electric lights, intelligent power supplies, AR / VR devices with networking functions, intelligent speakers, autonomous driving vehicles, PCs, smartphones, tablet computers, e-book readers, portable computers, and other devices with display functions.

[0072] As Figure 1 shown, the server 1 includes: a memory 11, a processor 12, and a network interface 13.

[0073] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 can be an internal storage unit of the server 1 in some embodiments, such as the hard disk of the server 1. The memory 11 can also be an external storage device of the server 1 in other embodiments, such as a plug-in hard disk equipped on the server 1, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0074] Furthermore, the memory 11 can also include an internal storage unit of the server 1 and an external storage device. The memory 11 can be used not only to store application software installed on the server 1 and various types of data, such as the code of the program 10 that maps three-dimensional motion gestures to key combinations, but also to temporarily store data that has been output or will be output.

[0075] The processor 12 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 11 or process data, such as executing the program 10 that maps three-dimensional motion gestures to key combinations.

[0076] The network interface 13 can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the server 1 and other electronic devices.

[0077] The network can be the Internet, a cloud network, a Wi-Fi network, a personal area network (PAN), a local area network (LAN), and / or a metropolitan area network (MAN). Various devices in the network environment can be configured to connect to the communication network according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols can include, but are not limited to, at least one of the following: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Li-Fi, 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocol, and / or Bluetooth communication protocol, or a combination thereof.

[0078] Optionally, the server may further include a user interface, and the user interface may include a display, an input unit such as a keyboard. Optionally, the user interface may further include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be referred to as a display screen or a display unit, and is used to display the information processed in the server 1 and to display a visual user interface.

[0079] Figure 1 Only the server 1 with components 11-13 and the program 10 that maps three-dimensional motion gestures to key combinations is shown. Those skilled in the art can understand that Figure 1 the shown structure does not constitute a limitation on the server 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0080] In this embodiment, the processor 12 can be used to call the program stored in the memory 11 that maps three-dimensional motion gestures to key combinations, and perform the following operations:

[0081] Perform multi-level signal processing on the raw three-dimensional motion data collected from the inertial measurement unit to obtain a standardized motion feature stream;

[0082] Perform basic gesture recognition on the standardized motion feature stream to obtain a basic gesture identifier time series;

[0083] Apply a preset sequence matching rule based on the current interaction context to the timing sequence of the basic gesture recognition identifiers, and obtain a modifier key and direction key combination command intention corresponding to the confirmation sequence;

[0084] Perform modifier key state control and combined key event triggering processing on the modifier key and direction key combination command intention, and combine the direction key information to obtain a target key event sequence;

[0085] Apply a verification rule including intention confirmation, context rationality verification, and repeated command suppression to the target key event sequence for command verification processing, and obtain a target key event sequence that passes the verification;

[0086] Perform target platform event conversion and injection processing on the target key event sequence that passes the verification, and simultaneously perform multimodal feedback signal generation processing corresponding to the modifier key and direction key combination command intention or its execution state, to obtain a standard input event injected into the target platform and the simultaneously generated multimodal feedback signal.

[0087] In one embodiment, the processor 12 can be used to call a program stored in the memory 11 that maps three-dimensional motion gestures to key combinations, and perform the following operations:

[0088] Apply a digital low-pass filter to the original three-dimensional motion data for filtering processing to obtain filtered three-dimensional motion data;

[0089] Apply a Kalman filter to the filtered three-dimensional motion data for sensor fusion processing to generate fusion motion data containing real-time attitude information;

[0090] Apply a motion state detection algorithm to the fusion motion data for processing. By calculating the short-time energy, signal amplitude variance, and main frequency components of the fusion motion data within a sliding time window, and comparing with a motion activation threshold dynamically adjusted based on the user's recent activity level, identify and screen out potential gesture signal segments representing the user's intended gesture execution;

[0091] Map the data of the potential gesture motion segment to a preset range to obtain the standardized motion feature stream.

[0092] In one embodiment, the processor 12 can be used to call a program stored in the memory 11 that maps three-dimensional motion gestures to key combinations, and perform the following operations:

[0093] Apply a dynamic segmentation algorithm based on signal variance and derivative threshold determination to the standardized motion feature stream for processing, and real-time define the start and end time points of a single basic gesture action to obtain gesture time segments;

[0094] Extract features from the fused motion data within the gesture time segment, and calculate a multi-dimensional feature set including time-domain statistics, frequency-domain distribution characteristics, and autoregressive model coefficients;

[0095] Apply the principal component analysis method to the multi-dimensional feature set for dimensionality reduction processing to obtain the gesture feature vector after dimensionality reduction;

[0096] Input the gesture feature vector after dimensionality reduction into a pre-trained linear kernel support vector machine classification model for classification processing to determine the basic gesture category corresponding to the gesture time segment;

[0097] Organize the continuously determined basic gesture categories in chronological order to obtain the time sequence of the basic gesture identifiers.

[0098] In one embodiment, the processor 12 can be used to call the program stored in the memory 11 that maps three-dimensional motion gestures to key combinations and perform the following operations:

[0099] Obtain the current active application information by querying the operating system application programming interface to determine the current interaction context;

[0100] According to the current interaction context, select and load an activated sequence matching rule library from multiple predefined context-specific sequence rule libraries, and the rule library contains multiple target gesture sequences and their corresponding command intents;

[0101] Append the latest identifier in the basic gesture identifier time sequence to the currently constructed user input sequence;

[0102] Perform real-time prefix matching processing on the currently constructed user input sequence and all target gesture sequences in the activated sequence matching rule library, and dynamically maintain a candidate sequence set;

[0103] Detect whether the candidate sequence set is reduced to a uniquely matched target gesture sequence, or detect whether a preset confirmation waiting time has passed after the last identifier in the time sequence is input to determine the final confirmation sequence;

[0104] According to the confirmation sequence, retrieve and obtain the corresponding modifier key and direction key combination command intent from the activated sequence matching rule library.

[0105] In one embodiment, the processor 12 can be used to call the program stored in the memory 11 that maps three-dimensional motion gestures to key combinations and perform the following operations:

[0106] Parse the command intent to identify the specific modifier key and specific direction key included therein;

[0107] Determine the modifier key trigger mode according to the command intention:

[0108] If the command intention requires continuously pressing the specific modifier key, activate the continuous trigger mode, and control the press and release states of the specific modifier key by continuously monitoring and maintaining the stability of the basic gesture signal of the specific modifier key and applying a hysteresis threshold for determination.

[0109] If the command intention requires an instantaneous trigger combination, activate the instantaneous trigger mode;

[0110] If it is detected that a specific sequence is used to switch the lock state, activate the lock trigger mode to switch and maintain the locked or unlocked state of the modifier key;

[0111] Generate the target key event sequence including the keyboard scan code and action according to the activated trigger mode and the parsed arrow key information, in the order and timing required by the target operating system.

[0112] In one embodiment, the processor 12 can be used to call the program stored in the memory 11 that maps three-dimensional motion gestures to key combinations, and perform the following operations:

[0113] Evaluate the classification confidence of the basic gesture on which the target key event sequence is generated and the matching confidence of the confirmation sequence for intention confirmation verification;

[0114] Compare the operation corresponding to the target key event sequence with the behavior rules preset based on the current interaction context for context rationality verification;

[0115] Detect whether the target key event sequence is exactly the same as the event sequence generated within a preset short time window for duplicate command suppression verification;

[0116] Based on the verification results of the intention confirmation, context rationality, and duplicate command suppression, obtain the target key event sequence that passes the verification.

[0117] In one embodiment, the processor 12 can be used to call the program stored in the memory 11 that maps three-dimensional motion gestures to key combinations, and perform the following operations:

[0118] Perform target platform event conversion processing on the target key event sequence that passes the verification to obtain a standard input event that conforms to the input specification of the target operating system;

[0119] Call the underlying input simulation interface function of the target operating system, and inject the standard input event data structure into the system event queue according to the parsing timing to obtain the standard input event injected into the target platform;

[0120] Query a preset feedback mapping table according to the combined command intention of the modifier key and the direction key or the execution status of the target key event sequence, and generate a feedback instruction including a specified feedback channel, a feedback mode, and a feedback intensity level, where the feedback channel includes at least one of a tactile channel, an auditory channel, and a visual channel, and the feedback mode includes at least one of a preset vibration waveform, an audio sample, and a visual element style;

[0121] Apply adaptive adjustment processing to the feedback instruction, adjust the intensity level or channel priority in the feedback instruction according to real-time environmental parameters obtained from an environmental sensor or preference parameters set by a user through a configuration interface, and then output through a corresponding feedback actuator to obtain the synchronously generated multimodal feedback signal.

[0122] Based on the hardware architecture of the device that maps three-dimensional motion gestures to key combinations as described above, an embodiment of the method of the present invention for mapping three-dimensional motion gestures to key combinations is proposed. Refer to Figure 2 , Figure 2 This is an embodiment of the method of the present invention for mapping three-dimensional motion gestures to key combinations. The method of mapping three-dimensional motion gestures to key combinations includes the following steps:

[0123] S10. Perform multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit to obtain a standardized motion feature stream. The original data here usually includes the three-axis acceleration and angular velocity information provided by the accelerometer and gyroscope, which reflects the motion state of the user's limb in three-dimensional space. In order to facilitate subsequent gesture recognition and mapping, these original data must be converted into a consistent and processable feature form.

[0124] This process is detailed in steps S11 to S14 below and is illustrated with examples.

[0125] In step S11, apply a digital low-pass filter to the original three-dimensional motion data for filtering to obtain filtered three-dimensional motion data. The original data often contains high-frequency noise, such as device jitter or environmental vibration interference, which will affect the accuracy of subsequent analysis. The digital low-pass filter smooths the data by retaining low-frequency signals (related to the user's intentional actions) and attenuating high-frequency components. For example, assume that the IMU samples data at a sampling rate of 100Hz, and the user performs a slow wrist rotation action. The original acceleration data may contain sharp high-frequency peaks. By setting a low-pass filter with a cut-off frequency of 5Hz, the fast jitter noise can be effectively filtered out, and a smooth motion curve can be retained. In this way, the filtered data can better reflect the user's true motion intention.

[0126] In step S12, the Kalman filter is applied to the filtered three-dimensional motion data for sensor fusion processing to generate fusion motion data containing real-time attitude information. The accelerometer and gyroscope data of the IMU each have limitations: the accelerometer is affected by gravity, and the gyroscope has drift. The Kalman filter estimates the real-time attitude of the device (such as Euler angles or quaternions) by integrating the advantages of both. Taking wrist rotation as an example, the filtered accelerometer data may show the inclination trend of the wrist, while the gyroscope data provides the rotational angular velocity. The Kalman filter combines the two to output a smooth and accurate attitude sequence, such as the yaw angle change at each moment. This fusion process ensures that subsequent analysis is based on a reliable motion trajectory rather than the unstable output of a single sensor.

[0127] In step S13, a motion state detection algorithm is applied to the fusion motion data for processing. By calculating the short-time energy, signal amplitude variance, and main frequency components within a sliding time window and comparing them with a dynamically adjusted motion activation threshold, gesture signal segments of the user's intention are identified. When the user uses the IMU, the data contains both intentional gestures and static or unintentional movements. The motion state detection algorithm aims to distinguish these states. For example, assume the user performs a "wave right" gesture. The fusion data shows a significant increase in short-time energy (sum of signal squares), a larger amplitude variance, and the main frequency components concentrated in 1 - 2 Hz (typical gesture frequency) within a certain time window (such as 0.5 seconds). By comparing with a threshold dynamically adjusted based on the user's recent activity level (such as the average energy in the past 5 seconds), it can be determined that this segment of data is a potential gesture signal. This dynamic threshold design adapts to individual differences and environmental changes, ensuring that the selected signal segments truly reflect the user's intention.

[0128] In step S14, the data of the potential gesture motion segment is mapped to a preset range to obtain a normalized motion feature stream. Since different users may execute the same gesture with different intensities, or the same user may also vary in intensity in different situations, directly using the values of the original amplitudes will cause trouble for the subsequent recognition model. Therefore, normalization processing is required to scale the data values (e.g., the acceleration or angular velocity in a certain axis) to a fixed interval, such as [-1, 1] or [0, 1], or to make it conform to the standard normal distribution (mean is 0, standard deviation is 1). Common methods include Min - Max normalization (formula: (value - min) / (max - min), and then adjusted to the target range) or Z - score normalization (formula: (value - mean) / std_dev). Here, min, max, mean, and std_dev are usually calculated based on the data within the currently recognized "potential gesture signal segment". For example, within a detected gesture segment, the X - axis angular velocity range is from - 150 degrees / second to + 200 degrees / second. After using Min - Max normalization to the [-1, 1] range, the point with the original value of 0 degrees / second will be mapped to ((0 - (-150)) / (200 - (-150)))*2 - 1 = (150 / 350)*2 - 1 ≈ 0.857 - 1 = - 0.143. After such processing, regardless of the original action amplitude, the relative form of its motion pattern is retained, but the values are unified into the standard range, forming a normalized motion feature stream.

[0129] For example, assume that the user wears an IMU bracelet and executes "wave right" to trigger the "Ctrl + right" command. Initially, the IMU collects raw data at 100Hz, recording the acceleration [x, y, z] and the angular velocity [gx, gy, gz]. In step S11, assume that the original acceleration x - axis data is [0.1, 0.3, 1.5, 2.0, 0.5] m / s 2 (including noise). After applying a 5Hz low - pass filter, the smoothed sequence [0.1, 0.2, 1.0, 1.8, 0.6] m / s is obtained 2 . Then in step S12, combining with the angular velocity data (e.g., gx = [0, 0.1, 0.5, 0.8, 0.2] rad / s), the Kalman filter fuses and outputs the attitude angle sequence, such as the yaw angle [0°, 2°, 10°, 18°, 6°]. In step S13, calculate the short - time energy within a 0.5 - second window (∑x 2= 5.24), variance (0.65), main frequency (1.5 Hz), compared with the dynamic threshold (energy 2.0), it is confirmed that this segment is a gesture signal. Finally, in step S14, the acceleration sequence is normalized, with a mean of 0.74 and a standard deviation of 0.65, and the normalized values [-1.0, -0.83, 0.4, 1.6, -0.22] are output to form a feature stream.

[0130] It can be understood that these steps significantly improve the conversion quality from raw data to feature stream through multi-level processing. The digital low-pass filter effectively removes noise to ensure that the data reflects real actions; the sensor fusion of the Kalman filter provides accurate postures to make up for the deficiencies of a single sensor; the motion state detection algorithm combines dynamic thresholds to accurately screen intention signals to adapt to user differences; the normalization process unifies the data scale to lay a foundation for subsequent recognition. This solution not only improves the signal quality and the robustness of gesture detection, but also enhances the universality of the system, enabling it to be reliably applied to diverse interaction scenarios.

[0131] S20. Perform basic gesture recognition on the normalized motion feature stream to obtain a sequence of basic gesture identifiers in time series. The core of this process is to decompose continuous motion data into discrete gesture units and accurately identify the action categories represented by each unit, and finally form an ordered gesture sequence.

[0132] The following gradually explains its implementation process through steps S21 to S25 and illustrates it with examples.

[0133] In step S21, apply a dynamic segmentation algorithm based on signal variance and derivative thresholds to the normalized motion feature stream to real-time define the start and end time points of a single basic gesture and obtain a gesture time segment. The normalized feature stream is a continuous signal stream that contains multiple gesture and non-gesture parts. The dynamic segmentation algorithm detects action boundaries by analyzing the statistical characteristics of the signal. For example, assume the feature stream is a set of normalized acceleration values, which are close to 0 during calmness and have large fluctuations when waving. Calculate the signal variance within a sliding window (e.g., the variance rises from 0.01 to 0.5 within a 0.1-second window) and the derivative (the rate of value change exceeds the threshold 0.2). When both the variance and the derivative exceed the preset threshold, it is marked as the start of a gesture; when both of them fall back, it is marked as the end. Taking "wave right" as an example, the feature stream may have a variance of 0.6 and a derivative of 0.3 at t = 1 s, and fall back at t = 1.5 s, segmenting out a gesture segment from 1 s to 1.5 s. This method ensures real-time performance and segmentation accuracy.

[0134] Next, in step S22, feature extraction is performed on the fused motion data within the gesture time segment to calculate a multi-dimensional feature set including time-domain statistics, frequency-domain distribution characteristics, and autoregressive model coefficients. The fused motion data (including pose information as described above) is the segmented input, and feature extraction aims to capture the unique patterns of the gesture. The time-domain statistics include mean, variance, peak value, etc.; the frequency-domain characteristics calculate the main frequency through the Fast Fourier Transform (FFT); the autoregressive (AR) model coefficients reflect the dynamic law of the time series. For example, the mean acceleration of the "wave right" segment may be 0.8, the variance is 0.5, the FFT shows the main frequency of 1.5 Hz, and the second-order coefficient of the AR model is [0.9, -0.2]. These features together form a high-dimensional vector that comprehensively describes the gesture characteristics.

[0135] Then, in step S23, principal component analysis (PCA) is applied to the multi-dimensional feature set for dimensionality reduction to obtain the reduced-dimensional gesture feature vector. Although the high-dimensional feature set is rich in information, too many dimensions will lead to complex calculations and overfitting. PCA retains the main variant components through linear transformation to reduce the dimension. For example, assume the original feature set has 10 dimensions (3 in the time domain, 3 in the frequency domain, and 4 in the AR coefficients). After PCA analysis, the first 3 principal components are retained, explaining 90% of the variance, and the output feature vector is like [1.2, -0.5, 0.3]. This not only reduces the computational burden but also highlights the key differences of the gesture, improving the subsequent classification efficiency.

[0136] In step S24, the reduced-dimensional feature vector is input into a pre-trained linear kernel support vector machine (SVM) classification model to determine the basic gesture category corresponding to the gesture segment. The linear kernel SVM distinguishes different gesture categories by finding the optimal hyperplane in the feature space. The model is pre-trained with labeled data. For example, "wave right" is labeled as R, and "raise hand up" is labeled as U. Assume the input vector is [1.2, -0.5, 0.3], and the SVM determines it belongs to the R category according to the trained boundary. This classification method is simple and efficient and is suitable for real-time applications.

[0137] Finally, in step S25, the continuously determined basic gesture categories are organized in chronological order to obtain the chronological sequence of basic gesture identifiers. After each gesture segment is classified, they are arranged in the order of occurrence. For example, if the user sequentially performs "wave right", "raise hand up", "wave right", and the classification results are R, U, R, the sequence is [R, U, R]. This provides a structured input for subsequent command mapping.

[0138] Exemplarily, assume that the user wears an IMU bracelet and performs "wave right" and "raise hand up" in sequence. In step S21, the normalized feature stream (the output of the previous case [-1.0, -0.83, 0.4, 1.6, -0.22]) is segmented, and it is detected that the variance is 0.6 and the derivative is 0.3 from t = 1s to 1.5s for "wave right", and the variance is 0.5 and the derivative is 0.25 from t = 2s to 2.4s for "raise hand up". In step S22, the features extracted for "wave right" are: mean 0.8, variance 0.5, main frequency 1.5Hz, AR coefficient [0.9, -0.2]; the features for "raise hand up" are mean 0.6, variance 0.4, main frequency 1Hz, AR coefficient [0.8, -0.3]. In step S23, after PCA dimensionality reduction, the vector for "wave right" is [1.2, -0.5, 0.3], and the vector for "raise hand up" is [0.9, 0.6, -0.1]. In step S24, the SVM classification determines as R and U. In step S25, the sequence output is [R, U].

[0139] It can be understood that these steps achieve an accurate conversion from continuous signals to gesture sequences through dynamic segmentation, feature extraction and dimensionality reduction, classification, and serialization. The segmentation based on variance and derivative captures the action boundaries in real time, the multi-dimensional feature set comprehensively describes the gesture characteristics, PCA dimensionality reduction optimizes the calculation efficiency, SVM classification ensures accuracy, and the time series organization lays the foundation for subsequent mapping. This solution significantly improves the robustness and real-time performance of gesture recognition, enabling the system to adapt to complex interaction requirements.

[0140] S30. Apply the preset sequence matching rules based on the current interaction context to interpret the time sequence of the basic gesture identifiers, and obtain the combination command intention of the modifier key and the direction key corresponding to the confirmation sequence. The key to this process is to convert a simple action sequence into a complex keyboard input, and use context information and dynamic matching to improve accuracy and usability.

[0141] The following gradually illustrates its implementation process through steps S31 to S36 and elaborates with examples.

[0142] In step S31, the current active application information is obtained by querying the operating system application programming interface (API) to determine the current interaction context. The user's operation intention is usually closely related to the application being used. For example, in a text editor, the user may need to press "Ctrl+Right" to move the cursor, while in a browser, the user may need to press "Alt+Down" to switch tabs. The API (such as GetForegroundWindow in Windows) can return the identifier of the current active window. Assuming the user is using "Microsoft Word" and the API returns "winword.exe", the context is determined to be a text editing scenario. This approach ensures that subsequent processing is aligned with the user's current task.

[0143] Next, in step S32, according to the current interaction context, an activated sequence matching rule library is selected and loaded from multiple predefined context-specific sequence rule libraries. This rule library contains multiple target gesture sequences and their corresponding command intentions. For example, the text editing rule library may define: [R,R] as "Ctrl+Right", [U,D] as "Ctrl+Down"; while the browser rule library may be: [R,R] as "Alt+Right", [U,D] as "Alt+Down". If the context is "Microsoft Word", the loaded rule library contains text editing mappings. This context-based rule selection makes the same gesture have different meanings in different scenarios, enhancing the flexibility of the system.

[0144] Then, in step S33, the latest identifier in the base gesture identifier time series is appended to the currently constructed user input sequence. The user's gestures are input step by step, and the system needs to update the sequence in real time to reflect the latest actions. For example, assuming the previous sequence [R] ( "wave right") is obtained and the user then performs "wave down" (recognized as D), the sequence becomes [R,D] after appending D. This step provides a dynamically updated input basis for subsequent matching.

[0145] In step S34, the currently constructed user input sequence is matched in real time with the target sequences in the activated rule library, and a candidate sequence set is dynamically maintained. The prefix match checks whether the current sequence is the beginning of a complete sequence in the rule library. For example, the rule library contains: [R,D] → "Ctrl+Down", [R,R] → "Ctrl+Right", the current sequence [R] matches {[R,D], [R,R]}, and after appending D, [R,D] only matches {[R,D]}. The candidate set is updated with the input. For example, there are more candidates when it is [R], and it shrinks when it is [R,D]. This real-time matching supports quick response to user actions.

[0146] Next, in step S35, it is detected whether the candidate sequence set is reduced to the uniquely matched target sequence, or whether a preset confirmation waiting time has elapsed after the last identifier of the timing sequence is input, so as to determine the final confirmation sequence. When there is only one match in the set, the intention is clear. For example, [R,D] corresponds to "Ctrl+down". If there are still multiple candidates (such as [R] may continue to expand), the system waits for a period of time (such as 0.5 seconds). If there is no new input, the current sequence is confirmed. Assume that the user pauses for 0.5 seconds after inputting [R,D], and the set is reduced to [R,D], and the confirmed sequence is [R,D]. This ensures the accuracy and timeliness of intention determination.

[0147] Finally, in step S36, according to the confirmed sequence, the corresponding modifier key and arrow key combination command intention is retrieved from the activated rule library. Taking [R,D] as an example, the rule library returns "Ctrl+down". If the sequence is [R,R], then "Ctrl+right" is returned. This step completes the mapping from the gesture sequence to the keyboard command and generates an executable intention.

[0148] Exemplarily, assume that the user inputs [R,D] in "Microsoft Word" to trigger "Ctrl+down". In step S31, the API confirms that the current application is "winword.exe" and the context is text editing. In step S32, the rule library is loaded: [R,D] → "Ctrl+down", [R,R] → "Ctrl+right". In step S33, the user inputs R, and the sequence is [R]. Then the user inputs D, and the sequence is updated to [R,D]. In step S34, [R] matches {[R,D],[R,R]}, and [R,D] matches {[R,D]}. In step S35, the user pauses for 0.5 seconds, and the confirmed sequence is [R,D]. In step S36, "Ctrl+down" is retrieved.

[0149] It can be understood that these steps achieve an accurate conversion from the gesture sequence to the complex command through the context awareness and dynamic matching mechanism. The API obtains the context to ensure the matching of the rule and the scenario. The rule library selection provides flexibility. The prefix matching and candidate maintenance achieve real-time performance. The confirmation waiting balances speed and accuracy. Finally, the retrieval generates a reliable intention. This solution significantly improves the adaptability and efficiency of the interaction, and is particularly suitable for scenarios that require diverse keyboard inputs.

[0150] S40. For the modifier key and arrow key combination command intention, perform modifier key state control and combined key event trigger processing and combine arrow key information to obtain a target key event sequence. This process aims to convert the abstract command intention into a specific keyboard input recognizable by the operating system, and at the same time flexibly handle different trigger requirements of the modifier keys.

[0151] The implementation process is gradually described below through steps S41 to S46 and explained with examples.

[0152] In step S41, the command intention is parsed to identify the specific modifier keys and direction keys it contains. The command intention is the result of the previous mapping, such as "Shift + down" or "Ctrl + right", and needs to be decomposed into specific elements. Taking "Shift + down" as an example, after parsing, the modifier key is identified as "Shift" and the direction key is "down". This step lays the foundation for subsequent processing to ensure that the system clearly understands the components of the operation.

[0153] Next, in step S42, the triggering mode of the modifier key is determined according to the command intention. The modifier key can be triggered in different ways: held continuously (such as for multi-selection operations), triggered instantaneously (such as for shortcut keys), or in a locked state (such as for fixed modifiers). The system determines the applicable mode based on the semantics of the intention. For example, "Shift + down" is usually instantaneously triggered, while "Shift held + down" may indicate continuous selection, and the determination result guides subsequent execution.

[0154] In step S43, if the command intention requires holding the modifier key continuously, the continuous triggering mode is activated. By continuously monitoring to maintain the stability of the technical gesture signal of the modifier key and applying a hysteresis threshold determination, the press and release states of the modifier key are controlled. For example, for the user intention "Shift held + down" to select multiple lines of text, the system monitors the angular velocity variance of the gesture (such as wrist depression). If it is below the threshold of 0.05, Shift is held down, and when the gesture changes (such as the variance rises to 0.2), it is released. The hysteresis threshold prevents frequent switching and ensures state stability.

[0155] Then, in step S44, if the command intention requires an instantaneous trigger combination, the instantaneous trigger mode is activated. This mode is suitable for single operations, such as "Shift + down" to select the next line. The system directly generates press and release events without continuous monitoring. Taking "Shift + down" as an example, the action is completed immediately after triggering, which is suitable for fast interaction.

[0156] In step S45, if a specific sequence is detected for switching the locked state, the locked trigger mode is activated to switch and maintain the locked or unlocked state of the modifier key. For example, the rule defines that "raising the hand upward + waving the hand downward" (sequence [U, D]) switches "Shift locked". After the user inputs [U, D], Shift remains pressed until the unlock is triggered again. This mode supports long-term modifier key requirements.

[0157] Finally, in step S46, according to the activated trigger mode and the parsed direction key information, a target key event sequence including keyboard scan codes and actions is generated in the order and timing required by the target operating system. The operating system requires specific scan codes (e.g., Shift = 0x10, Down = 0x28) and event order. Taking the instantaneous mode of "Shift + Down" as an example, the generated sequence is: [0x10 press, 0x28 press, 0x28 release, 0x10 release], with an interval of 10 ms, which conforms to the system specification.

[0158] For example, assume that the user enters "Shift + Down" in a text editor to select the next line. In step S41, "Shift + Down" is parsed to identify Shift and Down. In step S42, it is determined to be an instantaneous trigger. Steps S43 and S45 are not applicable, and step S44 is entered to activate the instantaneous mode. In step S46, an event sequence is generated: [Shift press (0x10), Down press (0x28), Down release (0x28), Shift release (0x10)], with a 10 ms interval for each step, adapting to Windows input.

[0159] It can be understood that these steps achieve an efficient conversion from commands to input through intention parsing, flexible mode determination, and specification event generation. The operation elements are clearly parsed, multiple trigger modes meet different scenarios, signal monitoring and hysteresis thresholds ensure the stability of continuous modes, and the scan code sequence ensures compatibility. This solution significantly improves the accuracy and diversity of command execution, especially suitable for complex interaction requirements.

[0160] S50. Apply a verification rule including intention confirmation, context rationality verification, and duplicate command suppression to the target key event sequence for command verification processing to obtain a verified target key event sequence. This process aims to filter out possible misoperations or unnecessary repeated inputs and generate a verified event sequence.

[0161] The following gradually explains its implementation process through steps S51 to S54 and illustrates it with examples.

[0162] In step S51, the confidence of the basic gesture classification on which the target key event sequence is generated and the confidence of the matching of the confirmation sequence are evaluated for intention confirmation verification. Both gesture classification (e.g., "wave right" is recognized as R) and sequence matching (e.g., [R, U] maps to "Ctrl + up") have confidence scores. For example, the confidence of the SVM classifier outputting R is 0.9, and the confidence of the sequence [R, U] matching the rule is 0.95. If the confidence threshold is set to 0.8, both pass the verification. This step ensures that the input originates from reliable gesture recognition and intention mapping, avoiding mis-triggering with low confidence.

[0163] Next, in step S52, the operations corresponding to the target key event sequence are compared with the behavior rules based on the current interaction context for context rationality verification. The context is determined by the API mentioned above. For example, in the "Notepad" scenario, the behavior rule allows moving the cursor with "Ctrl + Right Arrow", but closing the window with "Alt + F4" may be restricted. Suppose the sequence is [Ctrl pressed, Right Arrow pressed, Right Arrow released, Ctrl released], corresponding to "Ctrl + Right Arrow", which is consistent with the text editing rule and passes the verification. If the sequence is "Alt + F4", it may be rejected due to inconsistent context. This step ensures that the operations conform to the current application logic.

[0164] Then, in step S53, it is detected whether the target key event sequence is exactly the same as the event sequence generated within a preset short time window for duplicate command suppression verification. The user may repeatedly trigger the same command due to gesture jitter. For example, two "Ctrl + Right Arrow" are continuously generated within a 0.3 - second window. The system compares the current sequence [Ctrl pressed, Right Arrow pressed, Right Arrow released, Ctrl released] with the previous sequence. If they are exactly the same and the time interval is less than the threshold (such as 0.5 seconds), the latter one is suppressed. This avoids unnecessary repeated execution and improves the interaction fluency.

[0165] Finally, in step S54, based on the verification results of intention confirmation, context rationality, and duplicate command suppression, the target key event sequence that passes the verification is obtained. When all three verifications pass, the sequence is accepted; if any one fails, it is discarded or adjusted. For example, for the "Ctrl + Right Arrow" sequence, the classification confidence is 0.9 (pass), the context is reasonable (pass), and there is no duplication (pass), then the sequence that passes the verification is output. If the confidence is only 0.7 (lower than 0.8), it is rejected. This step integrates multi - dimensional verification to ensure reliable output.

[0166] Exemplarily, assume that the user inputs "Ctrl + Right Arrow" to move the cursor in "Notepad". In step S51, the classification confidences of gestures R and U are 0.9 and 0.85 respectively, and the confidence of the sequence [R, U] matching "Ctrl + Right Arrow" is 0.92, all exceeding the threshold of 0.8 and passing the verification. In step S52, "Ctrl + Right Arrow" is consistent with the text editing rule and passes the verification. In step S53, there is no same sequence [Ctrl pressed (0x11), Right Arrow pressed (0x27), Right Arrow released, Ctrl released] in the previous 0.5 seconds, passing the verification. In step S54, all three pass, and the sequence [0x11 pressed, 0x27 pressed, 0x27 released, 0x11 released] is output.

[0167] It can be understood that these steps significantly enhance the reliability and usability of the key event sequence through a multi-verification mechanism. Confidence assessment ensures the accuracy of gestures and intentions, context comparison guarantees the rationality of operations, duplicate suppression optimizes the user experience, and comprehensive judgment generates high-quality outputs. This solution effectively reduces misoperations and redundant inputs, and is particularly suitable for interactive scenarios that require precise control.

[0168] S60. For the target key event sequence that passes the verification, perform target platform event conversion and injection processing, and simultaneously perform multi-modal feedback signal generation processing corresponding to the modifier key and arrow key combination command intention or its execution status, to obtain the standard input event injected into the target platform and the simultaneously generated multi-modal feedback signal. This process not only completes the actual execution of keyboard input, but also enhances the user experience through feedback.

[0169] The following gradually explains its implementation process through steps S61 to S64 and illustrates it with examples.

[0170] In step S61, perform target platform event conversion processing on the target key event sequence that passes the verification to obtain a standard input event that conforms to the input specification of the target operating system. The target key event sequence (such as [Ctrl pressed (0x11), right pressed (0x27), right released, Ctrl released]) is represented using internal scan codes and needs to be converted into a format recognized by the operating system. For example, in Windows, the standard input event may be an INPUT structure, which contains key codes and status flags. After conversion, "Ctrl + right" generates an event list: [INPUT(Ctrl, pressed), INPUT(right, pressed), INPUT(right, released), INPUT(Ctrl, released)], ensuring compatibility with the system specification.

[0171] Next, in step S62, call the underlying input simulation interface function of the target operating system to inject the standard input event into the system event queue according to the parsing time sequence. Taking Windows as an example, use the SendInput function to inject the above INPUT structure sequence at 10ms intervals. After injection, the system recognizes "Ctrl + right" as a keyboard input and triggers the cursor to move to the right. This step realizes the transition from virtual commands to actual operations and ensures that events are executed as expected.

[0172] Then, in step S63, according to the execution status of the modifier key and arrow key combination command intention or the target key event sequence, query the preset feedback mapping table to generate a feedback instruction including a feedback channel, a mode, and an intensity. The feedback channels include tactile (such as vibration), auditory (such as sound), and visual (such as light); the mode may be a vibration waveform, an audio sample, or a visual style. For example, the "Ctrl+right" mapping table defines: tactile channel, short vibration (0.1 second), medium intensity; auditory channel, click sound, low intensity. The generated instruction is: [tactile, short vibration, medium], [auditory, click, low], providing multi-sensory confirmation.

[0173] Finally, in step S64, apply adaptive adjustment to the feedback instruction. After adjusting the intensity or channel priority according to the environmental sensor data or user preferences, output through the feedback actuator. For example, the environmental sensor detects a high noise level (such as 60 decibels), and the user prefers tactile feedback. The system will lower the auditory intensity to the minimum and increase the tactile to high level. The adjusted instruction is: [tactile, short vibration, high], [auditory, click, minimum], and output through the vibration motor and the speaker. This adaptability improves the comfort and effectiveness of the feedback.

[0174] For example, assume the user triggers "Ctrl+right" in "Notepad". In step S61, the sequence [0x11 pressed, 0x27 pressed, 0x27 released, 0x11 released] is converted into a Windows INPUT structure. In step S62, SendInput injects the event, and the cursor moves to the right. In step S63, the mapping table generates the instruction: [tactile, short vibration, medium], [visual, green light flashing, low]. In step S64, the environment is quiet and the user prefers tactile, so it is adjusted to: [tactile, short vibration, high], [visual, green light flashing, low], and the bracelet vibrates and the light flashes.

[0175] It can be understood that these steps achieve seamless input through event conversion and injection, and multi-modal feedback and adaptive adjustment enhance the user perception. The conversion ensures system compatibility, the injection completes the operation execution, the feedback mapping provides intuitive confirmation, and the adaptive adjustment optimizes the experience. This solution significantly enhances the reliability and friendliness of the interaction, and is particularly suitable for scenarios that require real-time feedback.

[0176] In addition, the embodiment of the present invention also proposes a device for mapping three-dimensional motion gestures to key combinations. The device for mapping three-dimensional motion gestures to key combinations includes:

[0177] An inertial measurement unit interface for receiving the original three-dimensional motion data collected from the inertial measurement unit;

[0178] A signal processing module connected to the inertial measurement unit interface, configured to perform multi-level signal processing on the original three-dimensional motion data to obtain a normalized motion feature stream;

[0179] A basic gesture recognition module, connected to the signal processing module, is configured to perform basic gesture recognition on the standardized motion feature stream to obtain a sequence of basic gesture recognition identifiers;

[0180] A sequence interpretation module, connected to the basic gesture recognition module, is configured to interpret the sequence of basic gesture recognition identifiers by applying a preset sequence matching rule based on the current interaction context to obtain a modifier key and direction key combination command intention corresponding to the confirmation sequence;

[0181] A status trigger module, connected to the sequence interpretation module, is configured to perform modifier key status control and combined key event trigger processing on the modifier key and direction key combination command intention and combine the direction key information to obtain a target key event sequence;

[0182] A command verification module, connected to the status trigger module, is configured to perform command verification processing on the target key event sequence by applying a verification rule including intention confirmation, context rationality verification, and duplicate command suppression to obtain a verified target key event sequence;

[0183] An event injection and feedback module, connected to the command verification module, is configured to perform target platform event conversion and injection processing on the verified target key event sequence, and simultaneously perform multimodal feedback signal generation processing corresponding to the modifier key and direction key combination command intention or execution status to output a standard input event injected into the target platform and the simultaneously generated multimodal feedback signal.

[0184] Wherein, the steps implemented by each functional module of the device for mapping three-dimensional motion gestures to key combinations can refer to the respective embodiments of the method for mapping three-dimensional motion gestures to key combinations of the present invention, which will not be elaborated here.

[0185] In addition, an embodiment of the present invention also proposes a computer-readable storage medium. The computer-readable storage medium can be any one or any combination of a hard disk, a multimedia card, an SD card, a flash card, an SMC, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, etc. The computer-readable storage medium includes a program 10 for mapping three-dimensional motion gestures to key combinations. The specific implementation manner of the computer-readable storage medium of the present invention is substantially the same as the specific implementation manners of the above method for mapping three-dimensional motion gestures to key combinations and the server 1, which will not be elaborated here.

[0186] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0187] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0188] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0189] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0190] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0191] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for mapping three-dimensional motion gestures to key combinations, characterized in that, Including: Performing multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit to obtain a standardized motion feature stream; Performing basic gesture recognition on the standardized motion feature stream to obtain a time sequence of basic gesture identifiers; Interpreting the time sequence of basic gesture identifiers by applying a preset sequence matching rule based on the current interaction context to obtain a combination command intention of a modifier key and a direction key corresponding to a confirmation sequence; Performing modifier key state control and combined key event triggering processing on the combination command intention of the modifier key and the direction key and combining direction key information to obtain a target key event sequence; Performing command verification processing on the target key event sequence by applying a verification rule including intention confirmation, context rationality verification, and repeated command suppression to obtain a verified target key event sequence; Performing target platform event conversion and injection processing on the verified target key event sequence, and simultaneously performing multi-modal feedback signal generation processing corresponding to the combination command intention or execution state of the modifier key and the direction key to obtain a standard input event injected into the target platform and the simultaneously generated multi-modal feedback signal; 2. The method of mapping a three-dimensional motion gesture to a key combination according to claim 1, wherein Performing multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit, including: Applying a digital low-pass filter to the original three-dimensional motion data for filtering processing to obtain filtered three-dimensional motion data; Applying a Kalman filter to the filtered three-dimensional motion data for sensor fusion processing to generate fusion motion data including real-time attitude information; Applying a motion state detection algorithm to the fusion motion data for processing, and by calculating the short-time energy, signal amplitude variance, and main frequency components of the fusion motion data within a sliding time window and comparing with a motion activation threshold dynamically adjusted based on the user's recent activity level, identifying and screening out potential gesture signal segments representing the user's intended gesture execution; Mapping the data of the potential gesture motion segment to a preset range to obtain the standardized motion feature stream; 3. The method of mapping three-dimensional motion gestures to key combinations according to claim 1, wherein Performing basic gesture recognition on the standardized motion feature stream, including: Applying a dynamic segmentation algorithm based on signal variance and derivative threshold determination to the standardized motion feature stream for processing, real-time defining the start and end time points of a single basic gesture action to obtain a gesture time segment; Performing feature extraction processing on the fusion motion data within the gesture time segment to calculate a multi-dimensional feature set including time-domain statistics, frequency-domain distribution characteristics, and autoregressive model coefficients; Applying a principal component analysis method to the multi-dimensional feature set for dimensionality reduction processing to obtain a dimensionality-reduced gesture feature vector; Inputting the dimensionality-reduced gesture feature vector into a pre-trained linear kernel support vector machine classification model for classification processing to determine the basic gesture category corresponding to the gesture time segment; Organizing the continuously determined basic gesture categories in chronological order to obtain the time sequence of basic gesture identifiers; 4. The method for mapping three-dimensional motion gestures to key combinations as claimed in claim 1, wherein Interpreting the time sequence of basic gesture identifiers by applying a preset sequence matching rule based on the current interaction context, including: Obtain the current active application information by querying the operating system application programming interface to determine the current interaction context; Select and load an activated sequence matching rule library from multiple predefined context-specific sequence rule libraries according to the current interaction context, where the rule library contains multiple target gesture sequences and corresponding command intents; Append the latest identifier in the basic gesture identifier time series sequence to the currently constructed user input sequence; Perform real-time prefix matching processing on the currently constructed user input sequence and all target gesture sequences in the activated sequence matching rule library, and dynamically maintain a candidate sequence set; Detect whether the candidate sequence set is reduced to a uniquely matched target gesture sequence, or detect whether a preset confirmation waiting time has elapsed after the last identifier in the time series sequence is input to determine the final confirmation sequence; Retrieve and obtain the corresponding modifier key and arrow key combination command intent from the activated sequence matching rule library according to the confirmation sequence; 5. The method for mapping three-dimensional motion gestures to key combinations according to claim 1, characterized in that, Perform modifier key state control and combined key event triggering processing on the modifier key and arrow key combination command intent and combine arrow key information to obtain a target key event sequence, including: Parse the command intent to identify the specific modifier key and specific arrow key contained therein; Determine the triggering mode of the modifier key according to the command intent: If the command intent requires continuously holding down the specific modifier key, activate the continuous triggering mode, maintain the stability of the basic gesture signal of the specific modifier key by real-time monitoring and apply a hysteresis threshold determination to control the press and release states of the specific modifier key; If the command intent requires an instantaneous trigger combination, activate the instantaneous trigger mode; If it is detected that a specific sequence is used to switch the locking state, activate the locking trigger mode to switch and maintain the locked or unlocked state of the modifier key; Generate the target key event sequence including keyboard scan codes and actions according to the activated trigger mode and the parsed arrow key information in the order and time series required by the target operating system; 6. The method for mapping three-dimensional motion gestures to key combinations according to claim 1, wherein, Perform command verification processing on the target key event sequence by applying verification rules including intent confirmation, context rationality verification, and duplicate command suppression, including: Evaluate the classification confidence of the basic gesture on which the target key event sequence is generated and the matching confidence of the confirmation sequence for intent confirmation verification; Compare the operation corresponding to the target key event sequence with the preset behavior rules based on the current interaction context for context rationality verification; Detect whether the target key event sequence is exactly the same as the event sequence generated within a preset short time window for duplicate command suppression verification; Integrate the verification results of intent confirmation, context rationality, and duplicate command suppression to obtain the verified target key event sequence; 7. The method for mapping three-dimensional motion gestures to key combinations according to claim 1, wherein Perform target platform event conversion and injection processing on the verified target key event sequence, and simultaneously perform multi-modal feedback signal generation processing corresponding to the modifier key and arrow key combination command intent or execution state, including: Perform target platform event conversion processing on the verified target key event sequence to obtain standard input events that conform to the input specifications of the target operating system; Call the underlying input simulation interface function of the target operating system, and inject the standard input event data structure into the system event queue according to the parsing time sequence to obtain the standard input events injected into the target platform; According to the combination command intention of the modifier key and the direction key or the execution status of the target key event sequence, query the preset feedback mapping table to generate a feedback instruction including a specified feedback channel, feedback mode, and feedback intensity level, where the feedback channel includes at least one of a tactile channel, an auditory channel, and a visual channel, and the feedback mode includes at least one of a preset vibration waveform, audio sample, and visual element style; Apply adaptive adjustment processing to the feedback instruction, adjust the intensity level or channel priority in the feedback instruction according to the real-time environmental parameters obtained from the environmental sensor or the preference parameters set by the user through the configuration interface, and then output through the corresponding feedback actuator to obtain the synchronously generated multi-modal feedback signal.

8. A device for mapping three-dimensional motion gestures to key combinations, characterized in that, Includes: An inertial measurement unit interface for receiving the original three-dimensional motion data collected from the inertial measurement unit; A signal processing module, connected to the inertial measurement unit interface, configured to perform multi-level signal processing on the original three-dimensional motion data to obtain a standardized motion feature stream; A basic gesture recognition module, connected to the signal processing module, configured to perform basic gesture recognition on the standardized motion feature stream to obtain a sequence of basic gesture recognition symbols; A sequence interpretation module, connected to the basic gesture recognition module, configured to interpret the sequence of basic gesture recognition symbols by applying a preset sequence matching rule based on the current interaction context to obtain the combination command intention of the modifier key and the direction key corresponding to the confirmation sequence; A state trigger module, connected to the sequence interpretation module, configured to perform modifier key state control and combined key event triggering processing on the combination command intention of the modifier key and the direction key and combine the direction key information to obtain a target key event sequence; A command verification module, connected to the state trigger module, configured to perform command verification processing on the target key event sequence by applying verification rules including intention confirmation, context rationality verification, and duplicate command suppression to obtain a verified target key event sequence; An event injection and feedback module, connected to the command verification module, configured to perform target platform event conversion and injection processing on the verified target key event sequence, and synchronously perform multi-modal feedback signal generation processing corresponding to the combination command intention or execution status of the modifier key and the direction key to output the standard input events injected into the target platform and the synchronously generated multi-modal feedback signal.

9. A device for mapping three-dimensional motion gestures to key combinations, characterized in that, Includes a memory, a processor, and a program for mapping three-dimensional motion gestures to key combinations stored on the memory and executable on the processor. When the processor executes the program for mapping three-dimensional motion gestures to key combinations, it implements the method for mapping three-dimensional motion gestures to key combinations as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A program for mapping three-dimensional motion gestures to key combinations is stored on the computer-readable storage medium. When the program for mapping three-dimensional motion gestures to key combinations is executed by a processor, the method for mapping three-dimensional motion gestures to key combinations according to any one of claims 1-7 is implemented.