Mouse and keyboard key control method based on inertial measurement unit
By performing multi-level preprocessing and mode switching gesture recognition on the three-dimensional motion data of the inertial measurement unit and dynamically switching the control mode, the flexibility and coordination problems of cursor control and key operation in the existing technology are solved, and efficient and precise mouse and keyboard key control is achieved.
Patent Information
- Application Number
- CN202510696406.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-16
AI Technical Summary
Existing human-computer interaction technology based on inertial measurement units lacks flexibility and coordination in cursor control and key operations, and is difficult to adapt to diverse operating scenarios, resulting in low operating efficiency and insufficient accuracy.
By acquiring the three-dimensional motion data of the inertial measurement unit, performing multi-level preprocessing and multi-scale dynamic time window segmentation, identifying mode switching gestures, dynamically switching control modes, and combining with the hierarchical motion primitive library to parse the mouse and keyboard buttons and generate instructions.
It achieves seamless integration of precise cursor control, reliable gesture-based key input, and collaborative control of keyboard modifier keys in different operating scenarios, improving the consistency and smoothness of operations.
Smart Images

Figure CN120653133A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and in particular to a method and device for controlling a mouse and keyboard buttons based on an inertial measurement unit, and a computer-readable storage medium. Background Art
[0002] As a sensor technology capable of capturing the three-dimensional motion and posture of an object, the inertial measurement unit (IMU) demonstrates significant potential in the field of human-computer interaction, particularly in replacing or enhancing traditional input devices such as mice and keyboards. Existing technologies typically utilize IMUs that integrate accelerometers, gyroscopes, and sometimes magnetometers to acquire raw motion data. These IMUs then use specific pre-processing algorithms and mapping logic to translate the user's intended movements in three-dimensional space into computer-understandable instructions, such as cursor movements on a two-dimensional screen or discrete key presses.
[0003] However, current IMU-based human-computer interaction technology still faces core challenges in achieving comprehensive, efficient, and natural computer control, which limits its widespread application and user acceptance. First, in terms of cursor control, many existing technical solutions tend to adopt a relatively single sensitivity mapping mode or fixed control logic to cope with all operating scenarios. This design is often difficult to take into account diverse interaction needs in practical applications. For example, users may need a fast-responding relative mapping mode when navigating a large-scale interface, while preferring a stable and accurate absolute mapping mode or low sensitivity setting when performing fine positioning tasks (such as clicking on small targets or performing graphic editing). Existing solutions generally lack the flexibility to dynamically adjust the cursor control strategy according to task characteristics or the user's immediate intentions, which may lead to low operational efficiency in some scenarios or difficulty in achieving the required control accuracy in other scenarios.
[0004] Secondly, existing technologies have obvious deficiencies in achieving a natural and efficient integration of mouse functions (such as button clicks) and keyboard auxiliary functions (especially the coordinated operation of modifier keys such as Ctrl, Shift, and Alt). After using the IMU to position the cursor, users often find it difficult to perform accompanying operations such as mouse clicks smoothly and accurately through gestures, and these gestures themselves can easily interfere with the stability of the cursor. Furthermore, when it is necessary to perform combined operations that rely on modifier keys (such as Shift+drag selection, Ctrl+click multiple selection), most IMU control solutions fail to provide an intuitive and reliable mechanism for users to simultaneously manage cursor movement, mouse clicks, and the activation status of modifier keys. This leads to a sense of fragmented operation. Users may need to switch between different input modes, or it is difficult to complete such complex input tasks that are extremely common in modern operating systems and applications through the IMU, resulting in a large gap between the functional completeness and operational consistency of IMU-based control solutions and traditional keyboards and mice.
[0005] Therefore, this field urgently needs an IMU human-computer interaction method that can intelligently adapt to different operating scenarios and seamlessly integrate precise cursor control, reliable gesture key input, and collaborative control of keyboard modifier keys. Summary of the Invention
[0006] The embodiments of the present application provide a mouse and keyboard button control method based on an inertial measurement unit, aiming to provide an IMU human-computer interaction method that can intelligently adapt to different operating scenarios and seamlessly integrate precise cursor control, reliable gesture button input, and collaborative control of keyboard modifier keys.
[0007] To achieve the above objectives, an embodiment of the present application provides a method for controlling a mouse and keyboard buttons based on an inertial measurement unit, comprising:
[0008] Obtain the raw 3D motion data of the inertial measurement unit and perform multi-level preprocessing to generate a standard motion feature stream containing 3D attitude parameters and 3D angular velocity parameters;
[0009] Applying multi-scale dynamic time window segmentation to the standardized motion feature stream to extract candidate gesture segments, and matching the candidate gesture segments with templates of a pre-constructed hierarchical motion primitive library to obtain a preliminary primitive sequence;
[0010] Analyzing the preliminary primitive sequence according to the definition of the mode switching gesture in the pre-built hierarchical motion primitive library to identify the occurrence or end state of the mode switching gesture, and determining whether the current control mode is the mouse control mode or the keyboard control mode according to the recognition result of the mode switching gesture;
[0011] If the current control mode is the mouse control mode, selecting parameters from the standardized motion data stream as basic cursor control input according to the current human-computer interaction context information, determining and applying a mapping algorithm corresponding to the currently effective sensitivity mapping mode to process the basic cursor control input to generate a mouse cursor movement instruction;
[0012] and analyzing the portion of the preliminary primitive sequence that is not recognized as the mode switching gesture according to the definition of the mouse button gesture in the pre-built hierarchical motion primitive library, so as to recognize the mouse button gesture and parse it into a mouse button instruction;
[0013] If the current control mode is the keyboard control mode, analyzing the portion of the preliminary primitive sequence that is not recognized as the mode switching gesture according to the definition of the keyboard key gesture in the pre-built hierarchical motion primitive library to recognize the keyboard key gesture and parse it into a keyboard key instruction;
[0014] Combining the identified mouse cursor movement instructions, the mouse button instructions, and the keyboard button instructions to generate a mixed mouse and keyboard button event sequence;
[0015] The mixed mouse and keyboard key event sequence is converted into standard mouse input events and standard keyboard key input events of the target operating system, and the standard mouse input events and the standard keyboard key input events are injected into the event processing queue of the target operating system.
[0016] To achieve the above-mentioned objectives, an embodiment of the present application also proposes a mouse and keyboard button control device based on an inertial measurement unit, comprising a memory, a processor, and a mouse and keyboard button control program based on an inertial measurement unit stored in the memory and runnable on the processor. When the processor executes the mouse and keyboard button control program based on the inertial measurement unit, the mouse and keyboard button control method based on the inertial measurement unit as described in any one of the above items is implemented.
[0017] To achieve the above-mentioned objectives, an embodiment of the present application also proposes a computer-readable storage medium, on which a mouse and keyboard button control program based on an inertial measurement unit is stored. When the mouse and keyboard button control program based on an inertial measurement unit is executed by a processor, the mouse and keyboard button control method based on an inertial measurement unit as described in any one of the above items is implemented.
[0018] The mouse and keyboard button control method based on an inertial measurement unit proposed in this application introduces a specific mode switching gesture performed by the user and explicitly determines the current control mode as mouse control mode or keyboard control mode based on the recognition result of the gesture. This fundamentally solves the pain point that in a single or ambiguous interaction context, the system has difficulty accurately distinguishing whether the user's intention is to point the cursor, click a mouse button, or enter a keyboard character. This means that the system no longer relies solely on subtle differences in motion characteristics to passively infer user intention, but instead gives the user a clear, active action (mode switching gesture) to clearly define how the subsequent IMU data stream should be parsed and mapped. When the mode switching gesture is recognized to indicate entering mouse control mode, subsequent IMU data and gesture fragments will be mainly used to parse cursor movement and mouse button operations; conversely, if the mode switching gesture is recognized to indicate entering keyboard control mode, subsequent data and gestures will be mainly used to parse keyboard key operations. This design based on active mode declaration allows users to intuitively and reliably separate different types of operation intentions, significantly improving control accuracy and user confidence when mixed use of mouse and keyboard functions is required.
[0019] Secondly, the technical solution of the present application adopts a special parsing and instruction generation logic for the core functions of each determined control mode. Specifically, in the mouse control mode, the method not only selects appropriate parameters (such as posture or angular velocity) from the standardized motion data stream as basic cursor control input based on the current human-computer interaction context information, and applies an adaptive sensitivity mapping mode and algorithm to generate smooth and accurate mouse cursor movement instructions, but also independently analyzes and identifies mouse button gestures from the preliminary primitive sequence (after excluding mode switching gestures) and parses them into mouse button instructions. Similarly, in the keyboard control mode, the method focuses on identifying keyboard button gestures (such as specific postures or actions for modifier keys) from the preliminary primitive sequence (also after excluding mode switching gestures) and parsing them into corresponding keyboard button instructions, including confirmation of keyboard modifier key gestures, state tracking and trigger mode selection. This differentiated data parsing and command generation mechanism based on the current activity mode ensures that even if the user performs hand movements with similar physical forms, the system can correctly interpret them as the expected operations in that mode (cursor movement, mouse clicks, or keyboard keys) based on the activated mode state, thereby effectively avoiding confusion and mutual interference between different types of control commands, and achieving clear separation and independent but coordinated control of mouse and keyboard functions.
[0020] Furthermore, the technical solution of the present application realizes the effective coordination and integration of the identified mouse cursor movement instructions, mouse button instructions and keyboard button instructions through the subsequent instruction combination and event injection mechanism. These different types of instructions that may be generated in parallel or successively are synchronized and integrated in time, and the instruction priority and conflict resolution rules are applied. The fused instruction intention is interpreted according to the context-aware instruction mapping rule library, and finally a unified mouse and keyboard key mixed event sequence is generated. This design ensures that the user can perform combined operations naturally. For example, in keyboard control mode, activate a modifier key (such as shift), and then quickly switch to mouse control mode through a mode switching gesture to execute a mouse click. The system can correctly understand that this is a shift+click combination intention. The instruction combination and verification mechanism ensures that these instructions from different sources can be reasonably synchronized and interpreted, filtering out potential conflicts or redundancies. This overcomes the problem of multimodal operation fragmentation and difficulty in natural and efficient coordination, significantly improves the consistency of operation and the smoothness of the overall interaction, and enables users to complete more complex input tasks close to traditional keyboard and mouse operations through IMU.
[0021] In addition, the technical solution of this application performs multi-level collaborative preprocessing on the raw three-dimensional motion data of the inertial measurement unit to generate a high-quality standardized motion data stream, and uses multi-scale dynamic time window segmentation and hierarchical motion primitive library template matching to obtain a preliminary primitive sequence. This provides a stable and reliable data foundation and an efficient recognition framework for all subsequent mode judgment, gesture recognition, and mapping algorithms, ensuring the accuracy of overall control and the stability of response from the source. Finally, the confirmed mixed event sequence of mouse and keyboard key presses is converted into standard input events of the target operating system and injected into the event processing queue, ensuring the wide compatibility of this method and its immediate availability in various operating systems and application environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0023] Figure 1 This is a module structure diagram of an embodiment of a mouse and keyboard button control device based on an inertial measurement unit of the present invention;
[0024] Figure 2 2 is a flow chart of an embodiment of a method for controlling a mouse and keyboard buttons based on an inertial measurement unit according to the present invention.
[0025] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0026] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0027] To better understand the above technical solutions, exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0028] It should be noted that in the claims, any reference signs placed between brackets shall not be construed as limiting the claims. The presence of "comprising" in the text does not exclude the presence of components or steps not listed in the claims. The quantifier "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim that lists several means, several of these means may be embodied by the same hardware item. The use of "first", "second", and "third" etc. does not indicate any order and these words may be interpreted as names.
[0029] like Figure 1 As shown, Figure 1 It is a structural diagram of a server 1 (also called a mouse and keyboard button control device based on an inertial measurement unit) in a hardware operating environment involved in an embodiment of the present invention.
[0030] The server of the embodiment of the present invention is a device with display function such as "Internet of Things devices", smart air conditioners, smart lights, smart power supplies with networking functions, AR / VR devices with networking functions, smart speakers, self-driving cars, PCs, smart phones, tablet computers, e-book readers, portable computers, etc.
[0031] like Figure 1 As shown, the server 1 includes: a memory 11 , a processor 12 and a network interface 13 .
[0032] The memory 11 includes at least one type of readable storage medium, including a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the server 1, such as a hard disk of the server 1. In other embodiments, the memory 11 may also be an external storage device of the server 1, such as a plug-in hard disk equipped on the server 1, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0033] Furthermore, the memory 11 may include both an internal storage unit of the server 1 and an external storage device. The memory 11 can be used not only to store application software installed on the server 1 and various data, such as the code of the mouse and keyboard button control program 10 based on the inertial measurement unit, but also to temporarily store data that has been output or is about to be output.
[0034] In some embodiments, the processor 12 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, used to run the program code stored in the memory 11 or process data, such as executing the mouse and keyboard button control program 10 based on the inertial measurement unit.
[0035] The network interface 13 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the server 1 and other electronic devices.
[0036] The network may be the Internet, a cloud network, a wireless fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN), and / or a metropolitan area network (MAN). Various devices in the network environment may be configured to connect to the communication network according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of the following: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Light Fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocol, and / or Bluetooth communication protocol, or a combination thereof.
[0037] Optionally, the server may further include a user interface, which may include a display and an input unit such as a keyboard. The optional user interface may also include a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display, which may also be referred to as a display screen or display unit, is used to display information processed in the server 1 and to display a visual user interface.
[0038] Figure 1 Only the server 1 having components 11-13 and the mouse and keyboard button control program 10 based on the inertial measurement unit is shown. It can be understood by those skilled in the art that Figure 1 The structure shown does not constitute a limitation on the server 1 , and the server 1 may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0039] In this embodiment, the processor 12 may be configured to call the inertial measurement unit-based mouse and keyboard button control program stored in the memory 11 .
[0040] Based on the hardware architecture of the aforementioned inertial measurement unit (IMU)-based mouse and keyboard control device, an embodiment of the present invention's IMU-based mouse and keyboard control method is proposed. This IMU-based mouse and keyboard control method aims to provide an IMU human-computer interaction method that intelligently adapts to different operating scenarios and seamlessly integrates precise cursor control, reliable gesture-based key input, and coordinated control of keyboard modifier keys.
[0041] Reference Figure 2 , Figure 2 This is an embodiment of a method for controlling a mouse and keyboard buttons based on an inertial measurement unit of the present invention. The method for controlling a mouse and keyboard buttons based on an inertial measurement unit comprises the following steps:
[0042] S10. Obtain the original three-dimensional motion data of the inertial measurement unit and perform multi-level preprocessing to generate a standard motion feature stream containing three-dimensional posture parameters and three-dimensional angular velocity parameters. The purpose of this step is to collect the original motion signal from the inertial measurement unit (IMU) and eliminate the interference factors such as noise and drift that may exist therein through a series of systematic processing methods, and finally extract a standardized data stream that can accurately reflect the posture information of the device in three-dimensional space (such as the direction and rotation angle of the device) and the angular velocity information of the motion (that is, the speed of the device rotation). This standardized motion feature stream provides a reliable and consistent data basis for the subsequent accurate recognition of user gestures, judgment of user intentions, and generation of corresponding control instructions. The source of the original three-dimensional motion data can be various devices integrated with IMU, such as smart watches, smart bracelets, VR / AR controllers, etc. This method does not limit the specific hardware implementation of the IMU or the type of device it is in.
[0043] In some embodiments, step S10 may be implemented by the following steps:
[0044] S11, performing preliminary filtering on the original three-dimensional motion data by applying an adaptive low-pass filter, wherein a cutoff frequency of the adaptive low-pass filter is dynamically adjusted according to signal energy of the original three-dimensional motion data;
[0045] S12, performing sensor fusion processing on the raw 3D motion data after preliminary filtering using a Kalman filter to generate the 3D posture parameters and the 3D angular velocity parameters;
[0046] S13. Normalize and offset-correct the three-dimensional posture parameters and the three-dimensional angular velocity parameters to obtain the standardized motion feature flow.
[0047] In step S11, an adaptive low-pass filter is used to remove high-frequency noise from the raw 3D motion data, retaining the low-frequency, valid signals associated with the user's actual movements. The key to the adaptive low-pass filter is its ability to dynamically adjust its cutoff frequency based on the signal's energy, striking a balance between noise removal and signal detail preservation. Signal energy is typically determined by calculating the root mean square value of acceleration or angular velocity data over a period of time. This dynamic adjustment mechanism enables the filter to adapt to changes in the intensity of user movements, such as rapid waving versus slow adjustments.
[0048] In step S12, the Kalman filter comprehensively estimates the device's attitude and angular velocity by fusing data from the accelerometer, gyroscope, and magnetometer. The accelerometer provides information about the device's tilt relative to gravity, the gyroscope calculates angular displacement through integration, and the magnetometer is used to correct heading drift. The Kalman filter uses state prediction and measurement update mechanisms to optimize the integration of multi-sensor data and generate accurate three-dimensional attitude parameters (such as quaternions or Euler angles) and three-dimensional angular velocity parameters. This fusion process significantly reduces the effects of noise and drift in single sensor data, providing a reliable data foundation for subsequent processing.
[0049] In step S13, normalization and offset correction further optimize the generated attitude and angular velocity parameters. Normalization scales parameters of different dimensions to a uniform range (such as [0, 1] or [-1, 1]) to eliminate dimensional differences and facilitate subsequent algorithm processing. Offset correction is used to eliminate zero-point drift of the sensor when it is stationary. For example, a gyroscope may output a non-zero angular velocity when there is no motion. Through these processes, the generated standardized motion feature stream has a consistent format and high precision, which can support accurate mapping of multiple control modes.
[0050] In step S11, adaptive filtering ensures initial data cleanliness. In step S12, Kalman filtering fuses multi-sensor information to generate precise parameters. In step S13, normalization and correction ensure data consistency and reliability. This multi-stage processing mechanism works together to transform raw data into a standard feature stream suitable for complex interactive scenarios, laying a solid foundation for subsequent cursor movement, scroll wheel control, or gesture recognition.
[0051] For example, suppose a user uses a smartwatch equipped with an IMU to interact with a Windows laptop. The IMU outputs raw 3D motion data at a sampling rate of 100Hz, including angular velocity [30, 15, 8] degrees / second and acceleration [0.1, 0.2, -9.7]m / s. 2In step S11, the system calculates the energy of the angular velocity signal and obtains a root mean square value of approximately 19.4 degrees / second. The cutoff frequency of the adaptive low-pass filter is determined to be 6.94Hz according to the formula f_c=0.1*E+5, and the noise above this frequency is filtered out, and the preliminary filtered data is output. In step S12, the Kalman filter fuses the filtered acceleration, angular velocity and magnetometer data to generate three-dimensional attitude parameters (for example, the pitch angle represented by quaternion is 12°, the roll angle is 6°, and the yaw angle is 2°) and angular velocity parameters [30,15,8] degrees / second. In step S13, the posture parameters are normalized using the formula θ_norm = (θ - θ_min) / (θ_max - θ_min). For example, a pitch angle of 12° is normalized to (12 + 180) / (180 - (-180)) = 0.533. The angular velocity is corrected for the offset (e.g., an offset of [0.3, 0.2, 0.1] degrees / second becomes [29.7, 14.8, 7.9] degrees / second after correction). This ultimately generates a standard motion feature stream containing the normalized postures [0.533, 0.517, 0.506] and the corrected angular velocities [29.7, 14.8, 7.9] degrees / second. This feature stream can be used for subsequent cursor control or gesture recognition, for example, to support precise positioning in graphic design scenarios or enable fast navigation in gaming scenarios.
[0052] As can be understood, through the dynamic frequency adjustment of the adaptive low-pass filter, the system balances noise removal and signal preservation at varying motion intensities, significantly improving data quality. Furthermore, by fusing multi-sensor data through a Kalman filter, the system overcomes the limitations of a single sensor and generates precise attitude and angular velocity parameters. Furthermore, through normalization and offset correction, the system ensures data consistency and reliability, providing a high-quality input foundation for subsequent mouse and keyboard control, thereby improving the accuracy and smoothness of interactions.
[0053] S20. Apply multi-scale dynamic time window segmentation to the standardized motion feature stream to extract candidate gesture segments, and match the candidate gesture segments with the template of the pre-built hierarchical motion primitive library to obtain a preliminary primitive sequence. This process aims to identify segments that may correspond to mouse or keyboard control gestures from the continuous standardized motion feature stream, and convert them into structured primitive sequences through template matching, providing a basis for subsequent mode switching, mouse button or keyboard button recognition. Through multi-scale time window segmentation and template matching, the system can capture user actions of different time spans and complexities, such as rapid jitters, continuous posture maintenance or sequential actions, laying the foundation for achieving mouse cursor movement, key operations or keyboard command generation.
[0054] In some embodiments, step S20 may be implemented by the following steps:
[0055] S21, applying a three-level parallel sliding window process comprising a micro-window, a medium-window, and a macro-window to the standardized motion feature stream, and dynamically adjusting the window size and overlap ratio of each level of windows according to the energy change rate and complexity index of the standardized motion feature stream to extract the candidate gesture segments;
[0056] S22. Extract time domain statistical features, frequency domain distribution characteristics and posture trajectory features from the motion data in the candidate gesture segment, and perform pattern matching between the extracted features and the primitive templates of the corresponding level in the pre-built hierarchical motion primitive library using a dynamic time warping algorithm or a support vector machine model to obtain the preliminary primitive sequence.
[0057] In step S21, a three-level parallel sliding window process captures potential gesture segments in the standardized motion feature stream by setting windows of different time spans (micro, medium, and macro). Micro windows (e.g., 0.1 seconds) are suitable for fast actions such as mouse clicks or single keyboard keystrokes; medium windows (e.g., 0.3 seconds) are suitable for actions of medium duration, such as double-clicks or short keyboard input sequences; and macro windows (e.g., 0.5 seconds) are used to identify longer actions, such as long presses or sustained keyboard modifier keys. The window size and overlap are dynamically adjusted based on the energy change rate (e.g., the root mean square value of the angular velocity change) and complexity metrics (e.g., signal entropy) of the feature stream. For example, when a sharp increase in the energy change rate is detected, the system may shorten the micro window length and increase its overlap to more accurately capture and define the start and end points of this rapidly changing action. Conversely, if the motion stream exhibits low complexity and gradual energy changes, the macro window size may be appropriately increased to fully encompass a longer, continuous action. Through this adaptive segmentation strategy, the system can effectively extract candidate gesture segments with a high probability of gesture from continuous data.
[0058] In step S22, time domain statistical features (such as mean, variance, peak), frequency domain distribution characteristics (such as the main frequency component after Fourier transform) and posture trajectory features (such as the continuous change trajectory of the posture angle) are extracted from the candidate gesture segment to comprehensively describe the motion pattern of the segment. The pre-built hierarchical motion primitive library contains templates at different levels, such as micro-motion primitives (such as single jitter) for fast movements, and combination primitives (such as continuous jitter or posture sequence) for complex gestures, such as keyboard modifier key combinations. The dynamic time warping (DTW) algorithm allows nonlinear time alignment by calculating the time series similarity between the segment features and the template, and adapts to individual differences in user movement speed; the support vector machine (SVM) model uses a classifier to identify the degree of match between the features and the template. The two are combined to generate a preliminary primitive sequence, which contains structured identifiers that may correspond to mouse or keyboard gestures, such as "fast jitter" or "gesture sequence".
[0059] Step S21 captures motion features at different time scales through multi-scale window segmentation, and step S22 converts the segments into structured sequences through multi-dimensional feature extraction and template matching. This combination enables the system to extract meaningful gestural segments from continuous motion data and initially identify their patterns, providing reliable input for subsequent mouse and keyboard control command generation.
[0060] For example, assuming a user is interacting with a Windows computer using an IMU-equipped smartwatch, the normalized motion feature stream contains normalized gesture parameters [0.528, 0.514, 0.503] and calibrated angular velocity [24.8, 11.9, 4.9] degrees / second. In step S21, the system detects two angular velocity peaks within 0.3 seconds ([50, 25, 10] degrees / second and [48, 22, 8] degrees / second), calculates the energy change rate (RMS value approximately 36 degrees / second) and complexity index (high signal entropy), adjusts the micro-window to 0.1 seconds (overlap ratio 0.5) and the medium window to 0.3 seconds (overlap ratio 0.3), and extracts candidate gesture segments containing these two peaks. In step S22, the system extracts the segment's time domain features (peak 50 degrees / second, variance 8), frequency domain features (dominant frequency 4 Hz), and gesture trajectory features (pitch angle change trajectory). Using the DTW algorithm, the segment features were matched with a double-click template in the primitive library (defined as two rapid angular velocity peaks separated by less than 0.5 seconds), achieving a similarity score of 0.94. The SVM model confirmed the match and generated a preliminary primitive sequence, labeled the "double-click primitive." This sequence can be used to subsequently identify double-clicks or keyboard combinations (such as Ctrl+C).
[0061] It can be understood that through multi-scale dynamic time window segmentation, the system flexibly captures gesture features at different time scales and adapts to the diversity of rapid or continuous movements, thereby improving the accuracy and coverage of extraction. In addition, by extracting time domain, frequency domain, and trajectory features and combining them with DTW or SVM for template matching, the system tolerates individual differences in user movement rhythms, accurately identifies complex gesture patterns, and generates reliable primitive sequences. At the same time, this combination of multi-dimensional feature extraction and hierarchical template matching significantly improves the robustness of gesture recognition, reduces the misrecognition rate, provides a high-quality input foundation for subsequent mouse and keyboard command generation, and optimizes the accuracy and naturalness of interaction.
[0062] S30. According to the definition of the mode switching gesture in the pre-built hierarchical motion primitive library, the preliminary primitive sequence is analyzed to identify the occurrence or end state of the mode switching gesture, and the current control mode is determined to be the mouse control mode or the keyboard control mode according to the recognition result of the mode switching gesture. This process aims to dynamically switch the control mode of the device by analyzing the user's gesture movements to adapt to different interaction needs, such as realizing cursor movement and key operations in mouse control mode, or triggering modifier keys (such as Ctrl, Shift) or function keys in keyboard control mode. The preliminary primitive sequence is a structured identifier generated by extracting and matching from the standardized motion feature stream, which contains motion patterns that may correspond to gestures. The hierarchical motion primitive library defines mode switching gesture templates, such as making a fist or a specific posture sequence, which are used to distinguish the triggering of mouse and keyboard control modes. By recognizing these gestures, the system can flexibly switch the control mode according to the user's intention, ensuring that the interaction method is highly matched with the current task scenario.
[0063] In some embodiments, based on the definition of the mode switching gesture in the pre-constructed hierarchical motion primitive library, the preliminary primitive sequence is analyzed to identify the occurrence or end state of the mode switching gesture, including: matching the preliminary primitive sequence with the mode switching gesture template predefined in the pre-constructed hierarchical motion primitive library using a dynamic time warping algorithm, and combining the intention prediction probability based on the initial part data of the preliminary primitive sequence to determine the occurrence or end of the mode switching gesture.
[0064] Here, a composite strategy combining deterministic matching and probabilistic prediction is adopted for the recognition of mode switching gestures to improve the recognition accuracy and response speed. First, the system matches the currently observed preliminary primitive sequence (or a fragment within a sliding time window) with a variety of predefined mode switching gesture templates stored in a pre-built hierarchical motion primitive library. This matching process preferably uses the dynamic time warping algorithm (DTW). The advantage of the DTW algorithm is that it can effectively compare two sequences that may have nonlinear stretching or compression on the time axis (that is, the speed at which the user performs the gesture may be different). As long as their overall shape and primitive order are similar, DTW can calculate a higher degree of matching. Each predefined mode switching gesture template is itself a specific primitive sequence, representing a specific user action such as switching from mouse mode to keyboard mode, or from keyboard mode back to mouse mode.
[0065] Secondly, to further enhance recognition robustness and potentially detect the user's switching intent in advance, the system also incorporates the intent prediction probability based on the initial portion of the preliminary primitive sequence. This means that before a gesture is fully executed or DTW matching is still in progress, the system analyzes the beginning of the current preliminary primitive sequence (e.g., the first few primitives or a short duration of primitive data). Based on these early signals, a pre-trained intent prediction model (e.g., a model based on statistical learning, machine learning, or specific heuristic rules) estimates the likelihood or probability that the user will next perform a specific mode switching gesture. This prediction probability can be fused with the DTW algorithm's matching results (e.g., matching score or cost). For example, a higher intent prediction probability can reduce the requirement for the DTW matching score or prioritize the mode switching gesture corresponding to the higher prediction probability among multiple possible DTW matches. Conversely, if the DTW matching results themselves are ambiguous, a lower intent prediction probability may lead the system to dismiss the mode switching as occurring. In this way, by comprehensively considering the complete sequence pattern matching and the probabilistic clues of early intentions, the system finally makes a judgment on the occurrence (e.g., a switch gesture is successfully recognized) or the end (e.g., the conditions of a previously activated switch gesture are no longer met, or a reverse switch gesture is recognized) of the mode switch gesture.
[0066] Through this recognition method that combines template matching of the dynamic time warping algorithm with the intention prediction probability based on early sequence data, step S30 can more intelligently and reliably analyze the user's mode switching intention, thereby accurately performing state transitions between mouse control mode and keyboard control mode.
[0067] For example, suppose the user is currently in mouse control mode and wishes to switch to keyboard control mode. The system predefines a mode-switch gesture, such as "quickly flip the wrist wearing the IMU upward twice." In the hierarchical motion primitive library, this gesture is represented as a specific mode-switch gesture template, which might consist of a "quick flip up" primitive repeated twice, separated by a short pause. When the user performs this action, the system generates a preliminary primitive sequence with a fragment similar to [...other primitives, quick flip up primitive_1, short pause primitive, quick flip up primitive_2, ...].
[0068] When recognizing a mode switch gesture, the system may first analyze the initial data after the appearance of the quick flip primitive_1. If the intent prediction model, based on historical data or learning results, determines that after a quick flip, the user has a high probability (e.g., 60%) of performing a second quick flip to complete the mode switch, this predicted probability is recorded. Subsequently, when the quick flip primitive_2 also appears, the system uses a dynamic time warping algorithm to match the primitive subsequence containing these two flips and the pause in between with the "two quick wrist flips upward" mode switch gesture template in the library. If the DTW algorithm calculates a high similarity (e.g., the matching cost is below a certain threshold), combined with the previously predicted increased intent probability (e.g., after the second flip, the prediction model's probability of completing this particular switch gesture may increase to 95%), the system will comprehensively determine that the mode switch gesture has successfully occurred. Then, according to step S30, the current control mode is changed from mouse control mode to keyboard control mode. Conversely, if you want to switch back from keyboard mode to mouse mode, you may need to perform another different predefined mode switching gesture, and the recognition process is similar.
[0069] As can be understood, by matching the observed preliminary primitive sequence with the predefined mode switching gesture templates in the pre-built hierarchical motion primitive library using a dynamic time warping algorithm, it is possible to effectively identify user-specific gestures that conform to the preset mode in terms of overall form and primitive order. This algorithm maintains high matching accuracy and robustness even when users perform these gestures with varying speeds or amplitudes, which can be varied from person to person or even instantaneously. Furthermore, by further combining this with the predicted probability of user intent derived from the initial data analysis of the preliminary primitive sequence, it assists in determining the final state or end state of the mode switching gesture. This enables the system to not only confirm the user's complete switch gesture execution but also to predict and track the user's potential switch intentions to a certain extent in the early stages of gesture execution. This not only improves the sensitivity of the mode switching command response and the smoothness of the interaction, but also provides the system with more intelligent and comprehensive decision-making information support when faced with incomplete, accidentally interrupted, or unclear switch gestures. This recognition strategy, which combines deterministic sequential pattern matching technology with a probabilistic early intent prediction mechanism, jointly improves the overall performance of mode switching gesture recognition, making the switching process between mouse control mode and keyboard control mode more reliable, natural, and better able to adapt to the diversity of user behavior.
[0070] S40. If the current control mode is the mouse control mode, parameters are selected from the standardized motion data stream as basic cursor control inputs based on the current human-computer interaction context information, and the mapping algorithm corresponding to the currently effective sensitivity mapping mode is determined and applied to process the basic cursor control inputs to generate mouse cursor movement instructions. This step specifically handles the cursor movement control logic when the system is in mouse control mode (based on the judgment of step S30). Its core idea is to implement a cursor control mechanism that is highly context-adaptive and user-dynamic behavior-adaptive. It first evaluates the user's current interaction environment, and then selects the most appropriate motion parameters from the standardized motion data stream based on this environment as the basis for driving the cursor. Next, it dynamically switches between different sensitivity mapping modes (for example, a mode suitable for precise pointing and a mode suitable for fast movement) based on the speed of the user's current action. Finally, combined with the user's existing mouse speed preference in the operating system, the selected mapping mode and algorithm are applied to process the basic cursor control input to generate smooth and user-expected mouse cursor movement instructions.
[0071] In some embodiments, step S40 can be implemented by following the steps below:
[0072] S41. Determine, through an application program interface or a user interface analysis tool provided by the operating system, the type of the currently active application or the type of the user interface element under the cursor as the current human-computer interaction context information;
[0073] S42. Selecting, according to a preset context mapping table, a specific axial combination of the three-dimensional gesture parameters or a specific axial combination of the three-dimensional angular velocity parameters corresponding to the type of the currently active application or the type of the user interface element under the cursor as a basic cursor control input;
[0074] S43, detecting the movement amplitude or movement speed related to the basic cursor control input in the standard movement data stream,
[0075] S44: When the movement amplitude or the movement speed exceeds a preset first threshold, determining the preset relative mapping mode as the currently effective sensitivity mapping mode,
[0076] S45. When the movement amplitude or the movement speed is lower than a preset second threshold, determining the preset absolute mapping mode as the currently effective sensitivity mapping mode;
[0077] S46. Read the physical mouse sensitivity setting of the operating system, and calculate the basic mouse sensitivity adjustment parameter according to the physical mouse sensitivity setting;
[0078] S47 uses the basic mouse sensitivity adjustment parameter as the gain adjustment factor of the mapping algorithm corresponding to the currently effective sensitivity mapping mode, processes the basic cursor control input, and when the currently effective sensitivity mapping mode changes from the sensitivity mapping mode of the previous control cycle, performs a smooth transition processing based on weighted average based on the current output of the preset mapping algorithm and the final output state data of the previous control cycle to generate the mouse cursor movement instruction.
[0079] In step S41, the system uses the application program interface (API) provided by the operating system (e.g., GetForegroundWindow in Windows) or user interface analysis tools (e.g., UI Automation) to identify the type of the currently active application (e.g., text editor, graphic design software) or the type of interface element under the cursor (e.g., button, text box). This contextual information helps the system understand the user's operation objectives, such as precise cursor positioning in graphic design and fast movement in gaming.
[0080] In step S42, a pre-set contextual mapping table defines the correspondence between different applications or interface elements and axial combinations of motion parameters. For example, in a text editor, the pitch and roll angle combinations of the 3D pose parameters are selected for precise cursor movement; in a gaming application, the X-axis and Y-axis components of the angular velocity parameters are selected for fast navigation. This context-adaptive selection mechanism ensures that control inputs match task requirements.
[0081] In step S43 , the system detects the movement amplitude (such as the root mean square value of the angular velocity) or movement speed of the basic cursor control input to quantify the intensity of the user action and provide a basis for subsequent mapping mode selection.
[0082] Steps S44 and S45 together implement a function that automatically switches between different sensitivity mapping modes based on the user's immediate motion intent. Step S44 stipulates that if the detected motion amplitude or speed exceeds a preset first threshold (this means that the user may be attempting a quick, large-scale cursor sweep), the system will determine to use the preset relative mapping mode as the currently effective sensitivity mapping mode. In this mode, the IMU device's motion rate or attitude change rate is typically mapped to the movement speed or relative displacement of the cursor on the screen. Conversely, step S45 stipulates that if the detected motion amplitude or speed is lower than a preset second threshold (this usually indicates that the user is making small-scale, slow, fine adjustments or aiming), the system will determine to use the preset absolute mapping mode as the currently effective sensitivity mapping mode. In this mode, the absolute attitude of the IMU device (for example, the tilt angle) is typically directly mapped to an absolute coordinate position on the screen. The setting of these two thresholds helps the system intelligently determine whether the user intends to "point and shoot" or "quickly swing."
[0083] In step S46, the system reads the physical mouse sensitivity setting of the operating system (for example, in the Windows system, the system obtains the mouse sensitivity value set by the user by calling the SystemParametersInfo function (using the SPI_GETMOUSESPEED parameter), which is usually in the range of 1 to 20, and the default value is 10, indicating medium sensitivity), and calculates the basic mouse sensitivity adjustment parameters through the conversion rule (such as the linear formula S_IMU=0.04*S_OS+0.2) to ensure that the IMU control is consistent with the user's traditional mouse habits.
[0084] In step S47, the basic mouse sensitivity adjustment parameter is used as a gain factor to adjust the mapping algorithm output ratio, for example, scaling angular velocity to pixel displacement in relative mapping mode. When the mapping mode is switched (e.g., from relative to absolute), the system uses a weighted average (e.g., C = 0.7 * C_prev + 0.3 * C_current) to smoothly transition between the current output and the previous cycle output to avoid cursor jumps.
[0085] When the system is in mouse control mode, steps S41 through S47 work together to form a highly adaptive and context-aware mouse cursor movement command generation process. This intelligently selects the optimal motion input source (S42) based on the user's current application scenario (S41), dynamically switches the most appropriate mapping logic (S44, S45) based on the user's immediate action speed (S43), and respects the user's existing system-level mouse speed preference (S46). On top of all this, through sophisticated algorithmic processing and necessary smooth transitions (S47), high-quality cursor control commands are ultimately output.
[0086] For example, suppose a user uses a smartwatch equipped with an IMU to operate on a Windows computer, which is currently in mouse control mode. The normalized motion feature stream contains the normalized posture [0.528, 0.514, 0.503] and angular velocity [24.8, 11.9, 4.9] degrees / second. In step S41, the system detects that the active application is "Notepad.exe" through GetForegroundWindow, indicating a text editing scenario. In step S42, according to the context mapping table, the pitch angle and roll angle of the posture parameters ([0.528, 0.514]) are selected as the basic cursor control input. In step S43, the system calculates the angular velocity amplitude as sqrt(24.8 2 +11.9 2 +4.9 2)≈27.6 degrees / second. In steps S44 and S45, the amplitude exceeds the first threshold of 20 degrees / second, and the relative mapping mode is selected. In step S46, the system reads the mouse sensitivity setting to 12 and calculates the basic sensitivity parameter S_IMU=0.04*12+0.2=0.68. In step S47, the system maps the angular velocity [24.8,11.9] degrees / second to the cursor displacement [0.68*24.8≈16.9,0.68*11.9≈8.1] pixels / second. If switched to absolute mode, it smoothly transitions to the absolute coordinates [475.2,480.6] through weighted averaging (weight 0.7) to generate cursor movement instructions, which are suitable for text cursor positioning.
[0087] It can be understood that by determining the user's currently active application type or the type of user interface element under the cursor to obtain the current human-computer interaction context information, subsequent cursor control strategies (such as the selection of motion parameters as input) can be made more targeted and intelligent, thereby better adapting to the specific operational requirements and user habits in different application scenarios. Moreover, according to a preset context mapping table, the parameter combination that matches the specific context information is selected as the basic cursor control input. This can match the most intuitive and efficient raw motion signal source for different types of interactive tasks (for example, precise drawing vs. quick browsing), significantly improving the naturalness of control and the efficiency of operation. Furthermore, by real-time detection of the motion amplitude or speed of the basic cursor control input in the standard motion data stream, and based on this, dynamically switching and determining the currently effective sensitivity mapping mode between the preset relative mapping mode suitable for rapid large-scale movement and the absolute mapping mode suitable for precise small-scale positioning, the system can intelligently adapt to the dynamic changes in user behavior under different operational intentions, significantly enhancing the flexibility, responsiveness and ability of cursor control to meet diverse task requirements. At the same time, by reading the physical mouse sensitivity settings configured by the user in the target operating system and calculating the basic mouse sensitivity adjustment parameters suitable for IMU control based on this, the parameters are uniformly applied to the currently effective mapping algorithm as gain adjustment factors. This ensures that the overall perception of movement speed of IMU-based mouse control is close to the user's existing traditional mouse usage habits, effectively reducing the user's learning threshold and adaptation cost. In addition, at the critical moment when the currently effective sensitivity mapping mode changes from the mode of the previous control cycle (i.e., mode switching), a smooth transition based on weighted average is performed based on the output states of the new and old mapping algorithms and the historical final output data. This can effectively eliminate or significantly reduce the sudden visual jumps of the mouse cursor, the sense of loss of control in operation, or the discontinuous interruption of response caused by sudden changes in mapping logic or sensitivity characteristics, thereby ensuring the continuity and stability of cursor control and the overall smoothness of the user experience during dynamic tasks and mode switching.
[0088] S50, and according to the definition of mouse button gestures in the pre-built hierarchical motion primitive library, analyze the part of the preliminary primitive sequence that is not recognized as the mode switching gesture to identify the mouse button gesture and parse it into a mouse button instruction. This step is specifically responsible for processing various gestures performed by the user through the IMU device in the mouse control mode to simulate mouse buttons (such as left click, right click, double click, long press, etc.). Its core task is to carefully analyze and match the remaining preliminary primitive sequence fragments after the gestures for mode switching have been excluded by step S30 to accurately identify the specific key action that the user intends to perform and convert it into a mouse button instruction that can be understood and executed by the computer system.
[0089] In some embodiments, step S50 may be implemented by the following steps:
[0090] S51. Applying a preset sequence matching rule to a portion of the preliminary primitive sequence that is not recognized as the mode switching gesture, the portion is matched with mouse button gesture templates defined in the pre-built hierarchical motion primitive library, where the mouse button gesture templates include a single-click gesture template, a double-click gesture template, and a long-press gesture template.
[0091] S52, performing duration verification, motion amplitude verification, and confidence threshold filtering based on current human-computer interaction context information on the successfully matched candidate mouse button gestures;
[0092] S53: Parse the verified candidate mouse button gesture into a corresponding left-click instruction, right-click instruction, double-click instruction, or a button press and release instruction sequence to generate the mouse button instruction.
[0093] In step S51, the system uses sequence matching rules (such as dynamic time warping (DTW)) to compare the non-mode switching portion of the preliminary primitive sequence with the mouse button gesture template. A single-click gesture template might be defined as a single rapid angular velocity peak (e.g., lasting 0.1 seconds); a double-click template as two peaks (intervals less than 0.5 seconds); and a long-press template as a sustained stability of gesture parameters (e.g., a change of less than 0.01 within 0.5 seconds). The DTW algorithm allows for nonlinear alignment on the timeline by calculating the similarity between the sequence and the template, adapting to differences in user motion rhythms.
[0094] In step S52, the candidate key gestures undergo multiple verifications to ensure reliability. Duration verification checks whether the gesture duration meets template requirements (e.g., a single click is less than 0.2 seconds). Motion amplitude verification confirms that the angular velocity or posture change exceeds a threshold (e.g., 50 degrees / second) to distinguish intentional motion from noise. The confidence threshold filter adjusts the verification criteria based on the human-computer interaction context (e.g., the application is a file manager). For example, a higher threshold (e.g., 0.9) is set in high-precision scenarios (e.g., graphic editing) to reduce false positives.
[0095] In step S53, the candidate gestures that pass the verification are parsed into standard mouse button commands. For example, a single-click gesture may generate MOUSEEVENTF_LEFTDOWN / UP (left-click) or MOUSEEVENTF_RIGHTDOWN / UP (right-click), a double-click generates two consecutive left-clicks, and a long-press generates a press-and-release sequence (such as a drag operation). This process ensures that the commands are compatible with the operating system's input protocol.
[0096] When the system is in mouse control mode, through the preliminary gesture recognition based on template matching in step S51, combined with the rigorous multi-dimensional and context-adaptive verification and screening in step S52, and the precise analysis of specific standard key events in step S53, step S50 can accurately and reliably extract the intentions of various mouse button operations from the user's continuous body movements, and convert them into actual computer control commands, thereby realizing rich key interaction functions without the need for a physical mouse.
[0097] For example, suppose a user is using an IMU-equipped smartwatch to navigate the Windows file manager. The initial primitive sequence includes the "rapid angular velocity peak" and "stable posture" flags. In step S51, the system detects two angular velocity peaks ([60, 30, 15] degrees / second and [55, 25, 12] degrees / second, separated by 0.3 seconds). DTW matches the double-click template (two peaks, separated by <0.5 seconds) with a similarity score of 0.95. In step S52, duration verification confirms that the 0.3-second interval meets the requirement, and amplitude verification confirms that the peak exceeds 50 degrees / second. Context analysis (the application is "explorer.exe") sets a confidence threshold of 0.9, and the matching score satisfies the requirements. In step S53, the double-click gesture is interpreted as two MOUSEEVENTF_LEFTDOWN / UP command sequences, triggering the file open operation.
[0098] As you can see, through sequence matching rules and the DTW algorithm, the system accurately identifies mouse button gestures and adapts to variations in movement. Furthermore, a multi-verification mechanism significantly reduces false positives by filtering by duration, amplitude, and contextual confidence. Furthermore, interpreting gestures as standard commands ensures compatibility with the operating system and improves the accuracy and fluidity of IMU-based mouse interactions in complex scenarios.
[0099] S60. If the current control mode is the keyboard control mode, then according to the definition of the keyboard key gesture in the pre-built hierarchical motion primitive library, analyze the part of the preliminary primitive sequence that is not recognized as the mode switching gesture to identify the keyboard key gesture and parse it into a keyboard key instruction. The purpose of this process is to interpret the user's body movements when the system confirms that it is in the keyboard control mode through the judgment of step S30, so as to realize the function of simulating keyboard input through the IMU device. Its core task is to analyze those valid fragments in the preliminary primitive sequence generated in the previous step (such as step S20) that are not used for mode switching, and refer to the specific gesture patterns defined for various keyboard keys (such as letters, numbers, symbols, function keys, and even modifier keys such as Ctrl, Shift, Alt, etc.) in the pre-built hierarchical motion primitive library for matching and identification. Once a keyboard key gesture is successfully identified, the system will parse it into the corresponding keyboard key instruction, allowing the user to perform keyboard input or control operations in a non-contact somatosensory manner.
[0100] In some embodiments, step S60 may be implemented by the following steps:
[0101] S61, matching the portion of the preliminary primitive sequence that is not recognized as the mode switching gesture with the keyboard gesture template in the pre-built hierarchical motion primitive library using a dynamic time warping algorithm, and confirming the matching result using a disambiguation mechanism that includes adjusting the recognition threshold or priority according to the current application scenario, to obtain a confirmed keyboard gesture recognition result;
[0102] S62. Apply a hierarchical state machine to track the state of the confirmed keyboard key gesture recognition result, distinguish between continuous keyboard key gestures and sequential keyboard key gestures, and select one of the instantaneous trigger mode, continuous trigger mode or locked trigger mode to parse it into the keyboard key instruction.
[0103] In step S61, the dynamic time warping (DTW) algorithm is used to match the non-mode switching part of the preliminary primitive sequence with the keyboard key gesture template. The template may include a single action (such as a rapid angular velocity peak corresponding to the Ctrl key trigger), a sequence action (such as two shakes corresponding to the Enter key) or a continuous posture (such as a 0.5-second stable posture corresponding to the Shift key hold). DTW allows nonlinear alignment on the time axis by calculating the similarity between the sequence and the template, and adapts to the differences in user action rhythm. To improve recognition accuracy, the system applies a disambiguation mechanism to adjust the recognition threshold or priority according to the current application scenario (such as a text editor or command line interface). For example, in a text editing scenario, the threshold of the Ctrl key template may be lowered to prioritize the recognition of shortcut keys, while in the command line, the priority of the Enter key template is increased. This scene-adapted disambiguation mechanism ensures that the gesture recognition results are consistent with the user's intentions. Figure 1 To.
[0104] In step S62, a hierarchical state machine tracks the states of confirmed keyboard key gestures, distinguishing between continuous gestures (e.g., a continuous press of the Shift key) and sequential gestures (e.g., a Ctrl+C sequence). The state machine selects a trigger mode based on the gesture's temporal characteristics and context (e.g., duration or sequence): a momentary trigger mode generates a single key press event (e.g., the Enter key); a continuous trigger mode generates a press-and-hold event (e.g., the Shift key); and a locked trigger mode generates a locked state event (e.g., CapsLock). Through state tracking and trigger mode selection, the system generates keyboard key commands that conform to operating system protocols.
[0105] Through DTW template matching combined with scene-adaptive disambiguation in step S61, preliminary accurate recognition and confirmation of keyboard key gestures are achieved; then, through the application of a hierarchical state machine in step S62, detailed state tracking and multi-mode trigger analysis are performed. Step S60 can, in keyboard control mode, convert the user's complex and diverse body movements into rich keyboard key instructions that conform to the actual keyboard usage logic.
[0106] For example, suppose a user is using an IMU-equipped smartwatch in a Windows text editor, currently in keyboard control mode. The initial primitive sequence includes the "rapid angular velocity peak" and "stability" flags. In step S61, the system detects a single angular velocity peak ([50, 25, 10] degrees / second, lasting 0.1 seconds). DTW matches it to the Ctrl key template (single peak, lasting <0.2 seconds) in the primitive library, achieving a similarity score of 0.92. The disambiguation mechanism lowers the Ctrl key threshold to 0.9 based on the text editor scenario, confirming that the gesture was triggered by the Ctrl key. In step S62, the hierarchical state machine tracks the gesture state, identifies it as a momentary trigger mode, and generates a KEYEVENTF_KEYDOWN / KEYUP command (virtual key code VK_CONTROL), triggering a Ctrl key event and supporting shortcut key operations (such as Ctrl+C to copy). If the sequence includes 0.5 seconds of stability (change <0.01), the Shift key template is matched, and the hierarchical state machine selects the continuous trigger mode, generating a continuous press command.
[0107] As can be seen, the system accurately identifies keyboard gestures through DTW matching and scenario-adaptive disambiguation mechanisms, adapting to different gestures and application scenarios. Furthermore, the hierarchical state machine distinguishes between continuous and sequential gestures through state tracking and trigger mode selection, significantly reducing the false recognition rate. Furthermore, the generation of standard keyboard commands ensures compatibility with the operating system, improving the accuracy and smoothness of IMU-based keyboard interactions in complex scenarios.
[0108] S70. Combine the identified mouse cursor movement instructions, the mouse button instructions and the keyboard button instructions to generate a mixed mouse and keyboard button event sequence. This process aims to integrate the various instructions generated in the previous steps to form a logically consistent and time-coordinated event sequence to ensure that the cursor movement, mouse button (such as single-click, double-click) and keyboard button (such as Ctrl, Enter) operations are seamlessly connected to reflect the user's complete intention in complex interactive scenarios. The mouse cursor movement instruction controls the displacement of the screen pointer, the mouse button instruction implements click or drag, and the keyboard button instruction triggers the modifier key or function key. Through time alignment, conflict resolution and priority management, the system generates a mixed event sequence, supports compound operations (such as Ctrl+click, Shift+drag), and improves the smoothness and reliability of interactions based on the inertial measurement unit (IMU).
[0109] In some embodiments, combining the recognized mouse cursor movement instruction, the mouse button instruction, and the keyboard button instruction to generate a mixed mouse and keyboard button event sequence includes:
[0110] S71, performing timestamp alignment and status information integration on the identified mouse cursor movement instruction, the mouse button instruction, and the keyboard button instruction;
[0111] S72. Apply predefined instruction priority rules and conflict resolution logic to process the mouse cursor movement instruction, the mouse button instruction and the keyboard button instruction to generate a mixed mouse and keyboard button event sequence, wherein the conflict resolution logic includes suppressing the mouse cursor displacement instruction caused by gesture action within the trigger time window of the mouse button instruction when the confidence of the mouse button instruction is higher than a preset button confirmation threshold, and increasing the button confirmation threshold for confirming the mouse button instruction when the mouse cursor movement instruction represents high-speed continuous movement.
[0112] In step S71, timestamp alignment ensures that various instructions are arranged in the correct timing. Each instruction (such as cursor movement [10,5] pixels, left click, Ctrl key press) is accompanied by a timestamp reflecting the moment of generation. The system aligns instructions by a unified time window (such as ±20 milliseconds) to eliminate timing deviations caused by sensor sampling or processing delays. State information integration records the contextual state of the instruction (such as the current cursor position and key status) to ensure consistent sequence logic, such as integrating the press, move, and release states in a drag operation.
[0113] In step S72, predefined instruction priority rules and conflict resolution logic coordinate multimodal instructions. For example, the priority rules may stipulate that keyboard modifier keys (such as Ctrl) take precedence over mouse movements to support shortcut key operations. The conflict resolution logic includes: when the confidence of a mouse button instruction (such as a single click) is higher than a threshold (such as 0.9), the cursor displacement instruction within its trigger time window (±50 milliseconds) is suppressed to avoid the key action from falsely triggering the movement; when the cursor movement instruction represents high-speed continuous movement (such as angular velocity >30 degrees / second), the key instruction confirmation threshold is increased (such as from 0.9 to 0.95) to reduce false triggering during high-speed movement. These mechanisms generate a logically clear hybrid event sequence.
[0114] Through step S71, various types of instructions are precisely time-aligned and key status information is integrated to ensure the correct association of different modal inputs in terms of timing and logic. Subsequently, step S72 uses preset instruction priority determination and conflict resolution logic including specific suppression and dynamic threshold adjustment mechanisms to deeply process and optimize the integrated instruction stream. The collaborative work of these two sub-steps enables step S70 to intelligently and conflict-free merge the outputs from the three independent modules of mouse cursor control, mouse button operation, and keyboard button control into a unified, highly reliable mouse and keyboard button mixed event sequence, providing a solid foundation for realizing complex cross-modal human-computer interaction (for example, holding down keyboard modifier keys while dragging the mouse to select).
[0115] For example, suppose a user uses an IMU-equipped smartwatch in a Windows text editor. The generated commands include a cursor movement ([16.9, 8.1] pixels / second, t=1000ms), a left-click (confidence 0.95, t=1005ms), and a Ctrl key press (t=1010ms). In step S71, the system aligns the timestamps (1000ms, 1005ms, 1010ms), integrates the states (e.g., cursor position [475.2, 480.6], key state), and confirms the Ctrl+click sequence. In step S72, the priority rule assigns high priority to the Ctrl key. Because the single-click confidence is greater than 0.9 (0.95), the conflict resolution logic suppresses movement commands within 1005ms ± 50ms, ensuring that the Ctrl+click takes precedence. This generates a mixed event sequence: Ctrl+click → left-click, triggering a copy operation (e.g., Ctrl+C).
[0116] As you can see, through timestamp alignment and state integration, the system ensures the timing and logical consistency of multimodal commands. Furthermore, priority rules and conflict resolution logic effectively coordinate mouse and keyboard commands, reducing false trigger rates and generating mixed event sequences that align with user intent. This significantly improves the accuracy and smoothness of IMU-based interactions in complex operation scenarios.
[0117] S80, converting the mixed event sequence of mouse and keyboard buttons into standard mouse input events and standard keyboard button input events of the target operating system, and injecting the standard mouse input events and the standard keyboard button input events into the event processing queue of the target operating system. This process aims to convert the mixed event sequence generated by the above steps, including mouse cursor movement instructions (such as two-dimensional displacement), mouse button instructions (such as single-click, double-click) and keyboard button instructions (such as Ctrl, Enter), into standard input events that can be directly recognized and executed by the operating system, thereby achieving a user interaction experience consistent with traditional mouse and keyboard operations. Through standard event conversion and injection, the system can seamlessly drive complex interactive tasks, such as file dragging, shortcut key operations or text input, ensure that the control method based on the inertial measurement unit (IMU) is compatible with the input mechanism of the target operating system, and support a variety of application scenarios, such as text editing, graphic design or virtual reality interaction.
[0118] In the implementation of step S80, the system first maps the mixed event sequence of mouse and keyboard buttons to the standard input events of the target operating system. For mouse events, in the Windows system, the cursor movement instruction is converted to MOUSEEVENTF_MOVE (relative displacement) or MOUSEEVENTF_ABSOLUTE (absolute coordinate) events, the key instruction is converted to MOUSEEVENTF_LEFTDOWN / UP (left button) or MOUSEEVENTF_RIGHTDOWN / UP (right button), and the wheel instruction (if applicable) is converted to MOUSEEVENTF_WHEEL (vertical scroll). For keyboard events, the keyboard key instruction is converted to KEYEVENTF_KEYDOWN / UP event, carrying virtual key codes (such as VK_CONTROL, VK_RETURN). In the macOS system, similar events are generated by the CGEventCreateMouseEvent and CGEventCreateKeyboardEvent functions. These standard events carry specific parameters (such as displacement, key code) and comply with the operating system input protocol. Subsequently, the system injects the event into the event processing queue through the input interface (such as SendInput of Windows or CGEventPost of macOS), and the operating system distributes it to the currently active application to trigger the corresponding interactive behavior.
[0119] To enhance the user experience, the system can provide feedback after event injection. For example, the device's haptic module generates vibration (e.g., a slight vibration lasting 0.2 seconds), the speaker emits a prompt sound (e.g., a short "beep"), or the interface displays cursor / button status animations. These feedback mechanisms help users confirm command execution, especially in contactless interactions, enhancing intuitive operation.
[0120] For example, suppose a user uses an IMU-equipped smartwatch to operate a text editor on a Windows computer. The mixed event sequence includes cursor movement ([16.9, 8.1] pixels, t=1000ms), left-click (t=1005ms), and Ctrl+Shift+press (t=1010ms). In step S80, the system converts the cursor movement into a MOUSEEVENTF_MOVE event (parameters [16.9, 8.1] pixels), the left-click into a MOUSEEVENTF_LEFTDOWN / UP event (t=1005ms), and the Ctrl+Shift+press into a KEYEVENTF_KEYDOWN event (virtual key code VK_CONTROL, t=1010ms). Through the SendInput function, these events are injected into the Windows event queue in timestamp order, triggering the Ctrl+C copy operation. After injection, the smartwatch generates 0.2 seconds of vibration feedback to notify the user that the operation is complete. This process ensures that the event sequence seamlessly drives shortcut key operations, providing an intuitive user experience.
[0121] As can be understood, by converting mixed event sequences into standard input events, the system achieves compatibility with traditional mouse and keyboard input, ensuring correct command execution. Furthermore, through event injection and multimodal feedback mechanisms, the system ensures the reliability and smoothness of interaction, while enhancing the user's perception of successful operation through tactile, auditory, or visual feedback. This significantly improves the practicality and user experience of IMU-based interaction in complex multimodal scenarios.
[0122] In addition, embodiments of the present invention further provide a computer-readable storage medium. The computer-readable storage medium can be any one of, or any combination of, a hard disk, a multimedia card, an SD card, a flash memory card, an SMC, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), or a USB memory device. The computer-readable storage medium includes an inertial measurement unit-based mouse and keyboard key control program 10. The specific implementation of the computer-readable storage medium of the present invention is substantially the same as the specific implementation of the inertial measurement unit-based mouse and keyboard key control method and server 1 described above, and will not be further described herein.
[0123] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0124] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0125] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A method for controlling a mouse and keyboard buttons based on an inertial measurement unit, characterized in that: include: Obtain the raw 3D motion data of the inertial measurement unit and perform multi-level preprocessing to generate a standard motion feature stream containing 3D attitude parameters and 3D angular velocity parameters; Applying multi-scale dynamic time window segmentation to the standardized motion feature stream to extract candidate gesture segments, and matching the candidate gesture segments with templates of a pre-constructed hierarchical motion primitive library to obtain a preliminary primitive sequence; Analyzing the preliminary primitive sequence according to the definition of the mode switching gesture in the pre-built hierarchical motion primitive library to identify the occurrence or end state of the mode switching gesture, and determining whether the current control mode is the mouse control mode or the keyboard control mode according to the recognition result of the mode switching gesture; If the current control mode is the mouse control mode, selecting parameters from the standardized motion data stream as basic cursor control input according to the current human-computer interaction context information, determining and applying a mapping algorithm corresponding to the currently effective sensitivity mapping mode to process the basic cursor control input to generate a mouse cursor movement instruction; and analyzing the portion of the preliminary primitive sequence that is not recognized as the mode switching gesture according to the definition of the mouse button gesture in the pre-built hierarchical motion primitive library, so as to recognize the mouse button gesture and parse it into a mouse button instruction; If the current control mode is the keyboard control mode, analyzing the portion of the preliminary primitive sequence that is not recognized as the mode switching gesture according to the definition of the keyboard key gesture in the pre-built hierarchical motion primitive library to recognize the keyboard key gesture and parse it into a keyboard key instruction; Combining the identified mouse cursor movement instructions, the mouse button instructions, and the keyboard button instructions to generate a mixed mouse and keyboard button event sequence; The mixed mouse and keyboard key event sequence is converted into standard mouse input events and standard keyboard key input events of the target operating system, and the standard mouse input events and the standard keyboard key input events are injected into the event processing queue of the target operating system.
2. The method for controlling a mouse and keyboard buttons based on an inertial measurement unit according to claim 1, wherein: The raw 3D motion data of the inertial measurement unit is acquired and multi-level preprocessing is performed to generate a standard motion feature stream containing 3D attitude parameters and 3D angular velocity parameters, including: Applying an adaptive low-pass filter to perform preliminary filtering on the original three-dimensional motion data, wherein a cutoff frequency of the adaptive low-pass filter is dynamically adjusted according to signal energy of the original three-dimensional motion data; Performing sensor fusion processing on the raw three-dimensional motion data after preliminary filtering using a Kalman filter to generate the three-dimensional posture parameters and the three-dimensional angular velocity parameters; The three-dimensional posture parameters and the three-dimensional angular velocity parameters are normalized and offset corrected to obtain the standardized motion feature flow.
3. The method for controlling a mouse and keyboard buttons based on an inertial measurement unit according to claim 1, wherein: Applying multi-scale dynamic time window segmentation to the standardized motion feature stream to extract candidate gesture segments, and matching the candidate gesture segments with templates of a pre-constructed hierarchical motion primitive library to obtain a preliminary primitive sequence, including: Applying a three-level parallel sliding window process comprising a micro-window, a medium-window, and a macro-window to the standardized motion feature stream, and dynamically adjusting the window size and overlap ratio of each level of windows according to an energy change rate and a complexity index of the standardized motion feature stream to extract the candidate gesture segments; The time domain statistical features, frequency domain distribution characteristics and posture trajectory features are extracted from the motion data in the candidate gesture segment, and the extracted features are matched with the primitive templates of the corresponding level in the pre-built hierarchical motion primitive library using a dynamic time warping algorithm or a support vector machine model to obtain the preliminary primitive sequence.
4. The method for controlling a mouse and keyboard buttons based on an inertial measurement unit according to claim 1, wherein: According to the definition of the mode switching gesture in the pre-built hierarchical motion primitive library, the preliminary primitive sequence is analyzed to identify the occurrence or end state of the mode switching gesture, including: The preliminary primitive sequence is matched with the predefined mode switching gesture template in the pre-built hierarchical motion primitive library using a dynamic time warping algorithm, and the occurrence or end of the mode switching gesture is determined in combination with the intention prediction probability based on the initial part data of the preliminary primitive sequence.
5. The method for controlling a mouse and keyboard buttons based on an inertial measurement unit according to claim 1, wherein: Selecting parameters from the standardized motion data stream as basic cursor control input according to current human-computer interaction context information, determining and applying a mapping algorithm corresponding to a currently effective sensitivity mapping mode to process the basic cursor control input to generate a mouse cursor movement instruction, including: Determining, through an application program interface or a user interface analysis tool provided by the operating system, the type of the currently active application or the type of the user interface element under the cursor as the current human-computer interaction context information; selecting, according to a preset context mapping table, a specific axial combination of the three-dimensional gesture parameters or a specific axial combination of the three-dimensional angular velocity parameters corresponding to the type of the currently active application or the type of the user interface element under the cursor as a basic cursor control input; detecting a motion amplitude or a motion speed associated with the basic cursor control input in the standard motion data stream, When the movement amplitude or the movement speed exceeds a preset first threshold, determining the preset relative mapping mode as the currently effective sensitivity mapping mode, When the movement amplitude or the movement speed is lower than a preset second threshold, determining the preset absolute mapping mode as the currently effective sensitivity mapping mode; Reading the physical mouse sensitivity setting of the operating system, and calculating the basic mouse sensitivity adjustment parameter according to the physical mouse sensitivity setting; The basic mouse sensitivity adjustment parameter is used as the gain adjustment factor of the mapping algorithm corresponding to the currently effective sensitivity mapping mode to process the basic cursor control input, and when the currently effective sensitivity mapping mode changes from the sensitivity mapping mode of the previous control cycle, a smooth transition processing based on weighted average is performed based on the current output of the preset mapping algorithm and the final output state data of the previous control cycle to generate the mouse cursor movement instruction.
6. The method for controlling a mouse and keyboard buttons based on an inertial measurement unit according to claim 1, wherein: According to the definition of mouse button gestures in the pre-built hierarchical motion primitive library, analyzing the portion of the preliminary primitive sequence that is not recognized as the mode switching gesture to recognize the mouse button gestures and parse them into mouse button commands, including: Applying a preset sequence matching rule to a portion of the preliminary primitive sequence that is not recognized as the mode switching gesture, the portion is matched with mouse button gesture templates defined in the pre-built hierarchical motion primitive library, where the mouse button gesture templates include a single-click gesture template, a double-click gesture template, and a long-press gesture template; Perform duration verification, motion amplitude verification, and confidence threshold filtering based on the current human-computer interaction context information on the successfully matched candidate mouse button gestures; The verified candidate mouse button gesture is parsed into a corresponding left-click instruction, right-click instruction, double-click instruction, or a button press and release instruction sequence to generate the mouse button instruction.
7. The method for controlling a mouse and keyboard buttons based on an inertial measurement unit according to claim 1, wherein: According to the definition of keyboard key gestures in the pre-built hierarchical motion primitive library, analyzing the portion of the preliminary primitive sequence that is not recognized as the mode switching gesture to recognize the keyboard key gestures and parse them into keyboard key commands, including: Matching the portion of the preliminary primitive sequence that is not recognized as the mode switching gesture with the keyboard key gesture template in the pre-built hierarchical motion primitive library using a dynamic time warping algorithm, and confirming the matching result by applying a disambiguation mechanism including adjusting the recognition threshold or priority according to the current application scenario to obtain a confirmed keyboard key gesture recognition result; A hierarchical state machine is applied to the confirmed keyboard key gesture recognition result for state tracking, to distinguish between continuous keyboard key gestures and sequential keyboard key gestures, and to select one of the instantaneous trigger mode, continuous trigger mode or locked trigger mode to parse it into the keyboard key instruction.
8. The method for controlling a mouse and keyboard buttons based on an inertial measurement unit according to claim 1, wherein: Combining the recognized mouse cursor movement instruction, the mouse button instruction, and the keyboard button instruction to generate a mixed mouse and keyboard button event sequence, including: Performing time stamp alignment and state information integration on the identified mouse cursor movement instructions, the mouse button instructions, and the keyboard button instructions; Predefined instruction priority rules and conflict resolution logic are applied to process the mouse cursor movement instruction, the mouse button instruction and the keyboard button instruction to generate a mixed mouse and keyboard button event sequence. The conflict resolution logic includes suppressing the mouse cursor displacement instruction caused by gesture action within the trigger time window of the mouse button instruction when the confidence of the mouse button instruction is higher than a preset button confirmation threshold, and increasing the button confirmation threshold for confirming the mouse button instruction when the mouse cursor movement instruction represents high-speed continuous movement.
9. A mouse and keyboard button control device based on an inertial measurement unit, characterized in that: The invention comprises a memory, a processor and a mouse and keyboard button control program based on an inertial measurement unit stored in the memory and executable on the processor. When the processor executes the mouse and keyboard button control program based on the inertial measurement unit, the mouse and keyboard button control method based on the inertial measurement unit as described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a mouse and keyboard button control program based on an inertial measurement unit. When the mouse and keyboard button control program based on an inertial measurement unit is executed by a processor, the mouse and keyboard button control method based on an inertial measurement unit as described in any one of claims 1-8 is implemented.