Method for mapping three-dimensional motion gesture into key combination
Through multi-level processing and timing correlation of IMU gesture data, combined with interactive context and semantic prediction models, efficient and accurate mapping of IMU gestures to letter keys and arrow key combinations is achieved, solving the problems of high learning costs and low recognition accuracy in the existing technology, and improving user experience and interaction efficiency.
Patent Information
- Application Number
- CN202510464938.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, when using inertial measurement unit (IMU) gesture recognition, it is difficult to efficiently and accurately generate a specific combination of letter keys and arrow keys. Especially in scenarios where continuous input is required, the existing methods have high learning cost and low recognition accuracy, and lack an intuitive mapping scheme.
By performing multi-level signal processing on the three-dimensional motion data collected by the IMU, basic gesture recognition and timing association are performed, combined with the current interactive context and semantic prediction model, intelligent pairing and command analysis of letter-direction gestures are realized, and the timing association engine and semantic prediction model are applied for command checksum feedback.
It significantly reduces user learning costs, improves the accuracy and logic of command generation, enhances the contextual relevance and user experience of interaction, and ensures efficient execution of combined commands.
Smart Images

Figure CN120295482A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and particularly to a method, device, equipment and computer-readable storage medium for mapping three-dimensional motion gestures to key combinations. Background Art
[0002] The continuous evolution of human-computer interaction technology aims to break the limitations of traditional input methods and provide users with more natural and efficient means of device control. Among many emerging interaction technologies, three-dimensional motion gesture recognition based on inertial measurement units (IMUs) has shown broad application potential in fields such as wearable computing, assistive technology, and virtual reality, due to its ability to directly capture users' limb movements and good portability. Existing research has explored using IMUs for various interaction tasks, such as controlling cursor movement through wrist or head postures, triggering system-level commands or application functions through specific dynamic gestures, and even attempting continuous text input. These technologies have enriched the interaction dimension to a certain extent, especially in scenarios where it is inconvenient to use traditional keyboards and mice, providing alternative input options for users.
[0003] However, when applying IMU gestures to scenarios that require precise keyboard input, especially operations involving combinations of letter keys and arrow keys, the existing technologies have obvious limitations. On the one hand, if independent and distinguishable IMU gestures are designed for each letter, users need to learn and remember a large set of gestures. This not only has a high learning cost, but also as the number of gestures increases, the distinguishability between them will significantly decrease, making it difficult to guarantee the recognition accuracy, especially in real-time and continuous input scenarios. On the other hand, although there have been some attempts at IMU-based text input (such as air handwriting simulation or specific pattern mapping), they are usually not efficient or mainly focus on generating single characters, and it is difficult to smoothly and efficiently achieve the combined output of letter keys and arrow keys. Such combinations are significant in specific applications (such as code navigation in integrated development environments, shortcut commands in certain strategy or role-playing games, customized data table operation interfaces), but currently, there is a lack of an intuitive mapping scheme based on IMU gestures specifically for such combined commands.
[0004] Therefore, there is a gap in the prior art in efficiently and accurately generating specific combinations of letter keys and direction keys using IMU gestures. Directly mapping a large number of combinations to independent gestures is impractical, and existing single-letter input methods or general command mapping methods cannot meet the requirements of such specific combination operations. This requires the development of a new technical paradigm that can intelligently interpret the user's continuous three-dimensional motion gestures, distinguish whether the intention is to input a certain letter or perform a certain direction action, and dynamically combine the two into a unified "letter key + direction key" key command based on their temporal relevance and the current interaction context. Exploring the basic gestures that can distinguish different functional categories (such as letter type and direction type) and establishing a method for judging the temporal correlation mechanism between them has important practical significance and research value for filling this technical gap and realizing a richer and more refined IMU gesture keyboard control function. Summary of the Invention
[0005] An embodiment of the present application provides a method for mapping three-dimensional motion gestures to key combinations, aiming to provide an efficient, accurate, intelligent, and user-friendly human-computer interaction solution.
[0006] To achieve the above object, an embodiment of the present application provides a method for mapping three-dimensional motion gestures to key combinations, including:
[0007] Performing multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit to obtain a standardized motion feature stream;
[0008] Performing basic gesture recognition processing including multi-set classification on the standardized motion feature stream to obtain a recognition result time series sequence including basic gesture category identifiers and belonging information of the predefined gesture set to which they belong;
[0009] Processing the recognition result time series sequence by applying a temporal correlation engine to obtain letter-direction gesture pairing information or a separate basic gesture command type, wherein processing by applying the temporal correlation engine determines the pairing relationship between adjacent letter gestures and direction gestures based on the belonging information of the gesture set and the time interval between gesture occurrences;
[0010] Applying mapping rules based on the current interaction context to the letter-direction gesture pairing information or the separate basic gesture command type for command parsing processing, and verifying or correcting the letter-related commands parsed by combining a semantic prediction model to obtain a target letter key and direction key combination command identifier or a separate key command identifier;
[0011] Applying rules including basic gesture quality verification, temporal correlation rationality verification, and context validity verification to the target letter key and direction key combination command identifier or the separate key command identifier for command verification processing to obtain a command identifier that passes the verification;
[0012] For the command identifier that passes the verification, perform target platform event conversion and injection processing to execute the corresponding instantaneous key combination or individual key action, and simultaneously perform multi-modal feedback signal generation processing corresponding to the command identifier, execution status, or semantic prediction result, to obtain the standard input event injected into the target platform and the simultaneously generated multi-modal feedback signal.
[0013] To achieve the above object, an embodiment of the present application further provides a device for mapping three-dimensional motion gestures to key combinations, including:
[0014] An inertial measurement unit interface, configured to receive the original three-dimensional motion data collected from the inertial measurement unit;
[0015] A signal processing module, connected to the inertial measurement unit interface, configured to perform multi-level signal processing on the original three-dimensional motion data to obtain a standardized motion feature stream;
[0016] A multi-set basic gesture recognition module, connected to the signal processing module, configured to perform basic gesture recognition processing including multi-set classification on the standardized motion feature stream to obtain a recognition result time series sequence including basic gesture category identifiers and belonging pre-defined gesture set attribution information;
[0017] A time series association module, connected to the multi-set basic gesture recognition module, configured to process the recognition result time series sequence by applying a time series association engine to determine the pairing relationship between adjacent letter gestures and direction gestures, to obtain letter-direction gesture pairing information or individual basic gesture command types;
[0018] A context mapping and semantic parsing module, connected to the time series association module, configured to perform command parsing processing on the letter-direction gesture pairing information or individual basic gesture command types by applying mapping rules based on the current interaction context, and perform verification or correction in combination with a semantic prediction model, to obtain a target letter key and direction key combination command identifier or an individual key command identifier;
[0019] A command verification module, connected to the context mapping and semantic parsing module, configured to perform command verification processing on the target command identifier by applying rules including basic gesture quality verification, time series association rationality verification, and context validity verification, to obtain a command identifier that passes the verification;
[0020] An event injection and feedback module, connected to the command verification module, is configured to perform target platform event conversion and injection processing on the verified command identifier to execute corresponding instantaneous key combinations or individual key actions, and simultaneously perform multi-modal feedback signal generation processing corresponding to the command identifier, execution status, or semantic prediction result, so as to output standard input events injected into the target platform and the simultaneously generated multi-modal feedback signals.
[0021] To achieve the above object, an embodiment of the present application further provides a device for mapping three-dimensional motion gestures to key combinations, including a memory, a processor, and a program for mapping three-dimensional motion gestures to key combinations stored on the memory and executable on the processor. When the processor executes the program for mapping three-dimensional motion gestures to key combinations, the method for mapping three-dimensional motion gestures to key combinations as described in any one of the above is implemented.
[0022] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, on which a program for mapping three-dimensional motion gestures to key combinations is stored. When the program for mapping three-dimensional motion gestures to key combinations is executed by a processor, the method for mapping three-dimensional motion gestures to key combinations as described in any one of the above is implemented.
[0023] The application of mapping three-dimensional motion gestures to combinations of letter keys and direction keys proposed by the present application has the following beneficial effects:
[0024] First of all, the present application uses a timing correlation engine to process the timing sequence of recognition results containing gesture set attribution information. This mechanism performs intelligent pairing by determining the tight time correlation between adjacent basic gestures belonging to the letter gesture set and the direction gesture set respectively. This application allows users to generate "letter key + direction key" combination commands by performing the intuitive action of a gesture corresponding to a letter followed immediately by a gesture corresponding to a direction, avoiding the problem of designing and memorizing independent complex gestures or long sequences for hundreds of possible letter-direction combinations, significantly reducing the user's learning cost and cognitive load, and making the way of generating such specific combination commands more natural and efficient.
[0025] Secondly, the present application introduces multi-set classification at the basic gesture recognition stage, clearly outputting the functional category (letter, direction, or control) to which each basic gesture belongs. This feature provides a key judgment basis for the subsequent timing correlation engine, enabling it to accurately identify the intention conforming to the "letter followed by direction" pattern, effectively distinguishing combination commands from individual letter inputs or control commands, and improving the accuracy and logic of command generation. Combined with the SVM classifier based on small sample training and incremental learning, it ensures the rapid customization and continuous adaptation to the user's personalized basic gestures.
[0026] Furthermore, in the command parsing stage, this application introduces mapping rules and semantic prediction models based on the current interaction context. Context mapping enables the same letter-direction gesture pairing to be mapped to the most suitable operation in different applications, enhancing the situational relevance and intelligence of the interaction. By combining the semantic prediction model to verify or correct commands involving letter keys, and using the statistical laws of language to assist in judging or correcting potential letter gesture recognition errors, the accuracy and robustness of combined command generation are significantly improved in scenarios such as text editing and programming that require letter input.
[0027] In addition, the command verification of this application includes verification for the rationality of temporal association, which is specifically used to determine whether the time interval between letter gestures and direction gestures is within a preset valid window. This helps filter out incorrect pairings caused by reaction lag or accidental consecutive actions. Combining the basic gesture quality verification and context validity verification further improves the reliability of the final output command. At the same time, the adaptive multi-modal feedback containing predictive information not only confirms the user's operation but also provides candidate information when there is recognition ambiguity or semantic prediction for correction, enabling users to participate in the interaction more actively and correct errors efficiently, optimizing the user experience and interaction efficiency. Finally, this application ensures that the generated combined command is executed in an instantaneous key-press manner, which conforms to the actual behavior pattern of letter and direction key combinations in most applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0029] Figure 1 It is a module structure diagram of an embodiment of the device for mapping three-dimensional motion gestures to key combinations in the present invention;
[0030] Figure 2 It is a flowchart of an embodiment of the method for mapping three-dimensional motion gestures to key combinations in the present invention.
[0031] The realization of the objectives, functional features, and advantages of the present invention will be further described in conjunction with the embodiments and with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] To better understand the above technical solutions, exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0034] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" as used herein does not exclude the presence of components or steps not listed in the claims. The quantifier "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware including several different components and by means of a properly programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of "first", "second", and "third", etc. does not denote any order and these words can be interpreted as names.
[0035] As Figure 1 shown, Figure 1 is a schematic structural diagram of server 1 (also called a device that maps three-dimensional motion gestures to key combinations) in the hardware operating environment related to the embodiment solution of the present invention.
[0036] The server in the embodiment of the present invention, such as "Internet of Things devices", intelligent air conditioners with networking functions, intelligent lights, intelligent power supplies, AR / VR devices with networking functions, intelligent speakers, autonomous driving vehicles, PCs, smartphones, tablets, e-book readers, portable computers, and other devices with display functions.
[0037] As Figure 1 shown, the server 1 includes: a memory 11, a processor 12, and a network interface 13.
[0038] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 11 can be an internal storage unit of the server 1 in some embodiments, such as the hard disk of the server 1. The memory 11 can also be an external storage device of the server 1 in other embodiments, such as a plug-in hard disk equipped on the server 1, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0039] Further, the memory 11 may also include the internal storage unit of the server 1 and external storage devices. The memory 11 can be used not only to store the application software installed in the server 1 and various types of data, such as the code of the program 10 that maps three-dimensional motion gestures to key combinations, but also to temporarily store the data that has been output or will be output.
[0040] In some embodiments, the processor 12 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips, and is used to run the program code stored in the memory 11 or process data, such as executing the program 10 that maps three-dimensional motion gestures to key combinations.
[0041] The network interface 13 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the server 1 and other electronic devices.
[0042] The network can be the Internet, a cloud network, a wireless fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN), and / or a metropolitan area network (MAN). Various devices in the network environment can be configured to connect to the communication network according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols can include, but are not limited to, at least one of the following: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Li-Fi, 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocol, and / or Bluetooth communication protocol or a combination thereof.
[0043] Optionally, the server may further include a user interface. The user interface may include a display, an input unit such as a keyboard, and the optional user interface may also include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be referred to as a display screen or a display unit, and is used to display the information processed in the server 1 and to display a visual user interface.
[0044] Figure 1Only server 1 with components 11 - 13 and program 10 that maps three - dimensional motion gestures to key combinations is shown. Those skilled in the art can understand that Figure 1 the shown structure does not constitute a limitation on server 1. It may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0045] In this embodiment, the processor 12 can be used to call the program stored in the memory 11 that maps three - dimensional motion gestures to key combinations and perform the following operations:
[0046] Perform multi - level signal processing on the raw three - dimensional motion data collected from the inertial measurement unit to obtain a normalized motion feature stream;
[0047] Perform basic gesture recognition processing including multi - set classification on the normalized motion feature stream to obtain a recognition result time - series sequence including basic gesture category identifiers and belonging information of the predefined gesture set to which they belong;
[0048] Apply a time - series correlation engine to process the recognition result time - series sequence to obtain letter - direction gesture pairing information or a separate basic gesture command type. Among them, applying the time - series correlation engine for processing is based on the gesture set belonging information and the time interval between gesture occurrences to determine the pairing relationship between adjacent letter gestures and direction gestures;
[0049] Apply mapping rules based on the current interaction context to perform command parsing processing on the letter - direction gesture pairing information or the separate basic gesture command type, and combine a semantic prediction model to verify or correct the parsed letter - related commands to obtain a target letter key and direction key combination command identifier or a separate key command identifier;
[0050] Apply rules including basic gesture quality verification, time - series correlation rationality verification, and context validity verification to perform command verification processing on the target letter key and direction key combination command identifier or the separate key command identifier to obtain a command identifier that passes the verification;
[0051] Perform target platform event conversion and injection processing on the command identifier that passes the verification to execute the corresponding instantaneous key combination or separate key action, and simultaneously perform multi - modal feedback signal generation processing corresponding to the command identifier, execution status, or semantic prediction result to obtain a standard input event injected into the target platform and the simultaneously generated multi - modal feedback signals.
[0052] In one embodiment, the processor 12 can be used to call the program stored in the memory 11 that maps three - dimensional motion gestures to key combinations and perform the following operations:
[0053] Apply a digital low-pass filter to the original three-dimensional motion data for filtering to obtain the filtered three-dimensional motion data;
[0054] Apply a Kalman filter to the filtered three-dimensional motion data for sensor fusion processing to generate fusion motion data containing real-time attitude information;
[0055] Apply a motion state detection algorithm to the fusion motion data. By calculating the short-time energy, signal amplitude variance, and main frequency components of the fusion motion data within a sliding time window and comparing them with a motion activation threshold dynamically adjusted based on the user's recent activity level, identify and screen out potential gesture signal segments representing the user's intention to execute gestures;
[0056] Map the data of the potential gesture motion segment to a preset range to obtain the standardized motion feature stream.
[0057] In one embodiment, the processor 12 can be used to call a program stored in the memory 11 that maps three-dimensional motion gestures to key combinations and perform the following operations:
[0058] Apply a dynamic segmentation algorithm based on signal variance and derivative thresholds to the standardized motion feature stream for processing, and real-time define the start and end time points of a single basic gesture action to obtain a gesture time segment;
[0059] Perform feature extraction processing on the fusion motion data within the gesture time segment, and calculate a multi-dimensional feature set including time-domain statistical quantities, frequency-domain distribution characteristics, rotation and displacement characteristics based on pose estimation, and autoregressive model coefficients;
[0060] Apply the principal component analysis method to the multi-dimensional feature set for dimensionality reduction processing to obtain a dimensionality-reduced gesture feature vector;
[0061] Input the dimensionality-reduced gesture feature vector into a pre-trained linear kernel support vector machine classification model for classification processing to determine the basic gesture category identifier corresponding to the gesture time segment and its belonging information in the gesture set;
[0062] Organize the continuously determined basic gesture category identifiers and the belonging information of the gesture set in chronological order to obtain the recognition result time series.
[0063] In one embodiment, the processor 12 can be used to call a program stored in the memory 11 that maps three-dimensional motion gestures to key combinations and perform the following operations:
[0064] Initialize and maintain a short-term memory buffer containing the recently recognized basic gesture identifier, gesture set belonging information, and time stamp;
[0065] When a new recognition result is received, determine the attribution information of the gesture set:
[0066] If it is determined to belong to the letter gesture set, update the buffer, cache the basic letter gesture information, and start a correlation determination time window timer with a preset duration.
[0067] If it is determined to belong to the direction gesture set, check the buffer and the current time to determine whether there is cached basic letter gesture with a time stamp within the preset correlation determination time window. If so, perform letter-direction gesture pairing processing to generate the letter-direction gesture pairing information containing the corresponding letter gesture identifier and direction gesture identifier, and clear the buffer and the timer. If not, ignore the direction gesture or generate a separate direction command type according to the preset rules.
[0068] If it is determined to belong to the control gesture set, directly generate the corresponding control command type, and clear the buffer and the timer;
[0069] Monitor the correlation determination time window timer. If the timer times out and the cached basic letter gesture in the buffer is not successfully paired, generate the separate basic gesture command type or ignore it according to the preset rules, and clear the buffer and the timer;
[0070] According to the results of the pairing processing, ignoring processing, or direct generation processing, obtain the letter-direction gesture pairing information, or the separate basic gesture command type, or the control command type.
[0071] In one embodiment, the processor 12 can be used to call the program stored in the memory 11 that maps three-dimensional motion gestures to key combinations, and perform the following operations:
[0072] Obtain the context information of the current active application by querying the operating system application programming interface;
[0073] According to the context information, select and activate a matching mapping rule library from multiple predefined context-specific mapping rule libraries;
[0074] Apply the activated mapping rule library to the letter-direction gesture pairing information or the separate basic gesture command type for command parsing processing, and convert it into a preliminary letter key and direction key combination command identifier or a separate key command identifier or a control command identifier;
[0075] Apply a pre-trained semantic prediction model to the part of the preliminary command identifier that involves letter keys, and calculate the occurrence probability of the part of the preliminary command identifier that involves letter keys in the current text context;
[0076] Verify or selectively correct the preliminary command identifier according to the occurrence probability and the classification confidence of the corresponding letter-based gesture. If the confidence is low and the semantic probability supports the alternative option, make a correction to obtain the target letter key and direction key combination command identifier or the single key command identifier or the control command identifier.
[0077] In one embodiment, the processor 12 can be used to call a program stored in the memory 11 that maps three-dimensional motion gestures to key combinations, and perform the following operations:
[0078] Evaluate whether the classification confidence score and the motion feature parameters of the basic gesture that constitutes the command identifier are within a preset normal range to perform basic gesture quality verification;
[0079] If the command identifier is derived from a letter-direction gesture pairing, check the time interval between the paired letter gesture and the direction gesture to perform temporal association rationality verification;
[0080] For the operation mapped by the command identifier, look up a preset behavior rule library according to the current interaction context to perform context validity verification;
[0081] Combine and determine the verification processing results of the basic gesture quality, temporal association rationality, and context validity to obtain the command identifier that passes the verification.
[0082] In one embodiment, the processor 12 can be used to call a program stored in the memory 11 that maps three-dimensional motion gestures to key combinations, and perform the following operations:
[0083] Perform target platform event conversion processing on the command identifier that passes the verification to obtain a standard input event data structure that conforms to the input specification of the target operating system;
[0084] Call the underlying input simulation interface function of the target operating system, and inject the standard input event data structure into the system event queue according to the parsing time sequence to obtain the standard input event injected into the target platform;
[0085] According to the verified command identifier, query a preset feedback mode mapping table to generate a feedback instruction including a specified feedback channel, a feedback mode, and a feedback intensity level, where the feedback channel includes at least one of a tactile channel, an auditory channel, and a visual channel, and the feedback mode includes at least one of a preset vibration waveform, an audio sample, and a visual element style;
[0086] Apply adaptive adjustment processing to the feedback instruction. According to the real-time environmental parameters obtained from the environmental sensor or the preference parameters set by the user through the configuration interface, adjust the intensity level or channel priority in the feedback instruction, and then output through the corresponding feedback actuator to obtain the synchronously generated multi-modal feedback signal.
[0087] Based on the hardware architecture of the device that maps three-dimensional motion gestures to key combinations described above, an embodiment of the method of the present invention for mapping three-dimensional motion gestures to key combinations is proposed. Refer to Figure 2 , Figure 2 This is an embodiment of the method of the present invention for mapping three-dimensional motion gestures to key combinations. The method of mapping three-dimensional motion gestures to key combinations includes the following steps:
[0088] S10. Perform multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit (IMU) to obtain a standardized motion feature stream. The original data here usually includes the three-axis acceleration and angular velocity information provided by the accelerometer and gyroscope, reflecting the motion state of the user's limb in three-dimensional space. In order to facilitate subsequent gesture recognition and mapping, these original data must be converted into a consistent and processable feature form.
[0089] The following details this process through steps S11 to S14 and illustrates it with examples.
[0090] In step S11, apply a digital low-pass filter to the original three-dimensional motion data for filtering to obtain the filtered three-dimensional motion data. The data collected by the IMU often contains high-frequency noise, such as device jitter or environmental interference, which will interfere with the accuracy of subsequent gesture recognition. The role of the digital low-pass filter is to retain the low-frequency signals (the slow changes related to the user's gestures) and filter out the high-frequency noise. For example, assume that the IMU samples data at a sampling rate of 100Hz, and the acceleration data at a certain moment is 5.2, -3.1, 9.8m / s 2 , which contains noise components. By setting a low-pass filter with a cut-off frequency of 10Hz, the data can be smoothed, and the filtered result such as 5.0, -3.0, 9.7m / s 2 can be output. This process ensures the smoothness of the data and lays a foundation for subsequent processing.
[0091] Next, in step S12, the Kalman filter is applied to the filtered three-dimensional motion data for sensor fusion processing to generate fused motion data containing real-time attitude information. An IMU typically includes an accelerometer and a gyroscope. The former measures linear acceleration, and the latter measures angular velocity. However, single-sensor data may have drift or errors. The Kalman filter estimates the real-time attitude of the device (such as Euler angles or quaternions) by fusing these two types of data. For example, assume the filtered acceleration data is [0, 0, 9.8] m / s 2 (at rest, gravity direction), and the gyroscope data is [0.1, 0.2, 0] rad / s (slight rotation). The Kalman filter combines the historical states and current measurements of both to output fused attitude data, such as the quaternion [0.999, 0.01, 0.02, 0], which reflects the spatial orientation of the device. This fusion improves the reliability and consistency of the data.
[0092] Then, in step S13, a motion state detection algorithm is applied to the fused motion data. By calculating the short-time energy, signal amplitude variance, and main frequency components within a sliding time window and comparing them with a dynamically adjusted motion activation threshold, potential gesture signal segments representing the user's intention to perform a gesture are identified. When the user uses the IMU, there may be unintentional movements (such as natural arm swings), and the intentional gestures need to be filtered out. The short-time energy reflects the signal strength, the amplitude variance indicates the degree of fluctuation, and the main frequency component reveals the periodicity of the movement. For example, assume the fused data records acceleration changes within a 1-second sliding window, and the calculated short-time energy is 50, the amplitude variance is 2.5, and the main frequency is 1 Hz. If the user's recent activity level is low and the dynamic threshold is set to 40, then this segment of data exceeds the threshold and is identified as a potential gesture signal segment. This method adapts to the user's behavior through the dynamic threshold to ensure the accuracy of the screening.
[0093] Finally, in step S14, the data of the potential gesture motion segment is mapped to a preset range to obtain a normalized motion feature stream. Since different users may apply different forces when performing the same gesture, or the same user may vary in force in different situations, directly using the values of the original amplitudes will cause trouble to the subsequent recognition model. Therefore, normalization processing is required to scale the data values (e.g., the acceleration or angular velocity in a certain axis) to a fixed interval, such as [-1, 1] or [0, 1], or to make it conform to the standard normal distribution (mean is 0, standard deviation is 1). Common methods include Min - Max normalization (formula: (value - min) / (max - min), and then adjusted to the target range) or Z - score normalization (formula: (value - mean) / std_dev). Here, min, max, mean, and std_dev are usually calculated based on the data within the currently recognized "potential gesture signal segment". For example, within a detected gesture segment, the X - axis angular velocity range is from - 150 degrees / second to + 200 degrees / second. After using Min - Max normalization to the [-1, 1] range, the point with an original value of 0 degrees / second will be mapped to ((0 - (-150)) / (200 - (-150)))*2 - 1=(150 / 350)*2 - 1≈0.857 - 1=-0.143. After such processing, regardless of the original action amplitude, the relative form of its motion pattern is retained, but the values are unified into the standard range, forming a normalized motion feature stream.
[0094] For example, assume that the user triggers the "letter A + right arrow key" command by performing a "wave right" gesture through the IMU. The initial data is collected from the IMU, and step S11 receives an acceleration data sequence, such as [0.1, 0.2, 9.8], [2.5, 0.3, 9.7], [5.0, 0.1, 9.6] m / s 2 (sampling frequency 100 Hz). After applying the low - pass filter, the data is smoothed to [0.1, 0.2, 9.8], [2.4, 0.3, 9.7], [4.8, 0.1, 9.6] m / s 2. In step S12, the acceleration and gyroscope data (such as [0,0,0], [1.2,0,0], [2.5,0,0] rad / s) are fused by a Kalman filter to output an attitude sequence, such as quaternions [0.999,0,0,0], [0.98,0.19,0,0], [0.95,0.31,0,0]. In step S13, the short-time energy (set to 60), amplitude variance (set to 3.0), and dominant frequency (set to 2 Hz) within the window are calculated and compared with the threshold of 40 to confirm that this segment is a gesture signal. In step S14, the acceleration is mapped to -1,1 -1,1 -1,1 to obtain a normalized feature stream, such as [0.01,0.02,0.98], [0.48,0.06,0.97], [0.96,0.02,0.96], for subsequent recognition use.
[0095] It can be understood that removing noise through a digital low-pass filter ensures the cleanliness of the data; sensor fusion by the Kalman filter improves the accuracy of attitude estimation; the motion state detection algorithm combined with dynamic thresholds accurately distinguishes intentional and unintentional actions; and data normalization unifies the feature scales. These steps work together to transform the original, potentially noisy and inconsistent IMU data into a high-quality normalized motion feature stream, providing a reliable basis for subsequent gesture classification and command mapping. Compared with directly processing the raw data, this multi-level processing scheme significantly improves the robustness and adaptability of the system, especially in complex gesture recognition scenarios.
[0096] S20. Perform basic gesture recognition processing including multi-set classification on the normalized motion feature stream to obtain a recognition result time series sequence including basic gesture category identifiers and the belonging information of the predefined gesture set to which they belong. Here, the normalized motion feature stream is a continuous signal stream obtained from IMU data through preprocessing. The goal is to identify the specific gestures performed by the user through a series of steps and classify them into predefined sets (such as letter classes, direction classes).
[0097] To achieve this process, step S20 can be completed step by step through S21 to S25.
[0098] In step S21, a dynamic segmentation algorithm based on signal variance and derivative threshold determination is applied to the standardized motion feature stream to process it, and the start and end time points of a single basic gesture action are defined in real time, obtaining a gesture time segment. The standardized feature stream is continuous, and independent gesture actions need to be segmented from it. Signal variance reflects the degree of data fluctuation, and the derivative represents the rate of change. The combination of the two can detect the start and end of an action. For example, assuming the feature stream data is [0.1, 0.2, 0.3], [0.5, 0.7, 0.9], [0.8, 0.4, 0.2] (each group is three-axis acceleration), after calculating the variance and derivative, it is found that the variance of the second group is 0.04 (higher than the threshold 0.02), and the derivative is [0.4, 0.5, 0.6] (higher than the threshold 0.3), which is marked as the start of the gesture; the variance of the third group drops to 0.09, and the derivative is -0.1, -0.3, -0.7 (lower than the threshold), which is marked as the end. In this way, from [0.5, 0.7, 0.9] to [0.8, 0.4, 0.2] is defined as a gesture time segment. This method ensures the precise division of the gesture boundary.
[0099] Next, in step S22, feature extraction processing is performed on the fused motion data within the gesture time segment, calculating a multi-dimensional feature set including time-domain statistics, frequency-domain distribution characteristics, rotation and displacement characteristics based on pose estimation, and autoregressive model coefficients. The fused motion data (such as quaternions or accelerations) contains rich information, and multi-faceted features need to be extracted to comprehensively describe the gesture. Time-domain statistics include the mean and variance, the frequency-domain characteristics obtain the main frequency through Fourier transform, the rotation and displacement are calculated from the pose data, and the autoregressive coefficients reflect the time dependence. For example, for the acceleration data [0.5, 0.7, 0.9], [0.6, 0.5, 0.8] within the time segment, the calculated mean is [0.55, 0.6, 0.85], the variance is [0.005, 0.01, 0.005], the main frequency obtained by Fourier transform is 1 Hz, the calculated rotation angle for the pose change is 5°, the displacement is 0.1 m, and the autoregressive coefficients are [0.8, 0.1]. These features together constitute a multi-dimensional feature set to capture the dynamic characteristics of the gesture.
[0100] Then, in step S23, the principal component analysis (PCA) method is applied to the multi-dimensional feature set for dimensionality reduction processing, obtaining a dimensionality-reduced gesture feature vector. The multi-dimensional feature set has a high dimension and may contain redundant information. PCA retains the main variance through linear transformation. For example, assuming the feature set is a 5-dimensional vector [0.55, 0.005, 1, 5, 0.8], after PCA calculation, the first 2 principal components are retained (such as 90% of the variance), and the dimensionality-reduced feature vector [0.62, 0.35] is output. This step reduces the computational complexity while retaining the key information.
[0101] In step S24, the dimension-reduced gesture feature vector is input into a pre-trained linear kernel support vector machine (SVM) classification model for classification processing to determine the basic gesture category identifier corresponding to the gesture time segment and the belonging information of the gesture set to which it belongs. The SVM differentiates different gesture categories through a hyperplane, and the linear kernel is suitable for fast classification. Suppose the model is pre-trained with two types of gestures, "letter A" (set: letter category) and "right" (set: direction category). When the input feature vector is [0.62, 0.35], the SVM outputs a classification result of "letter A" with a confidence of 0.9 and belonging to the "letter category". This process realizes the precise classification of gestures.
[0102] Finally, in step S25, the continuously determined basic gesture category identifiers and the belonging information of the gesture sets are organized in chronological order to obtain the recognition result time series sequence. Multiple gesture recognition results need to be arranged in the order of occurrence to form a sequence. For example, continuously recognizing "letter A" (t = 1s) and "right" (t = 2s), the output sequence is [("A", "letter category", 1s), ("right", "direction category", 2s)]. This provides a structured input for subsequent time series association.
[0103] For example, assume the user executes the "letter A + right arrow key" command, first waves to represent "A", and then swipes right to represent "right". After the standardized feature stream is input, step S21 divides it into two segments: the first segment is from [0.5, 0.7, 0.9] to [0.6, 0.5, 0.8] (t = 1.0s - 1.2s), and the second segment is from [0.1, 0.3, 0.2] to [0.8, 0.9, 0.1] (t = 2.0s - 2.2s). Step S22 extracts the features of the first segment: mean [0.55, 0.6, 0.85], main frequency 1Hz, rotation 5°; the features of the second segment: mean [0.45, 0.6, 0.15], main frequency 2Hz, displacement 0.2m. After step S23 reduces the dimension, the first segment is [0.62, 0.35], and the second segment is [0.75, 0.20]. Step S24 classifies the first segment as "letter A" (letter category, confidence 0.9), and the second segment as "right" (direction category, confidence 0.95). Step S25 generates the sequence [("A", "letter category", 1.0s), ("right", "direction category", 2.0s)].
[0104] It can be understood that the dynamic segmentation algorithm accurately defines the gesture boundaries, the multi-dimensional feature extraction comprehensively describes the action characteristics, the PCA dimension reduction optimizes the calculation efficiency, the SVM classification ensures high accuracy, and the time series sequence organization lays the foundation for subsequent pairing. These steps together transform the standardized feature stream into a structured gesture recognition result, significantly improving the robustness and practicality of multi-set classification, especially in complex continuous gesture scenarios.
[0105] S30. Apply a temporal correlation engine to process the recognition result time series to obtain letter-direction gesture pairing information or individual basic gesture command types. Among them, applying the temporal correlation engine for processing is based on the gesture set attribution information and the time interval of gesture occurrence, and determines the pairing relationship between adjacent letter gestures and direction gestures. Here, the time series is a series containing gesture category identifiers, set attribution information, and timestamps obtained from the previous steps. The goal is to determine whether they form a combined command (such as the pairing of letters and directions) by analyzing the time relationship and category of gestures. The core of the processing lies in determining the pairing relationship between adjacent letter gestures and direction gestures based on the gesture set attribution information and the occurrence time interval.
[0106] To achieve this process, step S30 can be gradually completed through S31 to S37.
[0107] In step S31, initialize and maintain a short-term memory buffer containing the identifiers of the most recently recognized basic gestures, gesture set attribution information, and timestamps. The role of the buffer is to temporarily store gesture information for subsequent analysis of its temporal correlation. For example, initialize a buffer with a capacity of 3, which is initially empty. As gesture recognition results arrive, the buffer is dynamically updated. For example, when receiving the gesture ("A", "letter type", 1.0s), the buffer becomes [("A", "letter type", 1.0s)]. This step provides a data basis for temporal correlation.
[0108] Next, in step S32, when a new recognition result is received, determine the gesture set attribution information. This is the core entry of the temporal correlation engine, and the subsequent processing logic is determined according to the set to which the gesture belongs (letter type, direction type, or control type). For example, when receiving ("A", "letter type", 1.0s), it is determined to belong to the letter type and enter S33; if receiving ("right", "direction type", 2.0s), it enters S34. This determination ensures the differential processing of different types of gestures.
[0109] In step S33, if it is determined to belong to the letter gesture set, update the buffer, cache the letter basic gesture information, and start a correlation determination time window timer with a preset duration. Letter gestures may be part of a combined command and need to wait for potential direction gestures. For example, when receiving ("A", "letter type", 1.0s), the buffer is updated to [("A", "letter type", 1.0s)], and the time window timer is started (such as 1 second). This mechanism provides a time constraint for pairing.
[0110] Then, in step S34, if it is determined to belong to the direction gesture set, the buffer and the current time are checked to determine whether there are cached letter gestures with timestamps within the preset associated determination time window. If there are, letter-direction gesture pairing information is generated and the buffer and timer are cleared; if not, it is processed according to the preset rules. For example, if the buffer is [("A","letter type",1.0s)] and ("right","direction type",1.5s) is received, the time interval of 0.5s is less than the window of 1s, and the pairing information ("A","right") is generated and the buffer is cleared. If ("right","direction type",2.5s) is received and the interval of 1.5s exceeds the window, a separate command "right" is generated or ignored. This logic realizes dynamic pairing.
[0111] In step S35, if it is determined to belong to the control gesture set, the corresponding control command type is directly generated, and the buffer and timer are cleared. Control gestures usually execute independently and do not participate in pairing. For example, if ("confirm","control type",3.0s) is received, the "confirm" command is directly output and the buffer is cleared. This processing ensures the immediacy of the control command.
[0112] Next, in step S36, the associated determination time window timer is monitored. If the timer times out and the letter gestures cached in the buffer are not paired, a separate command is generated or ignored according to the preset rules, and the buffer is cleared. For example, if the buffer is [("A","letter type",1.0s)] and the window times out after 1 second to 2.0s and no direction gesture is received, a separate command "A" is generated and the buffer is cleared. This mechanism avoids the hanging state of uncompleted pairing.
[0113] Finally, in step S37, according to the various results of the foregoing processing, letter-direction gesture pairing information, separate basic gesture command types, or control command types are obtained. For example, the pairing process outputs ("A","right"), the timeout process outputs "A", and the control process outputs "confirm". These results provide inputs for subsequent command parsing.
[0114] For example, assume that the user executes the commands "Letter A + Right Arrow Key" and "Confirm" separately. The timing sequence is [("A", "Letter type", 1.0s), ("Right", "Direction type", 1.5s), ("Confirm", "Control type", 3.0s)]. Step S31 initializes the buffer to be empty. Step S32 receives ("A", "Letter type", 1.0s), S33 updates the buffer to [("A", "Letter type", 1.0s)] and starts a 1 - second timer. Step S32 receives ("Right", "Direction type", 1.5s), step S34 checks the buffer, the interval 0.5s < 1s, generates ("A", "Right"), and clears the buffer. S32 receives ("Confirm", "Control type", 3.0s), step S35 generates the "Confirm" command and clears the buffer. Step S36 monitors that there is no timeout event. Step S37 outputs the result as [("A", "Right"), "Confirm"].
[0115] It can be understood that the short - term memory buffer and the time - window timer together implement the dynamic pairing of gestures. The decision logic based on set membership and time interval ensures the accuracy and flexibility of the pairing. The timeout handling and the direct generation of control commands enhance the robustness of the system. These steps transform the discrete gesture recognition results into meaningful combinations or individual commands, significantly improving the practicality and efficiency of complex gesture input, especially performing excellently in scenarios that require continuous operations.
[0116] S40. Apply mapping rules based on the current interaction context to the letter - direction gesture pairing information or individual basic gesture command types for command parsing and processing, and combine with a semantic prediction model to verify or correct the letter - related commands parsed, to obtain the target letter key and direction key combination command identifier or individual key command identifier. The input here is the pairing or individual command generated by the timing association engine, and the goal is to map these gestures to specific keyboard operations while ensuring that the results meet the requirements of the current scenario.
[0117] To achieve this process, step S40 can be completed step by step through S41 to S45.
[0118] In step S41, obtain the context information of the current active application by querying the operating system application programming interface (API). The context information reflects the user's current operating environment, such as the software being used or the interface state. For example, in the Windows system, obtain the handle of the current active window by calling the GetForegroundWindow() function, and then extract the window title using GetWindowText(), such as "Visual Studio Code". This indicates that the user may be editing code, and the context information is "code editor".
[0119] Next, in step S42, according to the context information, a matching mapping rule library is selected and activated from multiple predefined context-specific mapping rule libraries. Different application scenarios require different gesture mapping rules. For example, the system maintains three rule libraries: code editor (A + right = cursor move right), game (A + right = character move), general (A + right = undefined). If the context is "code editor", the corresponding rule library is activated. This selection ensures the pertinence of the mapping.
[0120] In step S43, for the letter-direction gesture pairing information or individual basic gesture command types, the activated mapping rule library is applied for command parsing processing, and converted into a preliminary combination command identifier of letter keys and direction keys or an individual key command identifier or a control command identifier. The rule library defines the mapping relationship from gestures to keys. For example, inputting the pairing information ("A", "right"), it is mapped to "A key + right arrow key" in the code editor rule library, and the output preliminary identifier is ("A", "Right"); inputting the individual command "A", it is mapped to the individual "A key", and the output is ("A"). This step realizes the preliminary conversion from gestures to keys.
[0121] Then, in step S44, for the part involving letter keys in the preliminary command identifier, a pre-trained semantic prediction model is applied for processing to calculate its occurrence probability in the current text context. The semantic prediction model (such as an LSTM-based language model) predicts the rationality of letters based on the current text. For example, in a code editor, the current input is "print(", and the preliminary identifier is ("A", "Right"). The model calculates that the probability of "A" appearing after "print(" is 0.1 (low, because it should be lowercase or a parenthesis), while the probability of "a" is 0.6. This step provides a basis for correction.
[0122] Finally, in step S45, according to the occurrence probability and the classification confidence of the letter-based gesture, the preliminary command identifier is verified or selectively corrected. If the confidence is low and the semantic probability supports an alternative option, correction is made to obtain the target combination command identifier of letter keys and direction keys or an individual key command identifier or a control command identifier. For example, the confidence of "A" in ("A", "Right") is 0.85, the probability 0.1 is lower than the threshold 0.5, while the probability of "a" is 0.6, and it is corrected to ("a", "Right"); if the confidence of the individual command "A" is 0.9 and the probability is 0.8, it is retained as ("A"). This ensures the accuracy of the command.
[0123] For example, assume that the user inputs "letter A + right arrow key" and "confirm" alone in the code editor. The input is [("A", "right"), "confirm"]. In step S41, the API is queried to obtain the context "Visual Studio Code". In step S42, the code editor rule library is activated. In step S43, ("A", "right") is parsed as ("A", "Right"), and "confirm" is parsed as ("Enter"). In step S44, after the text "print(", the model calculates the probability of "A" as 0.1, the probability of "a" as 0.6, and there is no prediction required for "Enter". In step S45, since the confidence of "A" is 0.85 but the probability is low, it is corrected to ("a", "Right"), and "confirm" remains as ("Enter"), and the output is [("a", "Right"), ("Enter")]. If the current cursor is after "print(", after execution, it becomes "print(a" and a new line is started.
[0124] It can be understood that querying the API to obtain the context ensures the adaptability of the mapping scenario, selecting a specific rule library improves the pertinence of the command, the initial mapping realizes the conversion from gesture to key, and semantic prediction and correction enhance the accuracy of letter input. These steps together transform the abstract gesture command into a precise key operation that conforms to the context, significantly improving the practicality of the system and the user experience, especially performing excellently in diverse application scenarios.
[0125] S50. Apply rules including basic gesture quality verification, temporal correlation rationality verification, and context validity verification to the combined command identifier of the target letter key and the direction key or the single key command identifier to perform command verification processing, and obtain a command identifier that passes the verification. The input here is the command identifier parsed from the previous steps, and the goal is to ensure its accuracy and rationality through multi-dimensional verification to avoid misoperations.
[0126] In some embodiments, step S50 can be gradually completed through S51 to S54.
[0127] In step S51, for the basic gesture constituting the command identifier, evaluate whether the classification confidence score and the motion feature parameters are within the preset normal range to perform basic gesture quality verification. Gesture recognition may have errors due to signal quality or non-standard user actions, and its reliability needs to be checked. For example, for the command identifier ("a", "Right"), the classification confidence of its basic gesture "A" is 0.85 (threshold 0.8), and the motion feature (such as the acceleration variance 0.04) is within the normal range [0.02, 0.1], and the verification passes. If the confidence of the single command ("Enter") is 0.95 and the variance is 0.03, it also passes. This step ensures the quality of the gesture itself.
[0128] Next, in step S52, if the command identifier is derived from an alphabet-direction gesture pair, the time interval between the paired alphabet gesture and direction gesture is checked for temporal association rationality verification. The paired command depends on the temporal correlation of the gestures, and too long an interval may indicate a non-intentional combination. For example, for ("a", "Right"), the timestamp of "A" is 1.0 s and that of "Right" is 1.5 s, with an interval of 0.5 s which is less than the maximum allowed interval of 1 s, so the verification passes. If the interval is 1.2 s, it fails and is split into separate commands. This check ensures the rationality of the pairing.
[0129] Then, in step S53, for the operation mapped by the command identifier, the preset behavior rule library is searched according to the current interaction context for context validity verification. The command needs to conform to the logic of the current application scenario. For example, in the context of "code editor", ("a", "Right") is mapped to "input a and move the cursor to the right", and the behavior rule library allows this operation, so the verification passes. If there is no such definition in "game", it fails. The single command ("Enter") is valid in most contexts and also passes. This verification ensures the applicability of the command.
[0130] Finally, in step S54, the verification results of the basic gesture quality, temporal association rationality, and context validity are combined for determination to obtain the command identifier that passes the verification. Considering the three results comprehensively, only the commands that pass all of them are accepted. For example, for ("a", "Right"), the quality (confidence 0.85, variance 0.04), timing (interval 0.5 s), and context (allowed in code editor) all pass, and ("a", "Right") is output; if the interval of ("b", "Up") is 1.2 s and fails, it is split into ("b"), and if the context passes again, finally ("b") is output. This step integrates multiple safeguards.
[0131] For example, assume the input is [("a","Right"),("Enter")] and the context is "code editor". Step S51 checks the confidence of "A" in ("a","Right") is 0.85 with variance 0.04, the confidence of "Right" is 0.9 with variance 0.05, and the confidence of ("Enter") is 0.95 with variance 0.03. All are within the range [0.02,0.1] and the confidence > 0.8, so it passes. Step S52 checks the interval of ("a","Right") is 0.5s < 1s, so it passes. Step S53 checks in the code editor rule library, ("a","Right") is a valid operation and ("Enter") is for line break, both pass. Step S54 makes a combined determination and outputs [("a","Right"),("Enter")]. If the interval of ("b","Up") is 1.2s > 1s, only ("b") passes the context check and ("b") is output.
[0132] It can be understood that quality check ensures the reliability of gesture recognition, timing check maintains the logical consistency of pairing, context check guarantees the applicability of operations in scenarios, and combined determination provides comprehensive guarantee. These steps together filter potential errors, optimize the command identifier into an accurate and executable output, significantly improving the stability of the system and user trust, especially prominent in complex interactions.
[0133] S60. For the command identifier that passes the check, perform target platform event conversion and injection processing to execute the corresponding instantaneous key combination or individual key action, and simultaneously perform multi-modal feedback signal generation processing corresponding to the command identifier, execution status, or semantic prediction result, to obtain the standard input event injected into the target platform and the simultaneously generated multi-modal feedback signal. The input here is the command identifier that has passed multiple checks, and the goal is to convert these commands into actual system operations and enhance the user experience through feedback.
[0134] In some embodiments, step S60 can be gradually completed through S61 to S64.
[0135] In step S61, perform target platform event conversion processing on the command identifier that passes the check to obtain a standard input event data structure that conforms to the input specification of the target operating system. The operating system (such as Windows or Linux) requires specific formats of events to simulate key presses. For example, the command identifier ("a","Right") is converted into an INPUT structure in Windows: a keyboard event for pressing "a" (key code 0x41), a keyboard event for pressing "Right" (key code 0x27), and the corresponding release events. The single command ("Enter") is converted into a single press and release event (key code 0x0D). This step ensures the compatibility of the command with the platform.
[0136] Next, in step S62, the underlying input simulation interface function of the target operating system is called to inject the standard input event data structure into the system event queue according to the parsing timing sequence, obtaining the standard input events injected into the target platform. For example, in Windows, the SendInput() function is used to inject the event sequence of ("a","Right"): first press and release "a", then press and release "Right"; ("Enter") is injected with one press and release. If executed in a code editor, "a" is entered after the cursor and shifted to the right, or a new line is created. This process realizes the actual execution of the key presses.
[0137] Then, in step S63, according to the command identifier that passes the verification, the preset feedback mode mapping table is queried to generate a feedback instruction including the specified feedback channel, feedback mode, and feedback intensity level. The feedback channels include touch (such as vibration), hearing (such as sound), and vision (such as light), and the mode and intensity are predefined. For example, the mapping table stipulates that ("a","Right") corresponds to tactile vibration (short pulse, intensity 5), auditory prompt tone (500Hz, intensity 3); ("Enter") corresponds to tactile long vibration (intensity 7). The generated instructions are [("tactile","short pulse",5),("auditory","500Hz",3)] and [("tactile","long vibration",7)]. This enhances the user's perception of the operation.
[0138] Finally, in step S64, adaptive adjustment processing is applied to the feedback instruction. After adjusting the intensity or channel priority according to the real-time environmental parameters or user preferences, it is output through the feedback actuator to obtain the synchronously generated feedback signal. For example, the environmental sensor detects that the noise level is 60dB, and the user prefers touch first. Adjust the ("a","Right") instruction: the auditory intensity is reduced to 1, and the tactile intensity is increased to 6. The output is vibration (intensity 6) and low sound (intensity 1); ("Enter") maintains the tactile intensity of 7. If the user is in a quiet environment, the original settings are maintained. The actuator (such as a motor, speaker) outputs the feedback synchronously.
[0139] For example, assume the input is [("a","Right"),("Enter")] and it is executed in a Windows code editor. S61 converts ("a","Right") into INPUT events: [("a", pressed),("a", released),("Right", pressed),("Right", released)], and ("Enter") into [("Enter", pressed),("Enter", released)]. S62 injects them using SendInput(), and the text “print(” becomes “print(a” and a new line is added. S63 queries the mapping table and generates feedback: for ("a","Right") it is [("Tactile","Short pulse",5),("Auditory","500Hz",3)], and for ("Enter") it is [("Tactile","Long vibration",7)]. S64 detects a noise of 50dB and adjusts it to [("Tactile","Short pulse",6),("Auditory","500Hz",2)] and [("Tactile","Long vibration",7)], and the user feels the vibration and a faint beep.
[0140] It can be understood that event conversion and injection achieve seamless execution from gestures to keystrokes. Feedback mode mapping and adaptive adjustment enhance the interaction perception through multi-modal signals, and environment and preference adaptation enhance the user experience. These steps together ensure the efficient execution of commands and personalized feedback, significantly improving the usability and comfort of the system, especially in long-term or complex operations.
[0141] In addition, an embodiment of the present invention also proposes a device for mapping three-dimensional motion gestures to key combinations. The device for mapping three-dimensional motion gestures to key combinations includes:
[0142] An inertial measurement unit interface for receiving the original three-dimensional motion data collected from an inertial measurement unit;
[0143] A signal processing module connected to the inertial measurement unit interface and configured to perform multi-level signal processing on the original three-dimensional motion data to obtain a standardized motion feature stream;
[0144] A multi-set basic gesture recognition module connected to the signal processing module and configured to perform basic gesture recognition processing including multi-set classification on the standardized motion feature stream to obtain a recognition result time series sequence including basic gesture category identifiers and belonging information of the predefined gesture set to which they belong;
[0145] A time series correlation module connected to the multi-set basic gesture recognition module and configured to process the recognition result time series sequence using a time series correlation engine to determine the pairing relationship between adjacent letter gestures and direction gestures, and obtain letter-direction gesture pairing information or a single basic gesture command type;
[0146] The context mapping and semantic parsing module is connected to the temporal correlation module and is configured to perform command parsing processing on the letter-direction gesture pairing information or the individual basic gesture command types by applying mapping rules based on the current interaction context, and combine with a semantic prediction model for verification or correction to obtain a target letter key and direction key combination command identifier or an individual button command identifier;
[0147] The command verification module is connected to the context mapping and semantic parsing module and is configured to perform command verification processing on the target command identifier by applying rules including basic gesture quality verification, temporal correlation rationality verification, and context validity verification to obtain a command identifier that passes the verification;
[0148] The event injection and feedback module is connected to the command verification module and is configured to perform target platform event conversion and injection processing on the command identifier that passes the verification to execute corresponding instantaneous button combinations or individual button actions, and synchronously perform multi-modal feedback signal generation processing corresponding to the command identifier, execution status, or semantic prediction result to output the standard input event injected into the target platform and the synchronously generated multi-modal feedback signals.
[0149] Among them, the steps implemented by each functional module of the device that maps three-dimensional motion gestures to button combinations can refer to the respective embodiments of the method for mapping three-dimensional motion gestures to button combinations of the present invention, which will not be elaborated here.
[0150] In addition, an embodiment of the present invention also proposes a computer-readable storage medium. The computer-readable storage medium can be any one or any combination of a hard disk, a multimedia card, an SD card, a flash card, an SMC, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, etc. The computer-readable storage medium includes a program 10 for mapping three-dimensional motion gestures to button combinations. The specific implementation manner of the computer-readable storage medium of the present invention is substantially the same as the specific implementation manners of the above method for mapping three-dimensional motion gestures to button combinations and the server 1, which will not be elaborated here.
[0151] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0152] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0153] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0155] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0156] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for mapping three-dimensional motion gestures to key combinations, characterized in that, Including: Performing multi-level signal processing on the original three-dimensional motion data collected from an inertial measurement unit to obtain a standardized motion feature stream; Performing basic gesture recognition processing including multi-set classification on the standardized motion feature stream to obtain a recognition result time series sequence including basic gesture category identifiers and the belonging information of the predefined gesture sets to which they belong; Processing the recognition result time series sequence using a time series correlation engine to obtain letter-direction gesture pairing information or a separate basic gesture command type, wherein processing using the time series correlation engine determines the pairing relationship between adjacent letter gestures and direction gestures based on the gesture set belonging information and the gesture occurrence time interval; Applying mapping rules based on the current interaction context to the letter-direction gesture pairing information or the separate basic gesture command type for command parsing processing, and verifying or correcting the letter-related commands parsed by combining a semantic prediction model to obtain a target letter key and direction key combination command identifier or a separate key command identifier; Applying rules including basic gesture quality verification, time series correlation rationality verification, and context validity verification to the target letter key and direction key combination command identifier or the separate key command identifier for command verification processing to obtain a command identifier that passes the verification; Performing target platform event conversion and injection processing on the command identifier that passes the verification to execute the corresponding instantaneous key combination or separate key action, and simultaneously performing multi-modal feedback signal generation processing corresponding to the command identifier, execution status, or semantic prediction result to obtain a standard input event injected into the target platform and the simultaneously generated multi-modal feedback signal; 2. The method for mapping a three-dimensional motion gesture to a key combination as claimed in claim 1, wherein Performing multi-level signal processing on the original three-dimensional motion data collected from an inertial measurement unit, including: Applying a digital low-pass filter to the original three-dimensional motion data for filtering processing to obtain filtered three-dimensional motion data; Applying a Kalman filter to the filtered three-dimensional motion data for sensor fusion processing to generate fusion motion data including real-time attitude information; Applying a motion state detection algorithm to the fusion motion data for processing, by calculating the short-time energy, signal amplitude variance, and main frequency components of the fusion motion data within a sliding time window, and comparing with a motion activation threshold dynamically adjusted based on the user's recent activity level, to identify and screen out potential gesture signal segments representing the user's intention to perform gestures; Mapping the data of the potential gesture motion segment to a preset range to obtain the standardized motion feature stream; 3. The method of mapping three-dimensional motion gestures to key combinations according to claim 1, characterized in that, Performing basic gesture recognition processing including multi-set classification on the standardized motion feature stream, including: Applying a dynamic segmentation algorithm based on signal variance and derivative thresholds to the standardized motion feature stream for processing to real-time define the start and end time points of a single basic gesture action and obtain a gesture time segment; Performing feature extraction processing on the fusion motion data within the gesture time segment to calculate a multi-dimensional feature set including time domain statistics, frequency domain distribution characteristics, rotation and displacement characteristics based on pose estimation, and autoregressive model coefficients; Apply the principal component analysis method to the multi-dimensional feature set for dimensionality reduction processing to obtain the gesture feature vector after dimensionality reduction; Input the gesture feature vector after dimensionality reduction into a pre-trained linear kernel support vector machine classification model for classification processing to determine the basic gesture category identifier corresponding to the gesture time segment and the belonging information of the gesture set to which it belongs; Organize the continuously determined basic gesture category identifiers and the belonging information of the gesture set in chronological order to obtain the recognition result time series sequence.
4. The method for mapping three-dimensional motion gestures to key combinations according to claim 1, wherein, Apply a time series correlation engine to the recognition result time series sequence, including: Initialize and maintain a short-term memory buffer containing the recently recognized basic gesture identifier, gesture set belonging information, and timestamp; When a new recognition result is received, determine the belonging information of the gesture set: If it is determined to belong to the letter gesture set, update the buffer, cache the letter basic gesture information, and start a correlation determination time window timer with a preset duration; If it is determined to belong to the direction gesture set, check the buffer and the current time to determine whether there is a cached letter basic gesture with a timestamp within the preset correlation determination time window. If so, perform the letter-direction gesture pairing process to generate the letter-direction gesture pairing information containing the corresponding letter gesture identifier and direction gesture identifier, and clear the buffer and the timer. If not, ignore the direction gesture or generate a separate direction command type according to the preset rules; If it is determined to belong to the control gesture set, directly generate the corresponding control command type and clear the buffer and the timer; Monitor the correlation determination time window timer. If the timer times out and the letter basic gesture cached in the buffer is not successfully paired, generate the separate basic gesture command type or ignore it according to the preset rules, and clear the buffer and the timer; Obtain the letter-direction gesture pairing information, or the separate basic gesture command type, or the control command type according to the results of the pairing process, ignoring process, or direct generation process.
5. The method of mapping three-dimensional motion gestures to key combinations according to claim 1, characterized in that, Apply the mapping rules based on the current interaction context to the letter-direction gesture pairing information or the separate basic gesture command type for command parsing processing, and combine the semantic prediction model to verify or correct the letter-related commands parsed, including: Obtain the context information of the current active application by querying the operating system application programming interface; Select and activate a matching mapping rule library from multiple predefined context-specific mapping rule libraries according to the context information; Apply the activated mapping rule library to the letter-direction gesture pairing information or the separate basic gesture command type for command parsing processing, and convert it into a preliminary combination command identifier of letter keys and direction keys, or a separate key command identifier, or a control command identifier; Apply a pre-trained semantic prediction model to the part of the preliminary command identifier that involves letter keys, and calculate the occurrence probability of the part of the preliminary command identifier that involves letter keys in the current text context; Verify or selectively correct the preliminary command identifier according to the occurrence probability and the classification confidence of the corresponding letter-based gesture. If the confidence is low and the semantic probability supports the alternative option, make a correction to obtain the target letter key and direction key combination command identifier, or the single key command identifier, or the control command identifier.
6. The method for mapping three-dimensional motion gestures to key combinations according to claim 1, characterized in that Apply rules including basic gesture quality verification, temporal correlation rationality verification, and context validity verification to the target letter key and direction key combination command identifier or the single key command identifier for command verification processing, including: Evaluate whether the classification confidence score and motion feature parameters of the basic gesture constituting the command identifier are within the preset normal range to perform basic gesture quality verification; If the command identifier is derived from a letter-direction gesture pairing, check the time interval between the paired letter gesture and direction gesture to perform temporal correlation rationality verification; For the operation mapped by the command identifier, search the preset behavior rule library according to the current interaction context to perform context validity verification; Combine and determine the verification results of the basic gesture quality, temporal correlation rationality, and context validity to obtain the command identifier that passes the verification.
7. The method for mapping three-dimensional motion gestures to key combinations according to claim 1, characterized in that, For the command identifier that passes the verification, perform target platform event conversion and injection processing to execute the corresponding instantaneous key combination or single key action, and simultaneously perform multi-modal feedback signal generation processing corresponding to the command identifier, execution status, or semantic prediction result, including: Perform target platform event conversion processing on the command identifier that passes the verification to obtain a standard input event data structure that conforms to the input specification of the target operating system; Call the underlying input simulation interface function of the target operating system, and inject the standard input event data structure into the system event queue according to the parsing time sequence to obtain the standard input event injected into the target platform; According to the command identifier that passes the verification, query the preset feedback mode mapping table to generate a feedback instruction including a specified feedback channel, feedback mode, and feedback intensity level, where the feedback channel includes at least one of a tactile channel, an auditory channel, and a visual channel, and the feedback mode includes at least one of a preset vibration waveform, audio sample, and visual element style; Apply adaptive adjustment processing to the feedback instruction, and adjust the intensity level or channel priority in the feedback instruction according to the real-time environmental parameters obtained from the environmental sensor or the preference parameters set by the user through the configuration interface, and then output through the corresponding feedback actuator to obtain the synchronously generated multi-modal feedback signal.
8. A device for mapping three-dimensional motion gestures to key combinations, characterized in that, Including: An inertial measurement unit interface for receiving the original three-dimensional motion data collected from the inertial measurement unit; A signal processing module, connected to the inertial measurement unit interface, configured to perform multi-level signal processing on the original three-dimensional motion data to obtain a standardized motion feature stream; A multi-set basic gesture recognition module, connected to the signal processing module, configured to perform basic gesture recognition processing including multi-set classification on the standardized motion feature stream to obtain a recognition result time series sequence including basic gesture category identifiers and belonging pre-defined gesture set attribution information; A time series correlation module, connected to the multi-set basic gesture recognition module, configured to process the recognition result time series sequence by applying a time series correlation engine to determine the pairing relationship between adjacent letter gestures and direction gestures, and obtain letter-direction gesture pairing information or a separate basic gesture command type; A context mapping and semantic parsing module, connected to the time series correlation module, configured to perform command parsing processing on the letter-direction gesture pairing information or the separate basic gesture command type by applying mapping rules based on the current interaction context, and perform verification or correction in combination with a semantic prediction model to obtain a target letter key and direction key combination command identifier or a separate key command identifier; A command verification module, connected to the context mapping and semantic parsing module, configured to perform command verification processing on the target command identifier by applying rules including basic gesture quality verification, time series correlation rationality verification, and context validity verification to obtain a verified command identifier; An event injection and feedback module, connected to the command verification module, configured to perform target platform event conversion and injection processing on the verified command identifier to execute corresponding instantaneous key combinations or separate key actions, and simultaneously perform multi-modal feedback signal generation processing corresponding to the command identifier, execution status, or semantic prediction result to output a standard input event injected into the target platform and the simultaneously generated multi-modal feedback signals.
9. A device for mapping three-dimensional motion gestures to key combinations, characterized in that, It includes a memory, a processor, and a program for mapping three-dimensional motion gestures to key combinations stored on the memory and executable on the processor. When the processor executes the program for mapping three-dimensional motion gestures to key combinations, it implements the method for mapping three-dimensional motion gestures to key combinations as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A program for mapping three-dimensional motion gestures to key combinations is stored on the computer-readable storage medium. When the program for mapping three-dimensional motion gestures to key combinations is executed by the processor, it implements the method for mapping three-dimensional motion gestures to key combinations as described in any one of claims 1-7.
Citation Information
Cited By
Virtual scene interaction system and method based on inertial measurement unit
CN121560167A
A virtual scene interaction system and method based on an inertial measurement unit
CN121560167B