Application control method for mapping three-dimensional motion data into letter keys
By preprocessing and feature extraction of the three-dimensional motion data obtained by the IMU sensor, combined with multi-layer classification and anti-touch detection, the problems of insufficient gesture recognition accuracy and high error-touch rate in the prior art are solved, and efficient and accessible letter key input is achieved to adapt to different users and environments.
Patent Information
- Application Number
- CN202510464903.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
When converting three-dimensional motion data into letter keys, existing IMU-based gesture recognition technology is susceptible to user motion instability, environmental interference and sensor drift, lacks recognition accuracy, is difficult to adapt to different gesture habits of different users, has high error touch rate, lacks context perception ability and semantic understanding, and is difficult to achieve intelligent prediction and error correction.
By obtaining three-dimensional motion data from the inertial measurement unit sensor, preprocessing and motion state detection, extracting feature sets and adjusting context-aware weights, combining multi-layer classification architecture, personalized adaptive mechanism and feature stability protection algorithm, multi-stage dimensionality reduction processing is performed, initial key maps are generated and multi-layer anti-touch detection is performed, and feedback signals are finally generated and applied to the target application.
It improves the system's recognition accuracy and response speed in complex environments, adapts to changes in user gesture habits, reduces the error operation rate, provides multimodal feedback to enhance operation confirmation, and expands the application scope.
Smart Images

Figure CN120295479A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and particularly to an application control method, device, equipment and computer-readable storage medium for mapping three-dimensional motion data to letter keys. Background Art
[0002] With the rapid development of interaction technology, the human-computer interaction method based on gesture recognition has gradually become a research hotspot, especially in the fields of wearable devices, augmented reality, virtual reality, and barrier-free technology. Traditional gesture recognition technologies mainly rely on cameras for visual capture. Although they perform well in static environments, they have obvious limitations in low-light, severely occluded, or moving scenarios. In recent years, gesture recognition technologies based on inertial measurement unit (IMU) sensors have gradually attracted the attention of the academic and industrial communities due to their advantages such as being unaffected by environmental light, being able to work in three-dimensional space, and having low hardware costs. Currently, there are already various products on the market that use IMU sensors to achieve simple gesture control, such as smart bracelets and motion-sensing game controllers. However, most of these products only support a limited number of preset gesture commands and are difficult to meet complex input requirements.
[0003] However, existing IMU-based gesture recognition technologies still face many challenges when converting three-dimensional motion data into fine inputs such as letter keys. First, the process of gesture data acquisition and processing is easily affected by factors such as unstable user movements, environmental interference, and sensor drift, resulting in insufficient recognition accuracy. Second, traditional gesture classification algorithms are difficult to adapt to the differences in gesture habits of different users and lack personalized adaptability. Third, the false touch rate is relatively high in complex environments (such as during walking or riding in a vehicle), affecting the user experience. In addition, existing systems usually separate the gesture recognition and letter mapping processes, lacking context awareness and semantic understanding, and are difficult to implement intelligent prediction and error correction functions. These problems severely restrict the application expansion of IMU gesture control technology in fine control scenarios such as text input.
[0004] Facing the above technical problems, there is an urgent need for a comprehensive solution that can accurately recognize three-dimensional gestures and intelligently map them to letter keys, so as to provide users with a more natural, efficient, and barrier-free interaction method. Summary of the Invention
[0005] An embodiment of the present application provides an application control method for mapping three-dimensional motion data to letter keys, aiming to provide users with a more natural, efficient, and barrier-free interaction method.
[0006] To achieve the above object, an embodiment of the present application provides an application control method for mapping three-dimensional motion data to letter keys, including:
[0007] Obtain the raw multi-modal three-dimensional motion data of the user from the inertial measurement unit sensor and perform preprocessing to obtain standardized three-dimensional motion data, and perform motion state detection on the standardized three-dimensional motion data to obtain potential gesture action data;
[0008] Extract a feature set from the potential gesture action data, and adjust the weights of the feature set based on context information to obtain context-weighted feature data, and perform multi-stage dimensionality reduction processing on the context-weighted feature data to obtain a highly discriminative feature vector;
[0009] Input the highly discriminative feature vector into a classification system for processing, and through a multi-layer classification architecture, a personalized adaptive mechanism, and a feature stability protection algorithm, perform comprehensive analysis in combination with application scenario information to obtain a gesture category identifier and its confidence score;
[0010] Based on the gesture category identifier and its confidence score, generate an initial key mapping through a multi-level mapping, and adjust the initial key mapping based on user preferences, context prediction, and application environment to obtain a final key event;
[0011] Perform multi-layer anti-misoperation detection on the final key event to obtain a verified target key event;
[0012] Generate multi-modal user feedback for the target key event to obtain a key event with a feedback signal, and apply the target key event with the feedback signal to a target application program to implement application program control.
[0013] To achieve the above object, an application control device for mapping three-dimensional motion data to alphabet keys according to an embodiment of the present application further includes:
[0014] A data acquisition and preprocessing module, configured to obtain the raw multi-modal three-dimensional motion data of the user from the inertial measurement unit sensor and perform preprocessing to obtain standardized three-dimensional motion data, and perform motion state detection on the standardized three-dimensional motion data to obtain potential gesture action data;
[0015] A feature extraction and dimensionality reduction module, connected to the data acquisition and preprocessing module, configured to extract a feature set from the potential gesture action data, adjust the weights of the feature set based on context information to obtain context-weighted feature data, and perform multi-stage dimensionality reduction processing on the context-weighted feature data to obtain a highly discriminative feature vector;
[0016] The gesture recognition and classification module, connected to the feature extraction and dimensionality reduction module, is used to input the high-discriminative feature vectors into a classification system for processing. Through a multi-layer classification architecture, a personalized adaptive mechanism, and a feature stability protection algorithm, combined with application scenario information for comprehensive analysis, gesture category identifiers and their confidence scores are obtained;
[0017] The key mapping module, connected to the gesture recognition and classification module, is used to generate an initial key mapping through multi-level mapping based on the gesture category identifiers and their confidence scores, and adjust the initial key mapping based on user preferences, context prediction, and application environment to obtain a final key event;
[0018] The accidental touch prevention module, connected to the key mapping module, is used to perform multi-layer accidental touch detection on the final key event to obtain a verified target key event;
[0019] The feedback and application control module, connected to the accidental touch prevention module, is used to generate multi-modal user feedback for the target key event to obtain a key event with a feedback signal, and apply the target key event with the feedback signal to a target application program to achieve application program control.
[0020] To achieve the above object, an application control device for mapping three-dimensional motion data into letter keys according to an embodiment of the present application further includes a memory, a processor, and an application control program for mapping three-dimensional motion data into letter keys stored on the memory and executable on the processor. When the processor executes the application control program for mapping three-dimensional motion data into letter keys, the application control method for mapping three-dimensional motion data into letter keys as described in any one of the above is implemented.
[0021] To achieve the above object, an embodiment of the present application further proposes a computer-readable storage medium. An application control program for mapping three-dimensional motion data into letter keys is stored on the computer-readable storage medium. When the application control program for mapping three-dimensional motion data into letter keys is executed by a processor, the application control method for mapping three-dimensional motion data into letter keys as described in any one of the above is implemented.
[0022] The application control method provided by the present invention for mapping three-dimensional motion data to letter keys has the following beneficial effects: By using an inertial measurement unit sensor to collect three-dimensional motion data and performing filtering, calibration, and standardization processing, the problem of environmental light limitation is solved, enabling the system to work normally in dark environments and under occlusion conditions; at the same time, this method analyzes the standardized three-dimensional motion data through a motion state detection algorithm, effectively distinguishing meaningful gesture actions from daily random motions, greatly reducing system resource consumption and improving system response speed. In terms of feature processing, the proposed context-aware feature extraction and multi-stage dimensionality reduction processing technology enables the system to dynamically adjust feature weights according to the user's current usage scenario and device location, solving the problem of decreased accuracy of traditional gesture recognition in complex environments; adopting a cascaded classification architecture combined with a personalized adaptive mechanism and a feature stability protection algorithm enables the system to adapt to changes in the user's gesture habits without losing the learned key features, overcoming the "catastrophic forgetting" problem in traditional incremental learning. At the application level, the multi-level mapping mechanism of the present invention combines frequency adaptive remapping and context prediction functions, automatically mapping frequently used letters to simpler and more stable gestures, reducing the user's learning cost; the multi-level anti-misoperation detection mechanism verifies from three dimensions: intention confirmation, context conflict, and timing logic, effectively reducing the misoperation rate; the multi-modal feedback system provides comprehensive feedback of vision, audition, and touch, enhancing the operation confirmation, especially being able to automatically adjust the best feedback method in different usage environments to ensure that the user always obtains clear feedback. In addition, by designing interfaces adapted to various application programs, this method can be flexibly deployed in various application scenarios, expanding the application scope of three-dimensional gesture control. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following described drawings are only some embodiments of the present invention, and for those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0024] Figure 1 It is a module structure diagram of an embodiment of an application control device for mapping three-dimensional motion data to letter keys according to the present invention;
[0025] Figure 2 It is a flowchart of an embodiment of an application control method for mapping three-dimensional motion data to letter keys according to the present invention.
[0026] The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not used to limit the present invention.
[0028] To better understand the above technical solutions, the exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0029] As Figure 1 shown, Figure 1 Figure 1 is a schematic structural diagram of a server 1 (also called an application control device for mapping three-dimensional motion data to letter keys) in the hardware operating environment related to the embodiment solution of the present invention.
[0030] The server in the embodiment of the present invention, such as "Internet of Things devices", intelligent air conditioners with networking functions, intelligent lights, intelligent power supplies, AR / VR devices with networking functions, intelligent speakers, autonomous driving vehicles, PCs, smart phones, tablets, e-book readers, portable computers, and other devices with display functions.
[0031] As Figure 1 shown, the server 1 includes: a memory 11, a processor 12, and a network interface 13.
[0032] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 can be an internal storage unit of the server 1 in some embodiments, such as the hard disk of the server 1. The memory 11 can also be an external storage device of the server 1 in other embodiments, such as a plug-in hard disk equipped on the server 1, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 11 can also include both the internal storage unit and the external storage device of the server 1. The memory 11 can be used not only to store application software installed on the server 1 and various types of data, such as the code of the application control program 10 that maps three-dimensional motion data to alphabet keys, etc., but also to temporarily store data that has been output or will be output. The processor 12 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 11 or process data, such as executing the application control program 10 that maps three-dimensional motion data to alphabet keys. The network interface 13 can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the server 1 and other electronic devices. The network can be the Internet, a cloud network, a wireless fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN), and / or a metropolitan area network (MAN). Various devices in the network environment can be configured to connect to the communication network according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols can include, but are not limited to, at least one of the following: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Li-Fi, 802.16, IEEE802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocol, and / or Bluetooth communication protocol or a combination thereof.
[0033] Optionally, the server may further include a user interface, which may include a display, an input unit such as a keyboard, and optionally, the user interface may further include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be referred to as a display screen or a display unit, and is used to display the information processed in the server 1 and to display a visual user interface.
[0034] Figure 1 Only the server 1 with components 11-13 and the application control program 10 that maps three-dimensional motion data to letter keys is shown. Those skilled in the art can understand that, Figure 1 The shown structure does not constitute a limitation on the server 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0035] In this embodiment, the processor 12 may be used to call the application control program stored in the memory 11 that maps three-dimensional motion data to letter keys, and perform the following operations:
[0036] Obtain the original multi-modal three-dimensional motion data of the user from the inertial measurement unit sensor and perform preprocessing to obtain standardized three-dimensional motion data, and perform motion state detection on the standardized three-dimensional motion data to obtain potential gesture action data;
[0037] Extract a feature set from the potential gesture action data, and perform weight adjustment on the feature set based on context information to obtain context-weighted feature data, and perform multi-stage dimensionality reduction processing on the context-weighted feature data to obtain a highly discriminative feature vector;
[0038] Input the highly discriminative feature vector into the classification system for processing, and through a multi-layer classification architecture, a personalized adaptive mechanism, and a feature stability protection algorithm, perform comprehensive analysis in combination with application scenario information to obtain a gesture category identifier and its confidence score;
[0039] Based on the gesture category identifier and its confidence score, generate an initial key mapping through multi-level mapping, and adjust the initial key mapping based on user preferences, context prediction, and application environment to obtain a final key event;
[0040] Perform multi-layer anti-mis-touch detection on the final key event to obtain a verified target key event;
[0041] Generate multimodal user feedback for the target key event to obtain a key event with a feedback signal, and apply the target key event with the feedback signal to the target application to achieve application control.
[0042] Based on the hardware of the application control device that maps three-dimensional motion data to letter keys as described above, referring to Figure 2 , Figure 2 This is an embodiment of the application control method of the present invention that maps three-dimensional motion data to letter keys. The application control method that maps three-dimensional motion data to letter keys includes the following steps:
[0043] S10. Obtain the multimodal three-dimensional motion raw data of the user from the inertial measurement unit sensor and perform preprocessing to obtain normalized three-dimensional motion data, and perform motion state detection on the normalized three-dimensional motion data to obtain potential gesture action data. The inertial measurement unit can capture data such as the acceleration, angular velocity, and direction of the user's hand movement. These data constitute the multimodal three-dimensional motion raw data, representing the user's gesture movement in three-dimensional space.
[0044] In some embodiments, step S10 can be implemented through steps S11 - S16:
[0045] In step S11, the high-frequency noise of the multimodal three-dimensional motion raw data is eliminated through a low-pass filter to obtain preliminarily filtered motion data. The system uses a low-pass filter to process the multimodal three-dimensional motion raw data obtained from the IMU sensor, aiming to filter out high-frequency noise. The low-pass filter allows low-frequency signals to pass through while suppressing high-frequency signals, which is very effective for eliminating high-frequency interferences such as minute hand tremors, electrical interference, and sensor noise.
[0046] In step S12, the preliminarily filtered motion data is subjected to multi-sensor data fusion through the Kalman filtering algorithm to obtain fused three-dimensional motion data. The Kalman filter is a recursive optimal estimation algorithm that can effectively process dynamic systems containing random noise. Through this algorithm, the system can complement the advantages of different sensors, reducing the limitations and drift problems of a single sensor.
[0047] In step S13, the fused three-dimensional motion data is reset for the sensor baseline through an automatic calibration algorithm to obtain calibrated three-dimensional motion data. The IMU sensor will experience baseline drift during long-term use, which will cause cumulative errors in the data. The automatic calibration algorithm identifies the reference benchmark point by detecting the stationary state or specific posture, and then resets the sensor baseline to eliminate the drift caused by temperature changes, mechanical stress, or electrical characteristic changes, ensuring data accuracy during long-term use.
[0048] In step S14, the calibrated three-dimensional motion data is segmented by the sliding window technique to obtain time segment data of a fixed length. The sliding window technique divides a continuous data stream according to a time window of a fixed length, and the windows can overlap to a certain extent. This method enables the system to process continuous data streams without the user explicitly marking the start and end of gestures. By setting an appropriate window size (such as 200 - 500 milliseconds) and sliding step, the system can capture complete gesture actions while avoiding processing excessive non-gesture data.
[0049] In step S15, the time segment data is subjected to Z-score normalization to obtain normalized three-dimensional motion data. Z-score normalization is a commonly used data normalization method that transforms the data into a distribution with a mean of 0 and a standard deviation of 1. The calculation formula is: Z = (X - μ) / σ, where X is the original value, μ is the mean of the data set, and σ is the standard deviation of the data set. Through this normalization process, the system can eliminate the differences in the ranges of different sensors and the differences in the amplitudes of different users' gestures, making the data comparable.
[0050] In step S16, the short-time energy and spectral characteristics of the normalized three-dimensional motion data are analyzed to distinguish gestures from random motions, obtaining potential gesture action data. Short-time energy refers to the change in signal energy within a short time window and can be used to detect the start and end of an action. Spectral characteristics are the signal frequency distribution characteristics obtained through methods such as Fourier transform, and different types of actions have different spectral characteristics. By setting an energy threshold and analyzing the spectral characteristic patterns, the system can effectively distinguish intentional gesture actions from daily random motions (such as walking, shaking while riding in a vehicle, etc.).
[0051] Through the above steps S11 - S16, a complete process from obtaining raw data from the IMU sensor to identifying potential gesture actions is achieved. The combined use of low-pass filtering and Kalman filtering effectively reduces the noise and interference in the data and significantly improves the signal quality. The automatic calibration algorithm solves the drift problem during the long-term use of the sensor, ensuring the long-term reliability of the data. The sliding window technique enables the system to process continuous data streams without the user explicitly marking the start and end of gestures, improving the practicality of the system. Z-score normalization makes the data from different users and different devices comparable, enhancing the generality of the system. The analysis of short-time energy and spectral characteristics can effectively distinguish meaningful gesture actions from random motions, reducing the misidentification rate.
[0052] S20. Extract a feature set from the potential gesture action data, adjust the weights of the feature set based on context information to obtain context-weighted feature data, and perform multi-stage dimensionality reduction on the context-weighted feature data to obtain a highly discriminative feature vector.
[0053] In some embodiments, step S20 may be implemented through steps S21 - S29:
[0054] In step S21, time-domain feature extraction is performed on the potential gesture action data through parallel computing to obtain time-domain features including mean, standard deviation, kurtosis, skewness, zero-crossing rate, waveform length, and signal amplitude area. Specifically, time-domain features are statistical and morphological features directly extracted from time-series data, which can intuitively reflect the basic characteristics of gesture actions. Parallel computing technology significantly improves the feature extraction efficiency by simultaneously processing multiple data axes and various feature calculations. The mean reflects the overall intensity level of the signal, and the calculation formula is μ = (1 / N)∑x_i, where N is the number of samples and x_i is the i-th sample value; the standard deviation measures the degree of signal fluctuation, and the calculation formula is σ = √[(1 / N)∑(x_i - μ) 2 ; kurtosis describes the sharpness of the signal distribution, and the calculation formula is Kurt = E[(x - μ) 4 / σ 4 ; skewness represents the asymmetry of the signal distribution, and the calculation formula is Skew = E[(x - μ) 3 / σ 3 ; the zero-crossing rate calculates the frequency at which the signal crosses the zero axis, reflecting the frequency characteristics of the signal, and the calculation formula is ZCR = (1 / (N - 1))∑sgn(x_i·x_{i - 1}), where sgn() is the sign function; the waveform length measures the signal complexity, and the calculation formula is WL = ∑|x_{i + 1} - x_i|; the signal amplitude area represents the cumulative intensity of the signal, and the calculation formula is MAV = (1 / N)∑|x_i|. These features together constitute a comprehensive description of the time-domain characteristics of gesture actions.
[0055] In step S22, frequency-domain feature extraction is performed on the potential gesture action data through fast Fourier transform to obtain frequency-domain features of the main frequency band energy distribution of the power spectral density. Frequency-domain features reflect the periodicity and frequency composition of gesture actions. The fast Fourier transform (FFT) is an efficient algorithm for calculating the discrete Fourier transform, which can convert a time-domain signal into a frequency-domain representation, and the calculation formula is X_k = ∑x_n·e^{-j2πkn / N}, where j is the imaginary unit, k is the frequency index, and n is the time index. Through the FFT conversion, the system calculates the power spectral density (PSD) of the signal, and the formula is PSD = |X_k| 2 / N, and then extracts the energy distribution characteristics of each frequency band from it. Typical frequency band divisions may include 0 - 2Hz (representing slow motion), 2 - 5Hz (medium motion), 5 - 10Hz (fast motion), etc.
[0056] In step S23, statistical feature extraction is performed on the potential gesture motion data by calculating the correlation matrix of each axis data and the principal component contribution rate, and statistical features reflecting the spatial distribution characteristics of the gesture are obtained. Statistical features focus on describing the spatial relationship and multi-dimensional distribution characteristics of the data, and are crucial for understanding the three-dimensional structure of the gesture. The correlation matrix calculates the Pearson correlation coefficient between each axis, and the formula is ρ_{xy} = cov(x,y) / (σ_x·σ_y), where cov(x,y) is the covariance of the x-axis and y-axis data, and σ_x and σ_y are the standard deviations of the x-axis and y-axis data respectively. These correlation coefficients reflect the cooperative motion patterns of the gesture in different spatial dimensions. The principal component contribution rate is obtained by performing eigenvalue decomposition on the correlation matrix, which reflects the degree of variation of the data along each principal component direction. For example, the writing gesture may have a large amount of variation in the xy plane and less variation in the z-axis direction, while the pointing gesture may have significant variations in all three dimensions.
[0057] In step S24, the time-domain features, the frequency-domain features, and the statistical features are merged to obtain a feature set. Feature merging is to combine features of different dimensions and different types into a unified feature vector. The merging process includes steps such as feature vector splicing, feature normalization, and redundancy detection. The splicing operation can be expressed as F_combined = [F_time; F_freq; F_stat], where F_time, F_freq, and F_stat are the time-domain, frequency-domain, and statistical feature vectors respectively. Since the dimensions and numerical ranges of different types of features may vary greatly, normalization processing such as Min-Max normalization or Z-score standardization is usually required before merging. In addition, the system may also perform preliminary redundancy detection to remove highly correlated features to reduce the computational complexity.
[0058] In step S25, the usage scenario and device location are detected for the user's current state through sensor data analysis to obtain context information. Context information refers to external factors such as the environmental state, activity type, and device wearing location of the user, which may significantly affect the characteristics of the gesture signal. The system identifies whether the user is stationary, walking, or in a vehicle by analyzing information such as the background acceleration pattern, posture change, and environmental noise; determines whether the user is indoors or outdoors; and determines whether the device is worn on the wrist, finger, or in the pocket, etc. For example, periodic vertical acceleration changes may indicate that the user is walking, and the calculation formula is step frequency = FFT peak frequency, and the typical step frequency range is 1.5 - 2.5 Hz; the average attitude angle of the device can be used to judge the wearing position, such as the attitude angle range is usually within ±45° when worn on the wrist.
[0059] In step S26, the feature weights of the feature set are dynamically adjusted according to the context information to obtain context-weighted feature data. Feature weight adjustment is a mechanism that adaptively changes the importance of different features according to specific scenarios, which can significantly improve the robustness of the system in complex environments. The adjustment process can be expressed as F_weighted = W⊙F_combined, where ⊙ represents element-level multiplication, and W is a weight vector that is dynamically generated according to context information. For example, in a walking scenario, since body movement introduces regular interference, the system may increase the weight of frequency domain features to better separate gesture signals from walking interference, and the formula may be W_freq = base_weight × (1 + walking_intensity); in an environment with insufficient lighting, due to the weakening of visual assistance, the system may increase the weight of statistical features to reduce dependence on the environment; when the device is loosely worn, the system may increase the weight of time domain features to capture more obvious motion features. This dynamic weight adjustment mechanism enables the system to maintain stable recognition performance in a changing environment.
[0060] In step S27, the context weighted feature data is subjected to the first stage of dimensionality reduction by principal component analysis to obtain the principal component features that retain 95% of the variance. Principal component analysis (PCA) is a commonly used unsupervised dimensionality reduction technique that maps the original features to a set of new, mutually orthogonal principal components through linear transformation, while retaining the variance information of the data as much as possible. PCA first calculates the feature covariance matrix C = (1 / N) ∑ (x_i-μ) (x_i-μ) ^ T, and then performs eigenvalue decomposition C = VΛV^T on the covariance matrix, where V is the eigenvector matrix and Λ is the eigenvalue diagonal matrix. The size of the eigenvalue reflects the variance contribution of the corresponding principal component. The system selects the first k principal components whose cumulative variance contribution rate reaches 95%, and the calculation formula is ∑_{i=1}^kλ_i / ∑_{i=1}^nλ_i≥0.95, where λ_i is the i-th largest eigenvalue and n is the total number of features. After selecting the principal components, the reduced-dimensional features are obtained through the projection transformation F_PCA=V_k^T×F_weighted. PCA dimensionality reduction not only reduces the feature dimension and computational complexity, but also eliminates the linear correlation between features, making subsequent classification more efficient.
[0061] In step S28, the second-stage dimensionality reduction of the principal component features is performed through linear discriminant analysis to obtain linear discriminant features that enhance class distinctiveness. Linear discriminant analysis (LDA) is a supervised dimensionality reduction technique aimed at finding the projection directions that can maximize the between-class variance while minimizing the within-class variance. Different from PCA which focuses on the overall variance, LDA focuses on improving class separability, and the formula is J(w) = (w^TS_Bw) / (w^TS_Ww), where S_B is the between-class scatter matrix and S_W is the within-class scatter matrix. S_B is calculated as S_B = ∑N_i(μ_i - μ)(μ_i - μ)^T, and S_W is calculated as S_W = ∑∑(x - μ_i)(x - μ_i)^T, where μ_i is the mean of the i-th class, μ is the global mean, and N_i is the number of samples in the i-th class. The system obtains the discriminant vector w by solving the generalized eigenvalue problem S_Bw = λS_Ww, and then projects the PCA features onto these vectors to obtain the LDA features F_LDA = W_LDA^T×F_PCA. The LDA dimensionality reduction further enhances the class distinctiveness of the features.
[0062] In step S29, the principal component features and the linear discriminant features are dynamically adjusted in weight ratio according to the confusion probabilities of different gesture types through an adaptive feature selection matrix to obtain high-distinctiveness feature vectors. Adaptive feature selection is a mechanism that intelligently combines PCA and LDA features and optimizes them specifically for gesture types that are prone to confusion. The system first constructs a gesture confusion matrix C, where C_{ij} represents the probability of misidentifying gesture i as gesture j. Then, for each pair of easily confused gestures (i, j), the system calculates the performance difference between the PCA features and the LDA features in distinguishing this pair of gestures. For example, the Fisher discriminant ratio F = |μ_i - μ_j|^2 / (σ_i^2 + σ_j^2) is used as an evaluation metric. Based on these evaluations, the system assigns the optimal weight combination of the PCA and LDA features to each gesture type. This adaptive feature selection strategy enables the system to adopt the optimal feature representation for different gesture types, significantly improving the overall recognition accuracy, especially for gestures with similar shapes.
[0063] Through the above steps S21 to S29, a systematic transformation from potential gesture action data to high - discriminative feature vectors is achieved. Multi - angle feature extraction (time domain, frequency domain, and statistical features) ensures a comprehensive description of gesture actions, capturing rich information of gestures in the time, frequency, and spatial dimensions. The context - awareness and dynamic weight adjustment mechanism enable the system to adapt to various complex environments and user states, significantly improving the robustness in actual usage scenarios. Multi - stage dimensionality reduction processing (PCA + LDA) not only reduces the computational complexity but also optimally reorganizes the features, enhancing the class - discrimination ability. The adaptive feature selection matrix is finely tuned for gesture types prone to confusion, further improving the system's ability to distinguish similar gestures.
[0064] S30. Input the high - discriminative feature vector into a classification system for processing. Through a multi - layer classification architecture, a personalized adaptive mechanism, and a feature stability protection algorithm, combined with application - scenario information for comprehensive analysis, obtain a gesture class identifier and its confidence score.
[0065] In some embodiments, step S30 can be completed through steps S31 - S39:
[0066] In step S31, the high - discriminative feature vector is classified at the first layer through the AdaBoost algorithm with decision stumps as weak learners and combined with a dynamic decay mechanism to obtain the first - layer classification result. AdaBoost (Adaptive Boosting) is an ensemble learning algorithm that iteratively trains a series of weak classifiers and combines them into a strong classifier. In this solution, decision stumps (single - layer decision trees) are selected as weak learners. This simple classifier makes decisions based on a single feature only, with fast calculation speed and low overfitting probability. This solution also introduces a dynamic decay mechanism, that is, the weight of a sample is adjusted according to its timestamp, making the most recent samples have a higher influence. The formula is w i = w i ·exp(-λ·Δt), where λ is the decay coefficient and Δt is the time interval of the sample.
[0067] In step S32, the high - discriminative feature vector is classified at the second layer through random forest and support vector machine to obtain the second - layer classification result. Random forest is an ensemble classifier composed of multiple decision trees, where each tree is independently trained and votes to determine the final classification result. Support Vector Machine (SVM) is an algorithm for finding the optimal separating hyperplane, especially good at dealing with high - dimensional feature spaces. For the case of linear inseparability, SVM uses kernel functions to map data into high - dimensional spaces. Commonly used kernel functions include polynomial kernel K(x,y)=(x·y + 1) d and radial basis function (RBF) kernel K(x,y)=exp(-γ||x - y|| 2)。In this solution, the random forest and SVM work in parallel, leveraging their different abilities to interpret the feature space: the random forest is good at capturing the non-linear interactions between features, while the SVM focuses on finding the key support vectors for the class boundaries.
[0068] In step S33, the highly discriminative feature vectors are classified at the third layer through a long short-term memory network to obtain the third-layer classification result. The long short-term memory (LSTM) network is a special type of recurrent neural network that can effectively learn the long-term dependencies in sequential data. In gesture recognition, the LSTM is particularly suitable for processing temporal features and can capture the temporal patterns and context dependencies in gesture actions.
[0069] In step S34, the first-layer classification result, the second-layer classification result, and the third-layer classification result are fused through a confidence-weighted voting mechanism to obtain the initial gesture classification result. The preferred solution for combining multiple classifiers uses a confidence-weighted voting mechanism, that is, different weights are assigned according to the confidence of each classifier in the current sample. The formula is C(x) = argmax∑w i ·I i (c|x), where w i is the weight of the i-th classifier, and I i (c|x) is an indicator function that is 1 when classifier i classifies sample x as c and 0 otherwise. The confidence can be calculated in various ways, such as the margin distance in AdaBoost, the decision function value in SVM, or the softmax output probability in LSTM. In addition, the system dynamically adjusts the weights of each classifier according to their performance on recent samples. The formula is where η is the learning rate, A i is the accuracy of classifier i, is the average accuracy of all classifiers.
[0070] In step S35, the initial gesture classification result is analyzed for user-specific error patterns through the user feedback collector to obtain user correction data. User feedback collection is a key step in achieving personalization. The system collects user feedback on the recognition results through multiple channels, including: explicit feedback (users actively correct misidentifications) and implicit feedback (analysis of user behavior patterns). The system records the detailed information of each recognition error, including the error type (position in the confusion matrix), feature vectors, context information, and user correction behavior. Through clustering analysis, the system identifies user-specific error patterns, such as the continuous confusion of certain gesture pairs (e.g., the "J" and "I" letter gestures) or the difficulty in recognition in specific scenarios (such as fine gestures while walking). These error patterns are quantified as an error frequency matrix E, where E(i,j) represents the frequency of misidentifying gesture i as gesture j. User correction data includes not only error pattern statistics but also the preferred gesture execution methods and usage frequency distributions of users.
[0071] In step S36, the initial gesture classification result and the user correction data are optimized for personalization through the model adaptation engine to obtain the gesture classification result adapted to the user. The model adaptation engine is a dynamic learning system that can continuously optimize the classification model according to user feedback. The specific implementation includes the following strategies: parameter fine-tuning, using user-specific data to fine-tune the pre-trained model, and the adjustment formula is where θ is the model parameter, α is the learning rate, and D_user is the user data; decision threshold adjustment, adjusting the decision boundary for gesture pairs that users are prone to confusing. If users often misidentify gesture A as B, then increase the confidence threshold required to classify a sample as A; sample weighting, increasing the weight of user error samples in training, and the formula is w(x) = 1 + β·error_count(x), where β is the weighting coefficient; personalized feature selection, re-evaluating the importance of features according to user data and highlighting the features that are helpful for distinguishing user-specific confusing gestures.
[0072] In step S37, the gesture classification result adapted to the user is protected for model update through the feature stability preservation algorithm to obtain the gesture classification result that prevents catastrophic forgetting. Catastrophic forgetting is a common problem in incremental learning, which refers to the situation where the model forgets what it has learned before when learning new knowledge. To prevent this problem, this solution adopts the feature stability preservation algorithm, and the core idea is to protect the key feature representations during the model update process. The specific implementation includes: knowledge distillation, using the output of the old model as the soft target of the new model, and the loss function is Among them, \(L_{CE}\) is the cross-entropy loss, \(L_{KL}\) is the KL divergence, and \(\lambda\) is the balance parameter; Elastic Weight Consolidation introduces a penalty term for model parameters to prevent important parameters from deviating from their original values, and the formula is \(L_{EWC}=L + (\lambda / 2)\cdot\sum F_i\cdot(\theta_i - \theta_i')\) 2 , where \(F_i\) is the parameter importance and \(\theta_i\) is the original parameter value; Memory Replay retains representative historical samples and mixes old and new samples when updating the model; Parameter Regularization restricts the model update amplitude by adding a regularization term to prevent overfitting to new samples.
[0073] In step S38, the application scenario is recognized for the user's current activity through sensor data and system state analysis to obtain application scenario information. The system infers the user's current activity state and usage environment through multiple information sources. Sensor data analysis includes: acceleration pattern recognition, such as judging whether the user is stationary, walking, or running by step frequency and gait characteristics, and the step frequency calculation formula is \(f_{step}=FFT_{peak}(a_{vertical})\); attitude estimation, calculating the spatial attitude of the device through the fusion algorithm of gyroscope and accelerometer to judge whether the user is standing, sitting, or lying; environmental feature extraction, such as judging whether the user is indoors or outdoors, in a quiet environment or a noisy environment by background noise level, light sensor, and GPS information. System state analysis includes: the application type of the current activity (text editing, gaming, browsing, etc.); device state (battery level, processor load, etc.); user interaction history (recent operation sequence, usage frequency, etc.). This information is integrated into a structured scenario description vector \(S\), which includes dimensions such as activity type, environmental conditions, and device state.
[0074] In step S39, the gesture classification results for preventing catastrophic forgetting are optimized for classification configuration according to the application scenario information through a context-aware classification strategy selector, obtaining a gesture category identifier and its confidence score. The classification strategy selector is a hybrid system based on rules and learning, capable of dynamically adjusting the classification strategy according to the scenario information. The specific implementation includes: classifier weight adjustment, optimizing the classifier fusion weights in step S34 according to the scenario, with the formula w_i = base_w_i · f(S, i), where f(S, i) is the fitness function of scenario S for classifier i; decision threshold adaptation, adjusting the confidence threshold for gesture recognition according to the scenario, increasing the threshold in scenarios requiring high precision (such as text editing) and decreasing the threshold in fault-tolerant scenarios (such as game control); feature importance adjustment, highlighting the role of specific features according to the scenario, such as increasing the weight of frequency domain features in a sports scenario; context constraint application, using application semantics to limit the possible gesture set, such as giving priority to letter gestures in a text editing application. Finally, the system outputs a gesture category identifier g and the corresponding confidence score c, with the formula (g, c) = argmax_g{P(g|x, S)}, where P(g|x, S) is the posterior probability of gesture g given the feature vector x and scenario S. The confidence score reflects the reliability of the recognition result and provides an important reference for subsequent processing.
[0075] Through the implementation of the above steps S31 to S39, an intelligent conversion process from high-distinguishability feature vectors to gesture category identifiers is achieved. The multi-layer classification architecture (AdaBoost, random forest, SVM, and LSTM) utilizes the complementary advantages of different algorithms to improve the overall recognition ability and robustness of the system. The personalized adaptive mechanism (user feedback collection and model adaptation engine) enables the system to learn the user's unique gesture patterns and error patterns, providing a tailored recognition experience. The feature stability protection algorithm solves the problem of catastrophic forgetting in incremental learning, ensuring that the system does not lose its general recognition ability while adapting to individual users. The context-aware classification strategy optimization intelligently adjusts the classification parameters according to different scenarios, enabling the system to maintain high performance in various complex environments.
[0076] S40. Based on the gesture category identifier and its confidence score, generate an initial key mapping through multi-level mapping, and adjust the initial key mapping based on user preferences, context prediction, and application environment to obtain the final key event.
[0077] In some embodiments, step S40 can be completed through steps S41 - S491:
[0078] In step S41, the gesture category identifier and its confidence score are subjected to a basic mapping of Latin letters through the base layer mapping to obtain the basic letter key mapping. The base layer mapping is the first layer of the entire mapping system and establishes the initial correspondence between gesture categories and Latin letters. This mapping is usually based on the principle of intuitive morphological similarity. For example, a gesture similar to "A" is mapped to the letter "A". The system maintains a base mapping table M_base, where M_base(g) = c indicates that gesture g is mapped to character c. For each recognized gesture category identifier g, the system queries the mapping table to obtain the corresponding letter while retaining the original confidence score.
[0079] In step S42, the gesture category identifier and its confidence score are subjected to letter certainty analysis through mapping confidence evaluation to obtain letter mapping confidence data. The mapping confidence evaluation is a further analysis of the reliability of the gesture recognition result, aiming to determine the certainty degree of the mapping result. The system not only considers the original confidence score of the gesture recognition but also combines multiple factors for comprehensive evaluation: the uniqueness index of the gesture category, calculated as U(g) = 1 - max_{g'≠g}(similarity(g,g')), which represents the complementary value of the maximum similarity of this gesture to other gestures; the historical accuracy rate of the user for this gesture, calculated as A(g,u) = correct_count(g,u) / total_count(g,u); the expected recognition accuracy under the current environmental conditions. For example, the recognition accuracy of some delicate gestures will decrease when walking. The system combines these factors into a mapping confidence score, with the formula C_map(g) = w1·C_recog(g) + w2·U(g) + w3·A(g,u) + w4·E(g,env), where C_recog(g) is the original recognition confidence, and w1 to w4 are weight coefficients. This score will be used for subsequent mapping validity verification to help the system distinguish between high-certainty and low-certainty mapping results.
[0080] In step S43, the basic letter key mapping is used to convert one-to-one gestures to letters through a mapping matrix, obtaining an initial letter mapping result. The mapping matrix is a flexible gesture-to-letter conversion mechanism that allows the system to dynamically adjust the mapping relationship according to different scenarios and user needs. The mapping matrix M can be represented as an n×m matrix, where n is the number of gesture categories and m is the number of target characters, and M[i,j] represents the probability or weight of gesture i mapping to character j. In the initial state, the mapping matrix may be an identity matrix (one-to-one mapping), but the system will gradually adjust the matrix values according to user feedback and usage patterns. The mapping process can be expressed as c = argmax_j(M[g*,j]), that is, select the character c with the highest mapping probability to the recognized gesture g*. This matrix representation method enables the system to implement complex mapping strategies, such as one-to-many mapping (one gesture may map to multiple characters, selected according to the context) or many-to-one mapping (multiple similar gestures map to the same character, improving fault tolerance). The initial letter mapping result contains the target letter and its mapping probability distribution, providing a basis for subsequent processing.
[0081] In step S44, the initial letter mapping result and letter mapping confidence data are verified for mapping validity through confidence threshold judgment, obtaining a verified letter mapping. Confidence threshold judgment is a quality control mechanism used to filter out mapping results with low confidence and prevent incorrect input. The system sets a basic confidence threshold θ_base (such as 0.7), and only results with a mapping confidence exceeding this threshold are considered valid mappings. In addition, the system also adopts an adaptive threshold strategy to dynamically adjust the threshold according to different scenarios and user states. The formula is θ_adaptive = θ_base + Δθ(context), where Δθ(context) is the threshold adjustment amount based on the context. For example, in an exact input scenario (such as password input), the system may increase the threshold to 0.85; while in a fault-tolerant scenario (such as daily chat), it may decrease the threshold to 0.65. For mapping results that do not pass the threshold verification, the system may adopt different strategies: completely reject the input; mark it as a low-certainty result and wait for user confirmation; or trigger a secondary recognition process and require the user to repeat the gesture. Through this verification mechanism, the system can provide a good user experience while maintaining high input accuracy.
[0082] In step S45, the letter usage frequencies of the user's historical input data are calculated through statistical analysis to obtain the user's letter usage frequency data. The letter usage frequency is an important basis for personalized mapping. The system constructs a personalized letter frequency model by analyzing the user's historical input behavior. The statistical analysis process includes: collecting the text input history of the user in different applications; calculating the occurrence frequency of each letter, with the formula f(c) = count(c) / total_chars; analyzing the context correlation of letters, such as bigram frequency f(c1,c2) and trigram frequency f(c1,c2,c3); identifying the high-frequency words and phrase patterns unique to the user. The system also considers the time factor and gives higher weights to recent inputs, with the formula f_weighted(c) = ∑(w_t·count_t(c)) / ∑(w_t·total_chars_t), where w_t is the time weight, which decays exponentially over time. These frequency data are organized into a user letter usage frequency model, including a single-letter frequency table, an n-gram frequency table, and a list of high-frequency words, providing a data basis for subsequent frequency-adaptive remapping.
[0083] In step S46, the correspondence between gestures and letters is adjusted for the verified letter mapping through a frequency-adaptive remapping mechanism according to language usage habits and the user's letter usage frequency data to obtain a frequency-optimized letter key mapping. Frequency-adaptive remapping is a mechanism to optimize the user input efficiency, and the core idea is to make high-frequency letters correspond to more easily executed gestures. The system first evaluates the execution difficulty of each gesture, considering factors such as: action complexity (such as the number of joint movements required); execution time; error rate; user subjective rating, etc. Then, the system calculates the optimal mapping relationship based on the letter frequency and gesture difficulty, which can be modeled as an assignment problem: min∑D(g_i)·f(c_j)·x_{ij}, where D(g_i) is the difficulty coefficient of gesture i, f(c_j) is the usage frequency of letter j, and x_{ij} is an indicator variable, which is 1 when gesture i is mapped to letter j and 0 otherwise. By solving this optimization problem, the system obtains a frequency-optimized mapping relationship, making high-frequency letters correspond to simple gestures and low-frequency letters correspond to complex gestures. To avoid user confusion caused by frequent changes, the system usually sets a stabilization period and a change amplitude limit, and only makes adjustments when the new mapping is significantly superior to the current mapping (such as the expected efficiency improvement exceeds 20%).
[0084] In step S47, the next letter prediction is performed on the input text and the application context through semantic analysis to obtain the letter prediction result. The next letter prediction is a context-based intelligent assistance function that predicts the next character most likely to be input by the user by analyzing the input text and the current application scenario. The semantic analysis process includes: the n-gram model, which predicts the next character based on the conditional probability P(c_n|c_{n - 1},c_{n - 2},...,c_{n - k + 1}), such as using the trigram model P(c_3|c_2,c_1) to calculate the probability distribution of the third character given the first two characters; the lexical completion model, which, when detecting that the user is inputting the beginning of a word, searches the dictionary to predict the complete word, such as predicting "elephant", "electron", etc. when the user inputs "ele"; the semantic understanding model, which analyzes the semantic content of the text based on deep learning language models (such as variants of BERT, GPT, etc.) to predict the next word that conforms to the context. The system synthesizes the outputs of these models to generate the probability distribution P(c_next) of the next letter, taking into account the influence of the application context. For example, in a search box, it may be biased towards keywords, and in a chat application, it may be biased towards colloquial expressions.
[0085] In step S48, the frequency-optimized letter key mapping is enhanced according to the letter prediction result through the context-sensitive semantic prediction function to obtain the letter key mapping with prediction suggestions. Context-sensitive semantic prediction is a mechanism that integrates the letter prediction result into the mapping process, enabling the system to intelligently adjust the mapping priority. The specific implementation includes: prediction probability fusion, which combines the letter prediction probability with the basic mapping probability, and the formula is P_combined(c|g) = α·P_map(c|g)+(1 - α)·P_pred(c), where α is a balance parameter that controls the relative importance of the basic mapping and the prediction result; threshold dynamic adjustment, which adjusts the letter acceptance threshold according to the prediction confidence, reducing the acceptance threshold for letters with high prediction probabilities, and the formula is θ_c = θ_base - β·P_pred(c), where β is an adjustment coefficient; candidate list reordering, which sorts the possible letter candidates according to the combined probability and preferentially displays the letters with high probabilities; prediction feedback enhancement, which indicates the high-probability predictions to the user through visual or tactile cues, such as highlighting the predicted letters. These mechanisms work together to enable the system to provide intelligent input suggestions while maintaining user control, reducing unnecessary gesture inputs, and improving the overall input efficiency.
[0086] In step S49, the current active application is identified for its type through the system interface to obtain application environment information. The application environment information refers to the type, status, and functional context of the currently running application, and these information are crucial for optimizing the key mapping strategy. The system obtains application information through the operating system API or the accessibility interface, including: application category (such as text editor, browser, game, etc.); current focus control type (such as text box, search box, command line, etc.); application status (such as edit mode, browse mode, full screen mode, etc.); specific functional context (such as formula input in document editing, specific language syntax in code editing, etc.). The system structures these information into an application environment description vector E, which includes dimensions such as application type, control type, status flag, etc. In addition, the system also maintains an application feature database, which records the special requirements and best practices of different applications. For example, some applications may require specific shortcut key combinations or command sequences. These environment information provide a specific application context for subsequent mapping strategy adjustment.
[0087] In step S491, the letter key mapping with prediction suggestions is adjusted for the mapping strategy according to the application environment information through the application-aware mapping switching function to obtain the final key event. The application-aware mapping switching is the last step of the entire mapping process, which finely adjusts the mapping strategy according to the specific application environment. The specific implementation includes: mapping mode switching, automatically switching different mapping modes according to the application type. For example, use the letter mapping mode in the text editor and the control command mapping mode in the media player; key combination generation, mapping a single gesture to a complex key combination or command sequence, such as mapping a specific gesture to the "Ctrl+C" copy command; application-specific optimization, adjusting the mapping parameters according to the application characteristics. For example, increasing the mapping threshold in the precision operation application and optimizing the response speed in the game application; functional context adaptation, adjusting the mapping behavior according to the specific functional context within the application. For example, optimizing the number input in the table editing state and optimizing the symbol input in the code editing state. The system generates the final key event through these adjustments, including key type (such as character key, function key, combination key, etc.), key value, modifier status, and target control information. These key events will be sent to the target application to achieve the final execution of the user's intention.
[0088] Through the implementation of the above steps S41 to S491, an intelligent conversion process from gesture category identifiers to final key events is achieved. The basic layer mapping and mapping matrix establish the basic correspondence between gestures and letters, providing a stable foundation for the entire mapping system. The mapping confidence evaluation and threshold judgment mechanism ensure the accuracy and reliability of the input, effectively filtering out low-confidence recognition results. The frequency adaptive remapping optimizes the correspondence between gestures and letters according to the user's language usage habits, improving the input efficiency of frequently used characters. The context-sensitive semantic prediction function utilizes the powerful capabilities of the language model to predict the possible input content of the user, providing intelligent assistance. The application-aware mapping switching dynamically adjusts the mapping strategy according to the characteristics and requirements of different applications, enabling the system to seamlessly adapt to various application scenarios.
[0089] S50. Perform multi-layer anti-misoperation detection on the final key event to obtain a verified target key event. Anti-misoperation detection is a key link to ensure the reliability of the system. Through a multi-level verification mechanism, it effectively filters out accidentally triggered key events and improves the user experience. This step adopts a three-layer verification architecture from gesture intention, context environment to semantic rationality to comprehensively evaluate the effectiveness of the key event.
[0090] In some embodiments, step S50 can be completed through steps S51 - S56:
[0091] In step S51, for the gesture that generates the final key event, the starting stability, execution fluency, and termination characteristics of the gesture are analyzed through an intent confirmation algorithm to obtain an intent confirmation score. The intent confirmation algorithm is an analysis method specifically for evaluating whether a user's gesture is executed consciously. By checking the integrity and quality characteristics of the gesture, it distinguishes intentional input from unconscious actions. The starting stability refers to the preparatory state before the gesture starts. The system detects the short stationary period (usually 100 - 200 ms) before the gesture and calculates the standard deviation σ_pre of the acceleration and angular velocity. The stability score S_start = exp(-k·σ_pre), where k is a scaling factor. The execution fluency refers to the coherence and smoothness during the gesture. The system calculates indicators such as the curvature change rate, speed distribution, and acceleration continuity of the gesture trajectory. The fluency score S_flow = w1·smoothness + w2·continuity + w3·rhythm, where w1, w2, and w3 are weight coefficients. The termination characteristic refers to a specific pattern at the end of the gesture, such as an obvious deceleration or a short pause. The system detects the speed change pattern and posture stability at the end of the gesture. The termination score S_end = sigmoid(α·(v_final - v_threshold)), where v_final is the final speed, v_threshold is the threshold speed, and α is a scaling parameter. The analysis results of these three aspects are synthesized into an intent confirmation score, and the formula is S_intent = w_start·S_start + w_flow·S_flow + w_end·S_end, where w_start, w_flow, and w_end are the weight coefficients of each part, and the sum is 1.
[0092] In step S52, the first - layer false - touch filtering is performed on the intent confirmation score through threshold judgment to obtain the key - press event after intent confirmation. Threshold judgment is a basic but effective filtering mechanism. The system sets an intent confirmation threshold θ_intent (such as 0.75), and only gestures with scores exceeding this threshold are regarded as valid intents. The system adopts an adaptive threshold strategy, dynamically adjusting the threshold according to different scenarios and user states. The formula is θ_adaptive = θ_base+Δθ(context), where Δθ(context) is the threshold adjustment amount based on the context. For example, in scenarios with high - precision requirements (such as financial applications), the system may increase the threshold to 0.85; while in casual scenarios (such as games), it may reduce the threshold to 0.65. For gestures that do not pass the threshold verification, the system directly discards the corresponding key - press events to avoid misoperations. For borderline cases close to the threshold, the system may take special handling, such as temporarily retaining but marking them as low - certainty and waiting for more evidence from the subsequent verification layer. Through this initial intent - based filtering, the system can effectively intercept false - touch events caused by unconscious actions, such as arm swings while walking or daily hand movements.
[0093] In step S53, context extraction is performed on the current application state and user activity scenario through sensor data and system - state analysis to obtain context information. Context extraction is a key step in understanding the user's current state and environment, providing a basis for subsequent conflict detection. The system collects context data through multiple sensors and system interfaces: analyzing the acceleration sensor and gyroscope data to identify the user's activity state, such as stationary, walking, running, or riding in a vehicle. Activity recognition can use feature - extraction and classification algorithms, such as decision trees or support vector machines; device - pose estimation determines the spatial position and orientation of the device, using sensor - fusion algorithms such as complementary filtering or Kalman filtering; environmental perception detects environmental characteristics, such as noise level, light conditions, and geographical location, through microphones, light sensors, and GPS, etc.; system - state monitoring collects device and application information, such as battery level, processor load, the currently active application, and screen state. These data are integrated into a structured context vector C, which includes dimensions such as activity type, environmental conditions, and device state. The system also maintains a context - history buffer to record recent context changes for state - transition detection and trend analysis.
[0094] In step S54, the key event after intention confirmation is subjected to a second - layer accidental touch filtering by the context conflict detector based on context information to obtain a key event verified by context. Context conflict detection is a verification mechanism based on environment and activity compatibility, aiming to identify key events inconsistent with the current context. The system maintains a context compatibility model that defines the likelihood scores of key events in different contexts, with the formula P(key|context) = f(key, context), representing the probability of a specific key event given a context. Compatibility evaluation considers multiple factors: activity compatibility, such as the low compatibility of complex and delicate gestures in the walking state; application - state compatibility, such as restrictions on certain operations in the locked - screen state; time - mode compatibility, such as the abnormality of consecutive identical operations within a short period; and spatial - location compatibility, such as the low likelihood of gesture input when the device is in the pocket. For detected conflict events, the system adopts different strategies according to the degree of conflict: severe conflicts are directly rejected; moderate conflicts require additional confirmation; and minor conflicts are marked as warnings but allowed to pass. Through this context - aware filtering mechanism, the system can identify abnormal operations that, although the gesture is executed well, do not match the current scenario, such as the situation where a user suddenly enters a complex password while running.
[0095] In step S55, the historical key sequence is evaluated for semantic rationality through language model analysis to obtain semantic verification rules. Semantic rationality evaluation uses natural language processing techniques to analyze the semantic coherence of the key sequence, providing a rule basis for subsequent temporal verification. The system uses a multi - level language model for analysis: the character - level n - gram model calculates the probability of a character sequence, such as P(c_n|c_{n - 1},c_{n - 2},...,c_{n - k+1}), for detecting uncommon character combinations; the lexical - level analysis checks whether the formed words exist in the dictionary or conform to word - formation rules; the syntactic analysis evaluates the rationality of the sentence structure, such as the probability of a verb being followed by a noun being higher than by an adverb; and the semantic coherence analysis checks the topic consistency and logical relationships of the text. The system combines these analysis results to generate a set of semantic verification rules, including: an allowed character conversion probability matrix T, where T[i,j] represents the probability of character i being followed by character j; a list W of common words and phrases for quick verification; and a set A of abnormal patterns that define specific sequence patterns that may indicate errors, such as repeating the same character consecutively multiple times. These rules are organized as a decision tree or a rule set for subsequent temporal logic verification and are continuously updated and optimized based on user feedback and new data.
[0096] In step S56, the key events after context verification are subjected to third-layer accidental touch filtering by a timing logic validator to analyze the semantic rationality of consecutive key events according to semantic verification rules, and the verified target key events are obtained. Timing logic verification is the last line of defense for the accidental touch prevention system. By analyzing the time pattern and semantic relationship of the key sequence, unreasonable input sequences are identified. The verification process includes multiple aspects: time interval analysis, checking whether the time distribution of key events conforms to the normal input pattern, and abnormally fast or slow input may indicate accidental touch; sequence probability evaluation, using the language model in step S55 to calculate the probability score P_seq of the current key sequence. If P_seq is lower than the threshold θ_seq, it is marked as a suspicious sequence; pattern matching, matching the current sequence with the predefined set of abnormal patterns A to detect known error patterns, such as a keyboard row sequence like "qwerty"; context coherence check, ensuring that the key sequence is consistent with the current task and application context, such as entering text in the URL format in the search box. For the detected abnormal sequences, the system may adopt different correction strategies: directly rejecting the latest key event; prompting the user to confirm the abnormal input; automatically replacing it with the most likely correct sequence; or triggering the automatic error correction function of the input method. Through this advanced verification based on semantics and timing, the system can capture complex accidental touch situations that may be missed by the first two layers of filtering, such as input sequences that are semantically unreasonable but have correct gesture execution.
[0097] It can be understood that this multi-level and multi-angle accidental touch detection mechanism significantly improves the reliability of the system and the user experience, enabling the IMU-based three-dimensional gesture input technology to work stably in various complex environments and usage scenarios, providing users with an accurate and reliable input experience, while minimizing the troubles and frustrations caused by misoperations.
[0098] S60. Generate multi-modal user feedback for the target key event to obtain a key event with a feedback signal, and apply the target key event with the feedback signal to the target application program to achieve application program control.
[0099] In some embodiments, step S60 can be completed through steps S61 - S66:
[0100] In step S61, a visual feedback is generated for the target key event through a dynamic indicator on the screen, displaying the current recognition status, candidate letters, and confidence level, to obtain a key event with visual feedback. Visual feedback is the most direct user interaction method, and the system displays the status information of gesture recognition and key mapping on the screen in real time. The dynamic indicator includes multiple visual elements: the status icon shows the current working mode of the system (such as recognizing, confirmed, rejected, etc.), and different colors and animation effects are used to represent different states; the candidate letter area shows the recognition results and alternatives, with the main option displayed prominently and the alternatives sorted by confidence level; the confidence indicator bar intuitively shows the reliability of the recognition result, for example, green is used to indicate high confidence (>0.8), yellow is used to indicate medium confidence (0.6 - 0.8), and red is used to indicate low confidence (<0.6); the visualization of the gesture trajectory maps the user's three-dimensional gesture actions onto the two-dimensional screen to help the user understand the recognition process of the system.
[0101] In step S62, an auditory feedback is generated for the key event with visual feedback through a beep of different tones, corresponding to different types of keys and system states, to obtain a key event with audio-visual feedback. Auditory feedback enhances the user experience through sound cues and is especially suitable for scenarios where visual attention is limited. The system has designed a set of sound languages: the successful recognition tone uses a short rising tone to indicate that the key event has been successfully accepted; the error prompt tone uses a short falling tone to indicate that the input has been rejected or there is an error; the warning prompt tone uses a repeated rhythm tone to indicate special situations that require the user's attention; the status transition tone uses a specific sound effect to indicate a system mode switch, such as switching from the letter mode to the number mode. The sound design takes into account multiple factors: short duration (usually <200ms) to avoid interference; moderate volume and automatic adjustment according to environmental noise; obvious timbre differences to ensure distinguishability; and a moderate frequency range (usually between 500 - 4000Hz) to ensure that most users can hear clearly. The auditory feedback works in coordination with the visual feedback to provide multi-channel operation confirmation for users.
[0102] In step S63, for the key event with audio-visual feedback, tactile feedback is generated through the vibration motor of the wearable device. Different modes of vibration are associated with specific input behaviors to obtain a key event with a multi-modal feedback signal. Tactile feedback is a private and environment-independent feedback method, which is particularly suitable for use in public places or noisy environments. The system generates fine tactile patterns through the vibration motor of the wearable device (such as a smart bracelet, glove or ring): a single short vibration (about 50 ms) indicates that the key event is accepted; a double short vibration indicates that the input is rejected; a long vibration (about 300 ms) indicates a special situation that requires user confirmation; a gradual change in vibration (gradual change in intensity) indicates the progress of continuous operation. The tactile pattern design takes into account the characteristics of human tactile perception. For example, the vibration intensity is moderate to avoid discomfort, there are sufficient differences between vibration patterns to ensure distinguishability, and the vibration duration and interval are optimized to improve the recognition rate. The system also supports the personalized customization of tactile patterns, and users can adjust the vibration intensity, duration and pattern according to their personal preferences. Tactile feedback, together with visual and auditory feedback, constitutes a complete multi-modal feedback system, providing users with a comprehensive operation perception.
[0103] In step S64, for the key event with a multi-modal feedback signal, system integration is performed through a layered architecture interface to obtain an integratable key event. The layered architecture interface is a software design pattern that organizes the complex key event processing logic into multiple abstract levels to ensure the modularity and scalability of the system. The interface architecture includes multiple levels: the hardware abstraction layer is responsible for communicating with specific input and feedback devices, hiding hardware details; the event processing layer is responsible for the standardized processing of key events, including event type recognition, parameter extraction and priority management; the application adaptation layer is responsible for converting standardized events into a format recognizable by specific applications and handling compatibility issues for different applications; the feedback coordination layer is responsible for managing the timing and coordination of multi-modal feedback to ensure the consistency of various feedback methods. This layered design enables the system to adapt to different hardware platforms and application environments while maintaining the consistency of core functions. Through standardized event formats and clear interface definitions, the system can be easily integrated into various operating systems and application frameworks.
[0104] In step S65, for the integratable key events, the integration mode is selected through the analysis of the current application type to obtain the key events adapted to the mode. The integration mode selection is to choose the most suitable event transmission and processing method according to the characteristics and requirements of the target application. The system selects different integration strategies by analyzing the application type and status: The direct injection mode is applicable to the standard text input scenario, and the system directly injects character events into the input focus control; The command mapping mode is applicable to the application control scenario, and the system maps gestures to application-specific commands or shortcut key combinations; The accessibility mode is applicable to the accessibility scenario, and the system interacts with the application through the accessibility API of the operating system; The custom protocol mode is applicable to deeply integrated applications, and the system directly communicates with the application through a predefined protocol. The system maintains an application database to record the best integration mode and special requirements of common applications. For example, some games may require specific key sequences or timing requirements. By selecting the most suitable integration mode, the system can ensure that key events can be effectively processed in different applications.
[0105] In step S66, the key events adapted to the mode are transmitted through the application programming interface, and the key events adapted to the mode are applied to the target application to achieve application control. The application programming interface (API) is the bridge for the system to communicate with the target application, responsible for accurately transmitting the key events to the application and ensuring correct execution. The interface implementation includes multiple mechanisms: Operating system-level APIs, such as the Send Input function in Windows or the Input Manager in Android, are used to simulate standard keyboard input; Accessibility APIs, such as UI Accessibility in iOS or AccessibilityService in Android, are used to control applications that do not support direct input; Application-specific APIs are deeply integrated with specific applications through SDKs or plugins; Network protocols communicate with web applications through protocols such as WebSocket or HTTP. The system performs a final verification before transmitting the event to ensure that the event format is correct, the target application is in a receivable state, and additional control information, such as event timestamps and source identifiers, is added if necessary. Through this step, the user's gesture input is finally converted into the actual operation of the application, completing the entire input process.
[0106] Through the implementation of the above steps S61 to S66, a complete closed-loop from the target key event to the application control is achieved. The multi-modal feedback system (visual, auditory, tactile) provides users with rich and intuitive operation confirmation, significantly reducing input uncertainty and operation anxiety. The hierarchical architecture interface and flexible integration mode enable the system to seamlessly adapt to various application scenarios and platform environments. It can be understood that this comprehensive feedback and intelligent integration mechanism not only improves the usability and user experience of gesture input, but also expands the application scope of 3D gesture input technology, enabling it to play a role in various application scenarios and providing users with a more natural and efficient human-computer interaction method.
[0107] In addition, the embodiment of the present invention also proposes an application control device for mapping 3D motion data to alphabet keys, and the application control device for mapping 3D motion data to alphabet keys includes:
[0108] A data acquisition and preprocessing module, configured to acquire the original multi-modal 3D motion data of the user from an inertial measurement unit sensor and perform preprocessing to obtain standardized 3D motion data, and perform motion state detection on the standardized 3D motion data to obtain potential gesture action data; a feature extraction and dimensionality reduction module, connected to the data acquisition and preprocessing module, configured to extract a feature set from the potential gesture action data, adjust the weights of the feature set based on context information to obtain context-weighted feature data, and perform multi-stage dimensionality reduction processing on the context-weighted feature data to obtain a highly discriminative feature vector; a gesture recognition and classification module, connected to the feature extraction and dimensionality reduction module, configured to input the highly discriminative feature vector into a classification system for processing, and perform comprehensive analysis in combination with application scenario information through a multi-layer classification architecture, a personalized adaptive mechanism, and a feature stability protection algorithm to obtain a gesture category identifier and its confidence score; a key mapping module, connected to the gesture recognition and classification module, configured to generate an initial key mapping through multi-level mapping based on the gesture category identifier and its confidence score, and adjust the initial key mapping based on user preferences, context prediction, and application environment to obtain a final key event; a mis-touch prevention module, connected to the key mapping module, configured to perform multi-layer mis-touch detection on the final key event to obtain a verified target key event; a feedback and application control module, connected to the mis-touch prevention module, configured to generate multi-modal user feedback for the target key event to obtain a key event with a feedback signal, and apply the target key event with the feedback signal to a target application program to achieve application control. Among them, the steps implemented by each functional module of the application control device for mapping 3D motion data to alphabet keys can refer to the various embodiments of the application control method for mapping 3D motion data to alphabet keys of the present invention, which will not be elaborated here.
[0109] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which may be any one or any combination of a hard disk, a multimedia card, an SD card, a flash memory card, an SMC, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, etc. The computer-readable storage medium includes an application control program 10 that maps three-dimensional motion data to letter keys. The specific implementation of the computer-readable storage medium of the present invention is substantially the same as that of the above-mentioned application control method for mapping three-dimensional motion data to letter keys and the specific implementation of the server 1, and will not be elaborated here.
[0110] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. An application control method for mapping three-dimensional motion data to alphabet keys, characterized in that, Including: Obtain the raw multi-modal three-dimensional motion data of the user from the inertial measurement unit sensor and perform preprocessing to obtain standardized three-dimensional motion data, detect the motion state of the standardized three-dimensional motion data, and obtain potential gesture action data; Extract a feature set from the potential gesture action data, adjust the weights of the feature set based on context information to obtain context-weighted feature data, and perform multi-stage dimensionality reduction processing on the context-weighted feature data to obtain a highly discriminative feature vector; Input the highly discriminative feature vector into the classification system for processing, and through a multi-layer classification architecture, a personalized adaptive mechanism, and a feature stability protection algorithm, combine with application scenario information for comprehensive analysis to obtain a gesture category identifier and its confidence score; Based on the gesture category identifier and its confidence score, generate an initial key mapping through a multi-level mapping, and adjust the initial key mapping based on user preferences, context prediction, and application environment to obtain a final key event; Perform multi-layer anti-misoperation detection on the final key event to obtain a verified target key event; Generate multi-modal user feedback for the target key event to obtain a key event with a feedback signal, and apply the target key event with the feedback signal to the target application program to achieve application program control.
2. The application control method for mapping three-dimensional motion data to alphabet keys as described in claim 1, characterized in that Obtain the raw multi-modal three-dimensional motion data of the user from the inertial measurement unit sensor and perform preprocessing to obtain standardized three-dimensional motion data, detect the motion state of the standardized three-dimensional motion data, and obtain potential gesture action data, including: Eliminate high-frequency noise from the raw multi-modal three-dimensional motion data through a low-pass filter to obtain preliminarily filtered motion data; Perform multi-sensor data fusion on the preliminarily filtered motion data through the Kalman filtering algorithm to obtain fused three-dimensional motion data; Reset the sensor baseline of the fused three-dimensional motion data through an automatic calibration algorithm to obtain calibrated three-dimensional motion data; Segment the calibrated three-dimensional motion data through a sliding window technique to obtain time segment data of a fixed length; Perform Z-score normalization processing on the time segment data to obtain standardized three-dimensional motion data; Analyze the short-time energy and spectral characteristics of the standardized three-dimensional motion data to distinguish gestures from random motions and obtain potential gesture action data.
3. The application control method for mapping three-dimensional motion data to alphabet keys as claimed in claim 1, wherein Extract a feature set from the potential gesture action data, adjust the weights of the feature set based on context information to obtain context-weighted feature data, and perform multi-stage dimensionality reduction processing on the context-weighted feature data to obtain a highly discriminative feature vector, including: Extract time-domain features from the potential gesture action data through parallel computing to obtain time-domain features including mean, standard deviation, kurtosis, skewness, zero-crossing rate, waveform length, and signal amplitude area; Extract frequency-domain features from the potential gesture action data through fast Fourier transform to obtain frequency-domain features of the main frequency band energy distribution of the power spectral density; Statistical feature extraction is performed on the potential gesture motion data by calculating the correlation matrix and the principal component contribution rate of each axis data to obtain statistical features reflecting the spatial distribution characteristics of the gesture; The time-domain features, the frequency-domain features and the statistical features are combined to obtain a feature set; The user's current state is detected for the usage scenario and device location through sensor data analysis to obtain context information; The feature weights of the feature set are dynamically adjusted through the context information to obtain context-weighted feature data; The context-weighted feature data is subjected to a first-stage dimensionality reduction through principal component analysis to obtain principal component features that retain 95% of the explained variance; The principal component features are subjected to a second-stage dimensionality reduction through linear discriminant analysis to obtain linear discriminant features that enhance class distinctiveness; The principal component features and the linear discriminant features are dynamically adjusted in weight ratio according to the confusion probability of different gesture types through an adaptive feature selection matrix to obtain a highly discriminative feature vector; 4. The application control method for mapping three-dimensional motion data to alphabet keys as described in claim 1, characterized in that, The highly discriminative feature vector is input into a classification system for processing, and through a multi-layer classification architecture, a personalized adaptive mechanism and a feature stability protection algorithm, combined with application scenario information for comprehensive analysis, to obtain a gesture class identifier and its confidence score, including: The first layer of classification is performed on the highly discriminative feature vector through the AdaBoost algorithm with decision stumps as weak learners and combined with a dynamic decay mechanism to obtain a first-layer classification result; The second layer of classification is performed on the highly discriminative feature vector through random forests and support vector machines to obtain a second-layer classification result; The third layer of classification is performed on the highly discriminative feature vector through a long short-term memory network to obtain a third-layer classification result; The first-layer classification result, the second-layer classification result and the third-layer classification result are fused through a confidence-weighted voting mechanism to obtain an initial gesture classification result; The initial gesture classification result is analyzed for the user's specific error patterns through a user feedback collector to obtain user correction data; The initial gesture classification result and the user correction data are subjected to personalized optimization through a model adaptation engine to obtain a gesture classification result adapted to the user; The gesture classification result adapted to the user is protected for model update through a feature stability preservation algorithm to obtain a gesture classification result that prevents catastrophic forgetting; The user's current activity is recognized for the application scenario through sensor data and system state analysis to obtain application scenario information; The gesture classification result that prevents catastrophic forgetting is optimized for classification configuration through a context-aware classification strategy selector according to the application scenario information to obtain a gesture class identifier and its confidence score.
5. The application control method for mapping three-dimensional motion data to alphabet keys as claimed in claim 1, characterized in that, Based on the gesture class identifier and its confidence score, an initial key mapping is generated through a multi-level mapping, and the initial key mapping is adjusted based on user preferences, context prediction and application environment to obtain a final key event, including: The basic mapping of Latin letters is performed on the gesture class identifier and its confidence score through the basic layer mapping to obtain a basic letter key mapping; Perform alphabet certainty analysis on the gesture category identifier and its confidence score through mapping confidence evaluation to obtain alphabet mapping confidence data; Perform one-to-one gesture-to-alphabet conversion on the basic alphabet key mapping through a mapping matrix to obtain an initial alphabet mapping result; Perform mapping validity verification on the initial alphabet mapping result and the alphabet mapping confidence data through confidence threshold judgment to obtain the verified alphabet mapping; Calculate the alphabet usage frequency of the user's historical input data through statistical analysis to obtain user alphabet usage frequency data; Adjust the correspondence between gestures and alphabets for the verified alphabet mapping through a frequency adaptive remapping mechanism according to language usage habits and the user alphabet usage frequency data to obtain a frequency-optimized alphabet key mapping; Predict the next alphabet for the input text and application context through semantic analysis to obtain an alphabet prediction result; Enhance the frequency-optimized alphabet key mapping according to the alphabet prediction result through a context-sensitive semantic prediction function to obtain an alphabet key mapping with prediction suggestions; Identify the type of the current active application program through a system interface to obtain application environment information; Adjust the mapping strategy for the alphabet key mapping with prediction suggestions according to the application environment information through an application-aware mapping switching function to obtain a final key event; 6. The application control method for mapping three-dimensional motion data to alphabet keys as described in claim 1, characterized in that Perform multi-layer anti-misoperation detection on the final key event to obtain a verified target key event, including: Analyze the starting stability, execution fluency, and termination characteristics of the gesture that generates the final key event through an intention confirmation algorithm to obtain an intention confirmation score; Perform the first layer of misoperation filtering on the intention confirmation score through threshold judgment to obtain a key event after intention confirmation; Extract the context for the current application state and the user activity scenario through sensor data and system state analysis to obtain context information; Perform the second layer of misoperation filtering on the key event after intention confirmation through a context conflict detector according to the context information to obtain a key event after context verification; Evaluate the semantic rationality of the historical key sequence through language model analysis to obtain a semantic verification rule; Perform the third layer of misoperation filtering on the key event after context verification through a temporal logic validator to analyze the semantic rationality of consecutive key events according to the semantic verification rule to obtain a verified target key event; 7. The application control method for mapping three-dimensional motion data to alphabet keys as claimed in claim 1, characterized in that Generate multi-modal user feedback for the target key event to obtain a key event with a feedback signal, and apply the target key event with the feedback signal to the target application program to achieve application program control, including: Generate visual feedback for the target key event through a dynamic indicator on the screen to display the current recognition status, candidate alphabets, and confidence to obtain a key event with visual feedback; Generate auditory feedback for the key event with visual feedback through a beep of different tones corresponding to different types of keys and system states to obtain a key event with audio-visual feedback; Generate haptic feedback for the key event with audio-visual feedback through the vibration motor of the wearable device, associate different modes of vibration with specific input behaviors, and obtain a key event with a multi-modal feedback signal; Integrate the key event with the multi-modal feedback signal through a hierarchical architecture interface to obtain an integrable key event; Select an integration mode for the integrable key event through the analysis of the current application type to obtain a key event adapted to the mode; Transmit the key event adapted to the mode through an application programming interface and apply the key event adapted to the mode to the target application to achieve application control.
8. An application control device for mapping three-dimensional motion data into letter keys, characterized in that, It includes: A data acquisition and preprocessing module for acquiring the original multi-modal three-dimensional motion data of the user from the inertial measurement unit sensor and performing preprocessing to obtain standardized three-dimensional motion data, and detecting the motion state of the standardized three-dimensional motion data to obtain potential gesture action data; A feature extraction and dimensionality reduction module connected to the data acquisition and preprocessing module for extracting a feature set from the potential gesture action data, adjusting the weights of the feature set based on context information to obtain context-weighted feature data, and performing multi-stage dimensionality reduction processing on the context-weighted feature data to obtain a highly discriminative feature vector; A gesture recognition and classification module connected to the feature extraction and dimensionality reduction module for inputting the highly discriminative feature vector into a classification system for processing, and performing comprehensive analysis by combining application scenario information through a multi-layer classification architecture, a personalized adaptive mechanism, and a feature stability protection algorithm to obtain a gesture category identifier and its confidence score; A key mapping module connected to the gesture recognition and classification module for generating an initial key mapping through multi-level mapping based on the gesture category identifier and its confidence score, and adjusting the initial key mapping based on user preferences, context prediction, and application environment to obtain a final key event; A false touch prevention module connected to the key mapping module for performing multi-level false touch detection on the final key event to obtain a verified target key event; A feedback and application control module connected to the false touch prevention module for generating multi-modal user feedback for the target key event to obtain a key event with a feedback signal, and applying the target key event with the feedback signal to the target application to achieve application control.
9. An application control device that maps three-dimensional motion data to letter keys, characterized in that, It includes a memory, a processor, and an application control program stored on the memory and executable on the processor for mapping three-dimensional motion data to alphabet keys. When the processor executes the application control program for mapping three-dimensional motion data to alphabet keys, it implements the application control method for mapping three-dimensional motion data to alphabet keys according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, An application control program for mapping three-dimensional motion data to alphabet keys is stored on the computer-readable storage medium. When the application control program for mapping three-dimensional motion data to alphabet keys is executed by the processor, it implements the application control method for mapping three-dimensional motion data to alphabet keys according to any one of claims 1-7.
Citation Information
Cited By
Virtual scene interaction system and method based on inertial measurement unit
CN121560167A