Virtual key interaction method based on three-dimensional motion gesture
Through multi-level signal processing and visual perception of three-dimensional motion data combined with application state model, high stability and accurate gesture interaction in complex dynamic interfaces are achieved, solving the problem of insufficient gesture recognition in the prior art and improving user experience.
Patent Information
- Application Number
- CN202510464910.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
The existing interactive technology based on three-dimensional motion gestures has insufficient stability and accuracy in complex dynamic interface applications, and it is impossible to dynamically perceive changes in virtual key functions on the screen, resulting in poor user experience.
By performing multi-level signal processing on the three-dimensional motion data collected by the inertial measurement unit, combining visual perception and application state model, dynamic segmentation and gesture intention recognition are realized, context-aware mapping and validity confirmation are carried out, standard input events are generated, and multimodal feedback is provided.
It improves the stability and accuracy of gesture interaction in complex dynamic interfaces, reduces the error operation rate, and enhances the user's operating experience and sense of control.
Smart Images

Figure CN120295480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and in particular to a virtual button interaction method, device, equipment and computer-readable storage medium based on three-dimensional motion gestures. Background Art
[0002] As a bridge connecting users and computing devices, the development of human-computer interaction technology continuously affects the expansion of user experience and application scenarios. Currently, direct operation interfaces represented by touch screens dominate, but with the rise of technologies such as wearable devices, smart homes, virtual reality, and augmented reality, the demand for new interaction methods is increasing. Interaction based on three-dimensional motion gestures, especially using sensors such as inertial measurement units (IMUs) to capture user limb movements for input, shows application potential in specific application fields due to its advantages such as non-contact, rich operation dimensions, and potential natural intuitiveness.
[0003] However, existing three-dimensional motion gesture-based interaction technologies still face several challenges in practical applications. Although IMU-based solutions are portable and insensitive to environmental light, they are vulnerable to sensor noise, drift, and individual user movement differences, making it difficult to ensure the stability and accuracy of gesture recognition. Especially when differentiating subtle or fast continuous actions, precise dynamic segmentation and intention recognition remain difficult points. Vision-based solutions can obtain richer information, but they usually have large computational requirements, are sensitive to light and occlusion, and may involve privacy issues. More critically, most current gesture interaction systems map the recognized gestures to predefined functions or fixed button commands, and this mapping relationship is usually static, lacking the ability to perceive real-time changes in the target application interface. When the virtual button layout, usability, or function in the application interface (especially complex applications such as games) changes dynamically with the context, fixed gesture mapping often leads to operation failures or unexpected results, resulting in a poor user experience.
[0004] Therefore, there is an urgent need to propose a closed-loop interaction method that can dynamically perceive available virtual buttons on the screen, understand their functions in the current state, and intelligently and accurately map user gestures to target virtual buttons, so as to improve the practicality and reliability of three-dimensional gesture interaction in complex dynamic interface applications. Summary of the Invention
[0005] Embodiments of this application provide a virtual button interaction method based on three-dimensional motion gestures, aiming to improve the practicality and reliability of three-dimensional gesture interaction in complex dynamic interface applications.
[0006] To achieve the above objective, embodiments of this application provide a virtual button interaction method based on three-dimensional motion gestures, including:
[0007] Perform multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit to obtain a standardized motion feature stream;
[0008] Perform dynamic segmentation and gesture intention recognition processing on the standardized motion feature stream to obtain gesture recognition results;
[0009] Perform visual perception and analysis on the screen display content of the current application to obtain screen button layout information, where the screen button layout information includes position information and category labels;
[0010] Perform context association and function parsing processing on the screen button information based on a predefined application state model to obtain a set of valid virtual buttons containing function annotations;
[0011] Perform context-aware mapping and validity confirmation on the gesture recognition results and the set of valid virtual buttons to obtain a determined target virtual button and an associated activation instruction;
[0012] Perform event generation and execution processing on the target virtual button and the activation instruction to obtain standard input events injected into the target platform and corresponding multimodal feedback signals.
[0013] To achieve the above object, an embodiment of the present application further proposes a virtual button interaction device based on three-dimensional motion gestures, including:
[0014] An inertial measurement unit interface for receiving the original three-dimensional motion data collected from the inertial measurement unit;
[0015] A signal processing module connected to the inertial measurement unit interface, configured to perform multi-level signal processing including signal filtering and data standardization on the original three-dimensional motion data to obtain a standardized motion feature stream;
[0016] A gesture recognition and intention parsing module connected to the signal processing module, configured to perform dynamic segmentation and gesture intention recognition processing on the standardized motion feature stream to obtain gesture recognition results;
[0017] A screen visual analysis module configured to perform visual perception and analysis on the screen display content of the current application to obtain screen button layout information, where the screen button layout information includes position information and category labels;
[0018] A context association and function parsing module configured to perform context association and function parsing processing on the screen button information based on a predefined application state model to obtain a set of valid virtual buttons containing function annotations;
[0019] A mapping decision and verification module, connected to the gesture recognition and intention parsing module and the context association and function parsing module, is configured to perform context-aware mapping and validity confirmation on the gesture recognition result and the set of valid virtual keys, so as to obtain a determined target virtual key and an associated activation instruction;
[0020] An event injection and feedback module, connected to the mapping decision and verification module, is configured to perform event generation and execution processing on the target virtual key and the activation instruction, so as to obtain a standard input event injected into the target platform and a corresponding multimodal feedback signal.
[0021] To achieve the above object, an embodiment of the present application further provides a virtual key interaction device based on three-dimensional motion gestures, including a memory, a processor, and a virtual key interaction program based on three-dimensional motion gestures stored on the memory and executable on the processor. When the processor executes the virtual key interaction program based on three-dimensional motion gestures, the virtual key interaction method based on three-dimensional motion gestures as described in any one of the above is implemented.
[0022] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, on which a virtual key interaction program based on three-dimensional motion gestures is stored. When the virtual key interaction program based on three-dimensional motion gestures is executed by a processor, the virtual key interaction method based on three-dimensional motion gestures as described in any one of the above is implemented.
[0023] The technical solution of the present application has the following beneficial effects:
[0024] First, by performing multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit (IMU), including specific filtering (such as Kalman filtering or complementary filtering) and normalization (such as Z-score normalization), the influence of sensor noise, drift, and individual data scale differences on subsequent processing is effectively suppressed. Combining the dynamic segmentation technology based on multi-index analysis and adaptive threshold used in composite gesture recognition processing, and the end point judgment mechanism based on the four-state transition model, it is possible to more accurately extract the gesture segments intentionally executed by the user from the continuous and interference-containing standardized motion feature stream.
[0025] Secondly, in the gesture recognition stage, through composite gesture recognition processing, especially by applying the user intention prediction algorithm to analyze the initial motion data to obtain the probability distribution of the user intention gesture type, and fusing and making decisions with the similarity scores between the complete gesture segments and the templates calculated based on the constrained dynamic time warping (DTW) algorithm, not only the overall shape information of the gesture is utilized, but also the intention clues contained in the early motion patterns are combined. This fusion method enables the gesture recognition result to better reflect the true intention of the user, distinguish similar gestures with different intentions, and thus obtain a more reliable gesture category identifier.
[0026] Furthermore, this application introduces the technical feature of visually perceiving and analyzing the display content of the application program screen to obtain the screen key layout information. This feature enables the system to have the ability to perceive the current user interface in real time. Through subsequent context association and function parsing processing based on the predefined application program state model, the system can combine the visually detected key information with the current running state of the application program, not only verifying the validity of these keys in the current context, but also accurately parsing the specific functions they represent at this moment, obtaining a set of valid virtual keys containing function annotations. This closed-loop visual and context perception mechanism is the key to realizing intelligent and adaptive interaction.
[0027] Furthermore, the mapping and validity confirmation processing based on context perception is the core advantage of the present invention. It is no longer a simple static mapping from gestures to commands, but dynamically matches the recognized gesture category identifier (representing the user intention function) with the set of valid virtual keys (representing the keys that are available and whose functions are understood on the current interface). At the same time, the application program state model is used to perform context execution validity verification to ensure that the key function triggered by the gesture is an allowed operation in the current state. This mapping decision under multiple constraints ensures that the user's gesture can be accurately and unambiguously associated with the only correct target virtual key on the interface, significantly reducing the misoperation rate in complex and dynamic interfaces (such as games), and solving the problems of disconnection between gestures and interfaces and rigid mapping relationships in the prior art.
[0028] Finally, standard input events are generated and injected into the target platform through event generation and execution processing, and corresponding multimodal feedback signals are generated at the same time. In particular, the generation of the feedback signal is based on a preset mapping table and can be adaptively adjusted according to the environment or user preferences, providing the user with clear, timely, and personalized operation confirmation, enhancing the closed-loop experience of the interaction and the user's sense of control. Description of the Drawings
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0030] Figure 1 FIG. is a module structure diagram of an embodiment of a virtual button interaction device based on three-dimensional motion gestures of the present invention;
[0031] Figure 2 FIG. is a flowchart of an embodiment of a virtual button interaction method based on three-dimensional motion gestures of the present invention.
[0032] The realization of the purpose, functional characteristics and advantages of the present invention will be further described with reference to the embodiments and the drawings. Detailed Embodiments
[0033] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0034] To better understand the above technical solutions, the exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0035] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" in the text does not exclude the presence of components or steps not listed in the claims. The singular number "a" or "an" before a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware including several different components and by means of a computer appropriately programmed. In the unit claims listing several devices, several of these devices can be embodied by the same hardware item. The use of "first", "second", and "third", etc. does not denote any order and these words can be interpreted as names.
[0036] As Figure 1 shown, Figure 1 FIG. is a schematic structural diagram of a server 1 (also called a virtual button interaction device based on three-dimensional motion gestures) of the hardware operating environment involved in the embodiment solution of the present invention.
[0037] The server in the embodiment of the present invention includes devices with display functions such as "Internet of Things devices", smart air conditioners with networking functions, smart lights, smart power supplies, AR / VR devices with networking functions, smart speakers, autonomous driving vehicles, PCs, smartphones, tablets, e-book readers, and portable computers.
[0038] Such as Figure 1 As shown, the server 1 includes: a memory 11, a processor 12, and a network interface 13.
[0039] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 11 can be an internal storage unit of the server 1 in some embodiments, such as the hard disk of the server 1. The memory 11 can also be an external storage device of the server 1 in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the server 1.
[0040] Furthermore, the memory 11 can also include both the internal storage unit and the external storage device of the server 1. The memory 11 can be used not only to store application software installed on the server 1 and various types of data, such as the code of the virtual button interaction program 10 based on three-dimensional motion gestures, but also to temporarily store data that has been output or will be output.
[0041] The processor 12 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 11 or process data, such as executing the virtual button interaction program 10 based on three-dimensional motion gestures.
[0042] The network interface 13 can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the server 1 and other electronic devices.
[0043] The network can be the Internet, a cloud network, a Wi-Fi network, a personal area network (PAN), a local area network (LAN), and / or a metropolitan area network (MAN). Various devices in the network environment can be configured to connect to a communication network according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols can include, but are not limited to, at least one of the following: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Li-Fi, 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocol, and / or Bluetooth communication protocol, or a combination thereof.
[0044] Optionally, the server may further include a user interface, and the user interface may include a display, an input unit such as a keyboard. Optionally, the user interface may further include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be referred to as a display screen or a display unit, and is used to display the information processed in the server 1 and to display a visual user interface.
[0045] Figure 1 Only the server 1 with components 11-13 and the virtual key interaction program 10 based on three-dimensional motion gestures is shown. Those skilled in the art can understand that Figure 1 The shown structure does not constitute a limitation on the server 1, and it may include fewer or more components than shown, or combine some components, or have different component arrangements.
[0046] In this embodiment, the processor 12 can be used to call the virtual key interaction program based on three-dimensional motion gestures stored in the memory 11 and perform the following operations:
[0047] Perform multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit to obtain a standardized motion feature stream;
[0048] Perform dynamic segmentation and gesture intention recognition processing on the standardized motion feature stream to obtain a gesture recognition result;
[0049] Visually perceive and analyze the screen display content of the current application to obtain screen button layout information, where the screen button layout information includes position information and category labels;
[0050] Perform context association and function parsing processing on the screen button information based on a predefined application state model to obtain a set of valid virtual buttons containing function annotations;
[0051] Perform context-aware mapping and validity confirmation on the gesture recognition result and the set of valid virtual buttons to obtain a determined target virtual button and an associated activation instruction;
[0052] Perform event generation and execution processing on the target virtual button and the activation instruction to obtain standard input events injected into the target platform and corresponding multimodal feedback signals.
[0053] In one embodiment, the processor 12 can be used to call the virtual button interaction program based on three-dimensional motion gestures stored in the memory 11 and perform the following operations:
[0054] Apply a digital low-pass filter to the original three-dimensional motion data for filtering processing to obtain filtered three-dimensional motion data;
[0055] Apply a Kalman filter to the filtered three-dimensional motion data for sensor fusion processing to generate fusion motion data containing real-time attitude information;
[0056] Apply a motion state detection algorithm to the fusion motion data. By calculating the short-time energy, signal amplitude variance, and main frequency components of the fusion motion data within a sliding time window and comparing them with a motion activation threshold dynamically adjusted based on the user's recent activity level, identify and filter out potential gesture signal segments representing the user's intention to execute a gesture;
[0057] Map the data of the potential gesture motion segment to a preset range to obtain the standardized motion feature stream.
[0058] In one embodiment, the processor 12 can be used to call the virtual button interaction program based on three-dimensional motion gestures stored in the memory 11 and perform the following operations:
[0059] Apply a sliding window to the standardized motion feature stream to calculate energy change rate indicators, direction consistency indicators, and motion complexity indicators to obtain a time series of gesture feature indicators;
[0060] Apply weighted combination and adaptive threshold comparison processing to the time series of gesture feature indicators to mark the start point of the gesture;
[0061] Apply a four - state transition model based on data statistical characteristics to the standardized motion feature stream for state probability estimation processing, and mark the end point of the gesture;
[0062] Segment a gesture segment from the standardized motion feature stream based on the start point of the gesture and the end point of the gesture;
[0063] Apply a user - intent prediction algorithm to the initial part of the data of the gesture segment to obtain a probability distribution of user - intent gesture types covering predefined gesture categories, where the user - intent prediction algorithm is based on the temporal pattern of the initial part of the data;
[0064] Apply a dynamic time warping algorithm with Sakoe - Chiba bandwidth constraint and global path constraint to calculate the time - series similarity score between the gesture segment and a predefined gesture template;
[0065] Apply a fusion decision function by combining the time - series similarity score and the probability distribution of user - intent gesture types to obtain the gesture recognition result, where the fusion decision function synthesizes the prediction probability of the initial part of the data and the overall shape similarity of the complete gesture segment.
[0066] In one embodiment, the processor 12 can be used to call the virtual button interaction program stored in the memory 11 and perform the following operations:
[0067] Obtain a screen capture of the current application program through an operating system interface or frame buffer reading method;
[0068] Input the screen capture into a pre - trained object detection model to obtain a list of detected virtual buttons;
[0069] Determine the position and size of each virtual button in the list of detected virtual buttons to obtain the bounding box position data of each virtual button;
[0070] Perform category recognition on each virtual button in the list of detected virtual buttons to obtain the category label of each virtual button;
[0071] Combine the bounding box position data and the category label to obtain the screen button layout information.
[0072] In one embodiment, the processor 12 can be used to call the virtual button interaction program stored in the memory 11 and perform the following operations:
[0073] Determine the current running state of the application program based on the predefined application program state model and the current interaction context information;
[0074] Retrieve the expected virtual key category and function annotation rules associated with the current running state from the application state model;
[0075] Apply the expected virtual key category and function annotation rules to perform matching, filtering, and function attribute association processing on each virtual key element included in the screen key layout information, and obtain the set of valid virtual keys with function annotations.
[0076] In one embodiment, the processor 12 can be used to call the virtual key interaction program based on three-dimensional motion gestures stored in the memory 11, and perform the following operations:
[0077] Query the preset gesture intention-key function mapping rules based on the gesture recognition result to determine the corresponding target interaction function;
[0078] Search for a virtual key in the set of valid virtual keys whose function annotation matches the target interaction function;
[0079] When there is a uniquely matching virtual key, perform context execution validity confirmation processing based on the current running state and the state transition rules in the application state model;
[0080] If the context execution validity confirmation processing passes, determine the uniquely matching virtual key as the target virtual key and generate the associated activation instruction.
[0081] In one embodiment, the processor 12 can be used to call the virtual key interaction program based on three-dimensional motion gestures stored in the memory 11, and perform the following operations:
[0082] Determine the type and coordinate parameters of the target platform input event according to the position information of the target virtual key and the activation instruction;
[0083] Call the event injection interface or auxiliary function interface of the target platform, and generate and inject the standard input event according to the type and coordinate parameters;
[0084] Query the preset reflection mode mapping table according to the execution state or type of the activation execution, and generate a feedback instruction including a specified feedback channel, feedback mode, and feedback intensity level, where the feedback channel includes at least one of a tactile channel, an auditory channel, and a visual channel, and the feedback mode includes at least one of a preset vibration waveform, audio sample, and visual element style;
[0085] Adjust the feedback intensity level or feedback channel priority in the feedback instruction according to the real-time environmental parameters obtained from the environmental sensor or the preference parameters set by the user, and obtain the adjusted feedback instruction;
[0086] Output the adjusted feedback instruction through a feedback actuator corresponding to the feedback channel to obtain the multimodal feedback signal.
[0087] Based on the above hardware architecture of the virtual button interaction device based on three-dimensional motion gestures, an embodiment of the virtual button interaction method based on three-dimensional motion gestures of the present invention is proposed. Refer to Figure 2 , Figure 2 This is an embodiment of the virtual button interaction method based on three-dimensional motion gestures of the present invention. The virtual button interaction method based on three-dimensional motion gestures includes the following steps:
[0088] S10. Perform multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit to obtain a standardized motion feature stream. The original data here usually includes the three-axis acceleration and angular velocity information provided by the accelerometer and gyroscope, reflecting the motion state of the user's limb in three-dimensional space. In order to facilitate subsequent gesture recognition and mapping, these original data must be converted into a consistent and processable feature form.
[0089] In some embodiments, step S10 can be implemented through steps S11-S14. Each step is explained in detail below:
[0090] S11. Apply a digital low-pass filter to the original three-dimensional motion data for filtering to obtain filtered three-dimensional motion data. The inertial measurement unit (IMU) usually includes an accelerometer and a gyroscope, and its original data often contains high-frequency noise. The digital low-pass filter can retain the low-frequency components in the signal (i.e., the actual gesture actions of the user) while removing high-frequency noise. Commonly used low-pass filters include Butterworth filters or FIR filters. By setting an appropriate cut-off frequency (e.g., 20 Hz), the jitter of the sensor and environmental interference can be effectively filtered out while retaining the main features of the gesture actions.
[0091] S12. Apply a Kalman filter to the filtered three-dimensional motion data for sensor fusion processing to generate fusion motion data containing real-time attitude information. The Kalman filter is a recursive estimation algorithm that can optimally fuse the data of the accelerometer and gyroscope. In this step, the system calculates the angle change using the gyroscope data and corrects it in combination with the accelerometer data to obtain a more accurate attitude estimation (including the direction and tilt angle of the device, etc.). This fusion processing can effectively overcome the limitations of a single sensor, such as the drift problem of the gyroscope and the insensitivity of the accelerometer to rapid motion.
[0092] S13. Apply a motion state detection algorithm to process the fused motion data. By calculating the short-time energy, signal amplitude variance, and main frequency components of the fused motion data within a sliding time window, and comparing them with a motion activation threshold dynamically adjusted based on the user's recent activity level, identify and filter out potential gesture signal segments representing the user's intention to execute a gesture. This step realizes the preliminary detection of meaningful gestures by calculating the short-time energy, signal amplitude variance, and main frequency components of the fused motion data within a sliding time window and comparing them with a motion activation threshold dynamically adjusted based on the user's recent activity level. The sliding time window is usually set to 100 - 300 milliseconds. Calculating these feature metrics within the window can effectively distinguish the user's conscious gestures from unconscious natural movements. The system also dynamically adjusts the threshold according to the user's recent activity level. For example, when the user is in a high-activity state, the activation threshold is increased to avoid false triggers.
[0093] S14. Map the data of the potential gesture motion segment to a preset range to obtain the standardized motion feature stream. This step eliminates the influence of individual differences and device differences by unifying the data generated by different users and different devices to the same numerical range (such as [-1, 1] or [0, 1]). Methods such as min-max normalization or Z-score normalization can be used for standardization processing to ensure that subsequent gesture recognition algorithms can process consistent data and improve the accuracy and robustness of recognition.
[0094] For example, assume the user executes a gesture of swiping right. The raw three-dimensional acceleration data collected by the IMU may show the following changes in the X-axis direction (unit: m / s 2 ): [0.2, 0.5, 1.2, 2.5, 3.8, 4.2, 3.5, 2.1, 1.0, 0.3, 0.1], accompanied by high-frequency noise fluctuations. In step S11, after applying a Butterworth low-pass filter with a cut-off frequency of 15 Hz, the data becomes: [0.18, 0.48, 1.15, 2.45, 3.75, 4.15, 3.48, 2.08, 0.98, 0.28, 0.09], and the high-frequency noise is effectively filtered out.
[0095] In step S12, the system combines gyroscope data (such as angular velocity data: [0.05, 0.12, 0.25, 0.35, 0.38, 0.36, 0.28, 0.15, 0.08, 0.03, 0.01] rad / s) through a Kalman filter for fusion to obtain a more accurate attitude estimate, including the direction change of the device in space. The state transition equation of the Kalman filter can be expressed as: X(k) = F·X(k - 1) + B·u(k) + w(k), where X represents the state vector, F is the state transition matrix, B is the control matrix, u is the control vector, and w is the process noise.
[0096] In step S13, the system calculates the short-time energy E = Σx 2 (i), the signal amplitude variance Var = Σ(x(i)-μ)2 / n, and the main frequency components obtained through FFT analysis. When the calculated eigenvalue (such as E = 45.6, Var = 2.3) exceeds the currently set threshold (such as E_threshold = 30.0, Var_threshold = 1.5), the system identifies that this segment of data may represent a conscious gesture.
[0097] Finally, in step S14, the system maps the identified gesture data to the range [-1, 1] through the min-max normalization method: x_norm = 2*(x - x_min) / (x_max - x_min) - 1, obtaining the normalized data: [-0.91, -0.81, -0.56, -0.10, 0.46, 0.60, 0.37, -0.17, -0.62, -0.88, -0.95]. The data processed in this way eliminates the amplitude difference, facilitating subsequent feature extraction and pattern recognition.
[0098] It can be understood that since the digital low-pass filter effectively removes high-frequency noise, the signal-to-noise ratio of the original data is improved, making subsequent processing more accurate; in addition, through sensor fusion using the Kalman filter, the advantages of the accelerometer and gyroscope are comprehensively utilized, overcoming the limitations of a single sensor and obtaining a more accurate attitude estimate; at the same time, the motion state detection algorithm based on a sliding window can intelligently identify the conscious gestures of users and adapt to the activity levels of different users through dynamic threshold adjustment, effectively reducing false triggers; on the other hand, data normalization processing eliminates the influence of individual differences and device differences, making the system have better adaptability and consistency for different users and devices. Through the synergistic effect of these multi-level signal processing steps, the system can extract a high-quality standardized motion feature stream from the noisy original data, laying a solid foundation for subsequent gesture recognition.
[0099] S20. Perform dynamic segmentation and gesture intention recognition processing on the standardized motion feature stream to obtain a gesture recognition result. This step accurately segments the continuous motion data stream into independent gesture units and identifies the specific gesture types represented by each gesture unit.
[0100] In some embodiments, step S20 is implemented through steps S21 - S27:
[0101] S21. Apply a sliding window to the standardized motion feature stream to calculate the energy change rate index, direction consistency index, and motion complexity index, obtaining a time series of gesture feature indices. Specifically, the system moves a fixed-size sliding window (usually 150 - 250 milliseconds) over the standardized motion feature stream and calculates three key indices at each window position: the energy change rate index reflects the change speed of the motion intensity, the direction consistency index measures the stability of the motion direction, and the motion complexity index characterizes the complexity of the motion trajectory. These indices together constitute a time series describing the dynamic characteristics of the gesture, providing a basis for subsequent gesture boundary detection.
[0102] S22. Apply weighted combination and adaptive threshold comparison processing to the time series of gesture feature indices to mark the gesture start point. Specifically, the system combines the three indices calculated in step S21 to form a comprehensive score, and the weights can be adjusted according to different application scenarios. For example, in scenarios where fast gestures need to be accurately captured, a higher weight can be assigned to the energy change rate index. The system also dynamically adjusts the threshold based on the user's historical interaction data. When the comprehensive score exceeds the adaptive threshold, this time point is marked as a potential gesture start point.
[0103] S23. Apply a four-state transition model based on the statistical characteristics of the data to the standardized motion feature stream for state probability estimation processing to mark the gesture end point. Specifically, the four-state transition model includes a stationary state, a start state, an execution state, and an end state. The system estimates the current most likely state by analyzing the statistical characteristics of the data (such as mean, variance, kurtosis, etc.). When the probability of the model transitioning from the execution state to the end state exceeds a preset threshold, the system marks this time point as the gesture end point.
[0104] S24. Segment the gesture segment from the standardized motion feature stream based on the gesture start point and the gesture end point. Specifically, the system extracts the data between the start point marked in step S22 and the end point marked in step S23 to form a complete gesture segment. This dynamic segmentation method can adapt to the operation habits of different users and the execution speeds of different gestures, ensuring accurate capture of the complete gesture action.
[0105] S25. Apply a user intention prediction algorithm to the initial part of the data of the gesture segment to obtain a probability distribution of user intention gesture types covering predefined gesture categories, where the user intention prediction algorithm is based on the temporal pattern of the initial part of the data. Specifically, the system only uses the first 30% - 40% of the data of the gesture segment and predicts the possible gesture types that the user may execute through a lightweight temporal model (such as a simplified LSTM network). This early prediction mechanism can speculate the user's intention in advance before the gesture is completed, providing the possibility for subsequent real-time response.
[0106] S26. Apply the dynamic time warping algorithm with Sakoe - Chiba bandwidth constraint and global path constraint to calculate the time - series similarity score between the gesture segment and the predefined gesture template. Specifically, the dynamic time warping (DTW) algorithm can handle time - series data of different lengths and speeds, and calculates the similarity by finding the optimal alignment path. The Sakoe - Chiba bandwidth constraint limits the search range of the alignment path, reducing the computational complexity; the global path constraint ensures that the alignment satisfies monotonicity and continuity, improving the rationality of the matching. The system compares the complete gesture segment with the predefined gesture template library to calculate the similarity score.
[0107] S27. Apply a fusion decision function by combining the time - series similarity score and the probability distribution of the user - intended gesture type to obtain the gesture recognition result, where the fusion decision function synthesizes the prediction probability of the initial part of the data and the overall shape similarity of the complete gesture segment. Specifically, the fusion decision function comprehensively considers the early prediction result in step S25 and the complete matching result in step S26, and obtains the final gesture type judgment through methods such as weighted average or Bayesian fusion. This fusion strategy utilizes both the real - time nature of the early prediction and ensures the accuracy based on the complete gesture.
[0108] For example, assume that the user executes a "swipe right" gesture. In step S21, the system calculates three metrics within the sliding window: the rate of energy change E_rate = Σ|E(i)-E(i - 1)| / n (where E(i) is the energy at the i - th sampling point), the direction consistency D_cons = |Σv(i)| / Σ|v(i)| (where v(i) is the velocity vector at the i - th sampling point), and the motion complexity C_comp = Σ|a(i)| (where a(i) is the acceleration vector at the i - th sampling point). For the "swipe right" gesture, the possible metric sequences are: E_rate = [0.1, 0.3, 0.8, 1.2, 0.9, 0.4, 0.2], D_cons = [0.65, 0.78, 0.92, 0.95, 0.93, 0.85, 0.7], C_comp = [0.2, 0.5, 0.9, 1.1, 0.8, 0.4, 0.2].
[0109] In step S22, the system performs a weighted combination of these three metrics. For example, using weights [0.3, 0.4, 0.3], it obtains a comprehensive score sequence [0.33, 0.54, 0.87, 1.07, 0.88, 0.57, 0.38]. Assume that the current adaptive threshold is 0.5, then the second time point (score 0.54) is marked as the gesture start point.
[0110] In step S23, the system applies a four-state transition model to analyze the data characteristics. For example, by calculating statistics such as variance and kurtosis within a short-time window, the system estimates the probabilities of being in four states at each time point. When the ending state probability at the sixth time point reaches 0.75 (exceeding the preset threshold of 0.7), this point is marked as the gesture ending point.
[0111] In step S24, the system extracts the data from the second time point to the sixth time point to form a complete gesture segment [0.54, 0.87, 1.07, 0.88, 0.57].
[0112] In step S25, the system uses only the first 40% of the data in the gesture segment [0.54, 0.87] to predict the possible gesture types through a pre-trained LSTM model. For example, the probability distribution output by the model may be: {"Swipe right": 0.65, "Swipe left": 0.15, "Swipe up": 0.10, "Swipe down": 0.05, "Other": 0.05}.
[0113] In step S26, the system performs DTW matching on the complete gesture segment [0.54, 0.87, 1.07, 0.88, 0.57] with a predefined gesture template library. For example, the DTW distance from the "Swipe right" template is 12.3, the distance from the "Swipe left" template is 28.7, the distance from the "Swipe up" template is 35.2, and so on. These distances can be converted into similarity scores: {"Swipe right": 0.85, "Swipe left": 0.42, "Swipe up": 0.31, "Swipe down": 0.28, "Other": 0.15}.
[0114] In step S27, the system combines the early prediction results and the complete matching results through a fusion decision function (such as weighted average with weights [0.4, 0.6]) to calculate the final probabilities: {"Swipe right": 0.4×0.65 + 0.6×0.85 = 0.77, "Swipe left": 0.4×0.15 + 0.6×0.42 = 0.31,...}. Finally, the system identifies this gesture as "Swipe right" with a confidence of 0.77.
[0115] It can be understood that, due to the use of a sliding window analysis with multiple feature indicators, the system can capture the dynamic characteristics of gestures more comprehensively, thereby improving the accuracy of gesture boundary detection; in addition, through the combination of an adaptive threshold and a four-state transition model, the system can adapt to different user operation habits and environmental changes, achieving more accurate gesture segmentation; at the same time, the fusion decision-making mechanism of early prediction and complete matching not only ensures the real-time performance of recognition but also guarantees the accuracy of the results, effectively balancing the requirements of response speed and recognition accuracy; on the other hand, the use of a constrained DTW algorithm not only improves the computational efficiency but also enhances the robustness to changes in the time scale, enabling the system to accurately recognize the same gestures at different speeds and durations. This multi-level and multi-angle dynamic segmentation and gesture recognition method significantly improves the accuracy of three-dimensional motion gesture recognition and the user experience.
[0116] S30. Perform visual perception and analysis on the screen display content of the current application to obtain screen button layout information, where the screen button layout information includes position information and category labels. This step is a key link in realizing the dynamic perception of screen virtual buttons, enabling the system to understand the layout and functions of the interactive elements in the current interface.
[0117] In some embodiments, step S30 is implemented through steps S31 - S35:
[0118] S31. Obtain a screen screenshot of the current application through an operating system interface or frame buffer reading method. Specifically, the system can obtain real-time screen content in various ways: in the Android system, the MediaProjection API or Screenshot API can be used; in the iOS system, the ReplayKit framework can be used; in the Windows system, the GDI+ or DirectX interface can be used; in the Linux system, the X11 or Wayland protocol can be used. In addition, for specific hardware platforms, the frame buffer of the graphics processor can be directly read. The system will select the most efficient and low-latency screenshot method according to the characteristics of the target platform to ensure real-time performance.
[0119] S32. Input the screen screenshot into a pre-trained object detection model to obtain a list of detected virtual buttons. Specifically, the system uses a deep learning model specially trained to recognize interface elements, such as an object detection network improved based on Faster R-CNN, YOLO, or SSD, etc. These models are trained with a large amount of UI interface data and can recognize interactive elements such as buttons, sliders, and switches in various applications. The model output includes the position information and preliminary classification results of each detected virtual button, forming a list of virtual buttons.
[0120] S33. Determine the position and size of each virtual button in the detected virtual button list to obtain the bounding box position data of each virtual button. Specifically, the system accurately calculates the coordinate positions (upper left and lower right coordinates) and sizes (width and height) of each virtual button on the screen. For the bounding boxes output by the detection model, the system will perform further refinement processes, such as edge alignment and size correction, to ensure that the bounding boxes accurately enclose the target button areas. These accurate position data are crucial for mapping gestures to the correct virtual buttons subsequently.
[0121] S34. Identify the category of each virtual button in the detected virtual button list to obtain the category label of each virtual button. Specifically, the system uses a dedicated classification model or the classification branch of an object detection model to perform fine-grained classification on each virtual button. The category labels usually include function types (such as "confirm button", "cancel button", "menu button", "direction control key", etc.) and interaction states (such as "available", "disabled", "selected", etc.). The system will make a comprehensive judgment by combining the visual features, context positions, and historical information of the buttons to improve the classification accuracy.
[0122] S35. Combine the bounding box position data with the category labels to obtain the screen button layout information. Specifically, the system integrates the position and size data obtained in step S33 with the category labels obtained in step S34 to form structured screen button layout information. These information are usually organized in JSON or a similar format, containing a complete description of each virtual button, providing basic data for subsequent context association and function parsing.
[0123] For example, assume that the user is running an action game, and the game interface includes direction control buttons, attack buttons, and a menu button. In step S31, the system obtains the current screen screenshot through Android's MediaProjection API, which is an RGB image with a resolution of 1920×1080 pixels.
[0124] In step S32, the system inputs the screenshot into a pre-trained YOLOv5 model. This model has been trained with a large amount of game interface data and can identify common game control elements. After the model processes the input, it outputs the detection results: four direction control buttons are detected in the lower left corner area, two attack buttons are detected in the lower right corner, and one menu button is detected in the upper right corner of the screen, for a total of 7 virtual buttons.
[0125] In step S33, the system precisely calculates the position and size of each virtual button. For example, the bounding box coordinates of the up arrow button are [150, 800, 250, 900] (indicating that the upper left corner coordinates are (150, 800) and the lower right corner coordinates are (250, 900)), and the size is 100×100 pixels; the bounding box coordinates of the attack button A are [1600, 750, 1700, 850], and the size is 100×100 pixels; the bounding box coordinates of the menu button are [1800, 50, 1880, 130], and the size is 80×80 pixels.
[0126] In step S34, the system performs category recognition on each virtual button. For example, the four arrow buttons are respectively recognized as "up arrow key", "down arrow key", "left arrow key", and "right arrow key", the two attack buttons are recognized as "main attack key" and "special attack key", and the upper right corner button is recognized as "menu key". At the same time, the system also recognizes that all buttons are currently in the "usable" state.
[0127] In step S35, the system combines the position data with the category labels to form the complete screen button layout information. For example, the information of the up arrow button can be expressed as: {"type": "up arrow key", "status": "usable", "position": [150, 800, 250, 900], "size": [100, 100]}. The complete button layout information contains similar structured data for all 7 virtual buttons.
[0128] It can be understood that due to the adoption of an efficient screen capture method, the system can capture changes in the application interface in real time, ensuring timely response to dynamic interfaces; in addition, through object detection by a specially trained deep learning model, the system can accurately identify virtual buttons in various applications, adapting to different interface styles and layouts; at the same time, precise position and size calculations ensure the accuracy of subsequent gesture mapping, improving the precision of interaction; on the other hand, fine-grained category recognition enables the system to understand the functions and states of virtual buttons, providing a semantic basis for intelligent mapping. This method of screen content understanding based on computer vision enables the system to dynamically adapt to interface changes in different applications, breaking through the limitations of traditional gesture interaction systems that rely on static mapping, and greatly enhancing the applicability and user experience of 3D gesture interaction in complex dynamic interfaces.
[0129] S40. Perform context association and functional parsing processing on the screen button information based on a predefined application state model to obtain a set of valid virtual buttons containing functional annotations. This step combines the button information obtained by visual recognition with the running context of the application to understand the actual functions and effectiveness of each button in the current interface.
[0130] In some embodiments, step S40 is implemented through steps S41 - S43:
[0131] S41. Determine the current running state of the application based on the predefined application state model and the current interaction context information. Specifically, the application state model is a structured model that describes various states that the application may be in and their transition relationships, usually represented by a finite state machine or a hierarchical state diagram. The system combines multiple context information to determine the current state, including: the previously identified interface state, the most recently executed operation sequence, a specific combination of elements displayed on the screen, the currently active application components, etc. For example, in a game application, the system may determine whether it is currently in the "main menu", "in combat", or "paused state" by identifying a specific combination of UI elements.
[0132] S42. Retrieve the expected virtual button categories and function annotation rules associated with the current running state from the application state model. Specifically, each application state has an associated set of expected virtual buttons and function annotation rules. The expected virtual button categories indicate the types of buttons that should appear in the current state. For example, in the game combat state, it is expected to have attack buttons, defense buttons, and skill buttons, etc. The function annotation rules define how to determine the specific functions of the buttons based on their visual characteristics, positions, and contexts. These rules can be rule - based logical expressions or machine learning models.
[0133] S43. Apply the expected virtual button categories and function annotation rules to perform matching, filtering, and functional attribute association processing on each virtual button element included in the screen button layout information, to obtain the set of valid virtual buttons including function annotations. Specifically, the system matches the screen button layout information obtained in step S30 with the expected button categories retrieved in step S42, and filters out the irrelevant or invalid buttons in the current state. For the matched buttons, the system applies the function annotation rules to determine their specific functions, such as "confirm selection", "return to the upper - level menu", "perform an attack action", etc. The finally formed set of valid virtual buttons includes not only the position and category information of the buttons, but also their specific function annotations in the current context.
[0134] For example, assume that the user is playing a role - playing game, and the system has already identified multiple virtual buttons on the screen through step S30, including four direction control buttons, three skill buttons, one backpack button, and one settings button.
[0135] In step S41, the system analyzes the current screen content and the recent interaction history, and finds that there are models of characters and enemies in the center of the screen, the background is a battle scene, and the operation of entering the battle has been recently executed. According to the predefined application state model, the system determines that it is currently in the "battle state". The state model of this game may include multiple states such as "main menu state", "exploration state", "battle state", "dialogue state", "store state", etc., and each state has specific UI features and available operations.
[0136] In step S42, the system retrieves the expected virtual button categories and function annotation rules associated with the "battle state" from the application state model. In the battle state, the expected virtual button categories include: "direction control keys" (for moving the character), "skill buttons" (for releasing skills), "menu buttons" (for pausing or viewing status). The function annotation rules may include: the four-way button located in the lower left corner of the screen is annotated as "movement control", the circular button located on the right side of the screen is annotated as "skill release", and the icon button located in the upper right corner of the screen is annotated as "battle menu", etc.
[0137] In step S43, the system matches and filters the identified virtual buttons with the expected categories. In this example, the four direction control buttons are matched to the "movement control" function, the three skill buttons are matched to the "skill release" function, and the settings button is matched to the "battle menu" function. The backpack button is usually unavailable (or has limited functionality) in the battle state, so it is marked as low priority or filtered out from the valid set. Finally, the system generates a valid set of virtual buttons with function annotations, such as:
[0138] Up arrow key: {"type":"up arrow key","position":[150,800,250,900],"function":"move forward","priority":"high"}
[0139] Skill button 1: {"type":"skill button","position":[1600,600,1700,700],"function":"release normal attack","priority":"high"}
[0140] Skill button 2: {"type":"skill button","position":[1600,750,1700,850],"function":"release special skill","priority":"high"}
[0141] Settings Button: {"type": "menu button", "position": [1800, 50, 1880, 130], "function": "Open combat menu", "priority": "medium"}
[0142] Inventory Button: {"type": "inventory button", "position": [1700, 50, 1780, 130], "function": "View items (restricted during combat)", "priority": "low"}
[0143] It can be understood that by adopting a predefined application state model, the system can accurately understand the current running state of the application, thereby more precisely interpreting the functions of interface elements. In addition, through state-related expected virtual key categories and function annotation rules, the system can assign different function interpretations to the same key in different states to adapt to the dynamic changes of the application. At the same time, the matching and filtering mechanism ensures that only the keys valid in the current state are included in the interaction consideration, reducing the possibility of misoperations. On the other hand, the function attribute association processing adds semantic-level function annotations to each virtual key, enabling the system to understand "what this button does" rather than just "there is a button here". This context-based function parsing method enables the system to understand the actual uses of interface elements like human users, greatly enhancing the intelligence and accuracy of gesture interaction, especially in applications with complex functions and variable interfaces.
[0144] S50. Perform context-aware mapping and validity confirmation on the gesture recognition result and the set of valid virtual keys to obtain a determined target virtual key and the associated activation instruction. This step is a key link in converting the user's gesture intention into specific virtual key operations, realizing an intelligent mapping from gestures to interface interactions.
[0145] In some embodiments, step S50 is implemented through steps S51 - S54:
[0146] S51. Query the preset gesture intention - key function mapping rules based on the gesture recognition result to determine the corresponding target interaction function. Specifically, the system maintains a mapping rule library between gesture intentions and key functions, and these rules define the possible interaction functions that various gestures may correspond to in different contexts. For example, a right swipe gesture may be mapped to functions such as "confirm", "next item", or "move right", depending on the context of the current application. The system queries this rule library according to the gesture recognition result (including gesture type and confidence) obtained in step S20, and combines the current application state to determine the most likely target interaction function.
[0147] S52. Search for a virtual button in the set of valid virtual buttons whose function annotation matches the target interaction function. Specifically, the system searches in the set of valid virtual buttons obtained in step S40 for a virtual button whose function annotation matches the target interaction function determined in step S51. The matching process takes into account the semantic similarity of functions, rather than just an exact string match. For example, if the target interaction function is "confirm selection", the system will search for virtual buttons with function annotations such as "confirm", "select", "determine", etc. that have related semantics. The system may find zero, one, or multiple matching virtual buttons.
[0148] S53. When there is a uniquely matching virtual button, based on the current running state and the state transition rules in the application state model, perform a context execution validity confirmation process. Specifically, if the system finds a uniquely matching virtual button in step S52, it will further confirm whether it is valid to activate the button in the current context. This confirmation is based on the state transition rules defined in the application state model, checking whether activating the button will result in a valid state transition. For example, in certain game states, even though a skill button is visible, due to the character's state (such as being stunned) or insufficient resources (such as not having enough magic points), activating the button may be invalid. The system analyzes the current state and possible transition conditions to confirm the validity of the operation.
[0149] S54. If the context execution validity confirmation process passes, determine the uniquely matching virtual button as the target virtual button and generate the associated activation instruction. Specifically, when the validity confirmation in step S53 passes, the system determines this virtual button as the final target virtual button and generates the corresponding activation instruction according to the type of the button and the current context. The activation instruction includes the operation type (such as click, long press, swipe, etc.) and operation parameters (such as coordinate position, duration, etc.), and these instructions will be converted into actual input events in subsequent steps.
[0150] For example, assume that the user is playing a role-playing game and is currently in a combat state and performs a gesture of "swipe quickly to the right".
[0151] In step S51, the system queries the gesture intention - button function mapping rule based on the gesture recognition result ("swipe quickly to the right", confidence 0.85). In the combat state, the gesture mapping rule may be: {"swipe quickly to the right": ["release skill", "use item", "move to the right"]}. Combining the current combat context and the execution characteristics of the gesture (fast speed, large amplitude), the system determines that the most likely target interaction function is "release skill".
[0152] In step S52, the system searches for a virtual button in the set of valid virtual buttons whose function annotation matches "Release Skill". Suppose there are three skill buttons in the current set of valid virtual buttons, and the function annotations are "Release Normal Attack", "Release Special Skill", and "Release Ultimate Skill" respectively. Through semantic matching, the system finds that all three buttons are related to the "Release Skill" function. By further analyzing the characteristics of the gesture (such as speed and amplitude) and the user's historical operation preferences, the system determines that the "Release Special Skill" button is the best match.
[0153] In step S53, the system performs context execution validity confirmation based on the current running state and the application state model. The system checks whether the character currently has enough magic value to use the special skill, whether the skill is in the cooldown state, and whether the character is in a state where the skill can be released (not stunned, silenced, etc.). Suppose the confirmation result shows that the character has enough magic value, the skill is not in the cooldown, and the character's state is normal, so activating this button is effective.
[0154] In step S54, the system determines the "Release Special Skill" button as the target virtual button and generates an associated activation instruction. The activation instruction may be: {"type":"tap","position":[1650,800],"duration":100}, indicating a click operation for 100 milliseconds at the coordinates (1650,800).
[0155] Another example, if the user performs a "swipe up" gesture, but there is no valid virtual button in the current combat state that matches the "move up" function (possibly because the current scene restricts moving up), the system will not generate an activation instruction to avoid invalid operations.
[0156] It can be understood that due to the adoption of the preset gesture intention - button function mapping rule, the system can flexibly interpret the intention of the gesture according to different contexts, enabling the same gesture to trigger different functions in different scenarios; in addition, by searching for buttons with matching functions in the set of valid virtual buttons, the system realizes the intelligent mapping from abstract functions to specific interface elements without the need for the user to precisely locate the buttons on the screen; at the same time, the context execution validity confirmation mechanism ensures that only operations that are truly effective in the current state will be executed, avoiding invalid operations and potential errors; on the other hand, this context - aware mapping method enables the system to adapt to dynamically changing interfaces. Even if the button position changes, as long as the function remains the same, the user can still use the same gesture to complete the operation. This intelligent mapping and validity confirmation mechanism significantly improves the intuitiveness and reliability of gesture interaction, enabling users to interact with complex applications more naturally.
[0157] S60. Generate and execute event processing for the target virtual button and the activation instruction to obtain the standard input events injected into the target platform and the corresponding multimodal feedback signals. This step is the execution link that converts the mapping decision into actual operations and provides perceptual feedback, completing the entire interaction closed-loop.
[0158] In some embodiments, step S60 is implemented through steps S61 - S65:
[0159] S61. Determine the type and coordinate parameters of the target platform input event according to the position information of the target virtual button and the activation instruction. Specifically, the system converts the position information (bounding box coordinates) of the target virtual button determined in step S50 into specific screen coordinate points, usually selecting the center point of the button as the operation coordinate. At the same time, according to the operation type (such as click, long press, slide, etc.) specified in the activation instruction, determine the type of input event to be generated. For example, for a click operation, two events of press and release need to be generated; for a slide operation, a series of events of press, move, and release need to be generated. The system also adjusts the coordinate system and event parameters according to the characteristics of the target platform to ensure that the generated events can be correctly recognized by the target platform.
[0160] S62. Call the event injection interface or auxiliary function interface of the target platform, and generate and inject the standard input event according to the type and coordinate parameters. Specifically, the system converts the event type and coordinate parameters determined in step S61 into standard input events according to the API interface provided by the target platform and injects them into the operating system or application. In the Android system, AccessibilityService or InputManager can be used; in the iOS system, UIAutomation can be used; in the Windows system, the SendInput function or UI Automation framework can be used. These interfaces allow the system to simulate user touch, click, or keyboard input, thus realizing the activation operation of the target virtual button.
[0161] S63. Query a preset albedo mode mapping table according to the execution status or type of the activation execution, and generate a feedback instruction including a specified feedback channel, a feedback mode, and a feedback intensity level, where the feedback channel includes at least one of a tactile channel, an auditory channel, and a visual channel, and the feedback mode includes at least one of a preset vibration waveform, an audio sample, and a visual element style. Specifically, the system maintains a feedback mode mapping table that defines the feedback methods corresponding to different types of operations and execution statuses. The feedback channels include a tactile channel (such as vibration), an auditory channel (such as sound), and a visual channel (such as a flash or animation). The feedback modes include a preset vibration waveform (such as short, long, or rhythmic vibration), an audio sample (such as a click sound, a confirmation sound, or an error prompt sound), and a visual element style (such as highlighting, flashing, or color change). The system queries an appropriate feedback instruction from the mapping table according to the type of the current operation and the execution result (success, failure, partial success, etc.).
[0162] S64. Adjust the feedback intensity level or the feedback channel priority in the feedback instruction according to the real-time environmental parameters obtained from the environmental sensor or the preference parameters set by the user, and obtain an adjusted feedback instruction. Specifically, the system obtains parameters such as the light intensity, the noise level, and the device motion state of the current environment through environmental sensors (such as an ambient light sensor, a microphone, an accelerometer, etc.), or reads the preference settings preset by the user, and intelligently adjusts the feedback instruction generated in step S63. For example, enhance the vibration feedback and weaken the sound feedback in a noisy environment; enhance the brightness of the visual feedback in a low-light environment; adjust the vibration intensity according to the user's preference. This environment-adaptive feedback mechanism ensures that the user can clearly perceive the operation result under various conditions.
[0163] S65. Output the adjusted feedback instruction through a feedback actuator corresponding to the feedback channel to obtain the multimodal feedback signal. Specifically, the system activates the corresponding hardware actuator, such as a vibration motor, a speaker, an LED indicator, or a screen display element, to generate a multimodal feedback signal according to the feedback instruction adjusted in step S64. These feedback signals are transmitted to the user simultaneously or sequentially through multiple perception channels, providing a rich operation confirmation experience. The collaborative design of multimodal feedback ensures the redundant transmission of information, and even if a certain channel is restricted, the user can still perceive the operation result through other channels.
[0164] For example, assume that the user is playing an action game and executes a gesture of "swiping quickly to the right", and the system maps it to the activation operation of the "release special skill" button through the previous steps.
[0165] In step S61, the system calculates the center point coordinates (1650, 800) based on the position information of the target virtual button (bounding box coordinates [1600, 750, 1700, 850]), and determines that two touch events, press (ACTION_DOWN) and release (ACTION_UP), need to be generated according to the activation instruction type "tap", with the operation duration being 100 milliseconds.
[0166] In step S62, the system calls the InputManager.injectInputEvent() method of the Android platform to generate and inject the following standard input event sequence:
[0167] MotionEvent(ACTION_DOWN, timestamp: t0, x: 1650, y: 800);
[0168] MotionEvent(ACTION_UP, timestamp: t0 + 100ms, x: 1650, y: 800);
[0169] These events are injected into the input event stream of the operating system, and the operating system passes them to the currently active game application, triggering the release action of the special skill.
[0170] In step S63, the system queries the preset feedback mode mapping table and finds that the feedback instructions corresponding to the successful execution of the "release special skill" operation are:
[0171] Tactile channel: A short vibration of medium intensity, followed by a long vibration of weak intensity;
[0172] Auditory channel: Special skill release sound effect, medium volume;
[0173] Visual channel: The button briefly highlights and then returns to normal;
[0174] In step S64, the system detects that the current ambient noise level is high (e.g., 85 decibels), the light is normal, the device is in a stable state, and the user preference is set to "strong vibration feedback". Based on these parameters, the system adjusts the feedback instructions:
[0175] Increase the vibration intensity of the tactile channel (from medium intensity to high intensity);
[0176] Increase the volume of the auditory channel (from medium volume to high volume);
[0177] Keep the visual channel feedback unchanged.
[0178] In step S65, the system outputs the adjusted feedback instructions through the corresponding feedback actuators:
[0179] Generate a high-intensity vibration of 50 milliseconds by the device's vibration motor, followed by a medium-intensity vibration of 200 milliseconds;
[0180] Play the sound effect of special skill release through the speaker at a high volume;
[0181] Display a short highlighting effect (100 milliseconds) of the special skill button on the screen;
[0182] These multimodal feedback signals together provide the user with a clear perception that the operation has been successfully executed.
[0183] It can be understood that since the system accurately calculates the operation coordinates of the target virtual button and determines the appropriate event type, it can accurately simulate the user's direct touch operation, ensuring the functional equivalence of gesture interaction and traditional touch interaction; in addition, by calling the platform-standard event injection interface, the system achieves seamless compatibility with various applications and supports gesture control without the need for the application to be specifically adapted; at the same time, the multimodal feedback mechanism provides the user with rich perception channels, enhancing the confirmation sense of operation and the immersive experience; on the other hand, the feedback adjustment based on environmental perception enables the system to adapt to various usage scenarios and ensure clear and effective feedback under different conditions. This intelligent event generation and multimodal feedback mechanism not only completes the conversion from gesture to actual operation, but also enhances the user's perception of the system response through closed-loop feedback, significantly improving the usability and user experience of three-dimensional gesture interaction.
[0184] In addition, the embodiment of the present invention also proposes a virtual button interaction device based on three-dimensional motion gestures, and the virtual button interaction device based on three-dimensional motion gestures includes:
[0185] An inertial measurement unit interface for receiving the original three-dimensional motion data collected from the inertial measurement unit;
[0186] A signal processing module connected to the inertial measurement unit interface and configured to perform multi-level signal processing including signal filtering and data normalization on the original three-dimensional motion data to obtain a normalized motion feature stream;
[0187] A gesture recognition and intention parsing module connected to the signal processing module and configured to perform dynamic segmentation and gesture intention recognition processing on the normalized motion feature stream to obtain a gesture recognition result;
[0188] A screen visual analysis module configured to perform visual perception and analysis on the screen display content of the current application to obtain screen button layout information, and the screen button layout information includes position information and category labels;
[0189] A context - related and function - parsing module, configured to perform context - related and function - parsing processing on the screen key information based on a predefined application - state model, and obtain a set of valid virtual keys containing function annotations;
[0190] A mapping - decision and verification module, connected to the gesture recognition and intention - parsing module and the context - related and function - parsing module, configured to perform context - aware mapping and validity confirmation on the gesture recognition result and the set of valid virtual keys, and obtain a determined target virtual key and an associated activation instruction;
[0191] An event - injection and feedback module, connected to the mapping - decision and verification module, configured to perform event generation and execution processing on the target virtual key and the activation instruction, and obtain a standard input event injected into the target platform and a corresponding multimodal feedback signal.
[0192] Among them, the steps implemented by each functional module of the virtual - key interaction device based on three - dimensional motion gestures can refer to each embodiment of the virtual - key interaction method based on three - dimensional motion gestures of the present invention, which will not be elaborated here.
[0193] In addition, an embodiment of the present invention also proposes a computer - readable storage medium. The computer - readable storage medium can be any one or any combination of a hard disk, a multimedia card, an SD card, a flash card, an SMC, a read - only memory (ROM), an erasable programmable read - only memory (EPROM), a portable compact disc read - only memory (CD - ROM), a USB memory, etc. The computer - readable storage medium includes a virtual - key interaction program 10 based on three - dimensional motion gestures. The specific implementation manner of the computer - readable storage medium of the present invention is substantially the same as the specific implementation manners of the above - mentioned virtual - key interaction method based on three - dimensional motion gestures and the server 1, which will not be elaborated here.
[0194] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer - program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer - program product implemented on one or more computer - usable storage media (including but not limited to disk memories, CD - ROMs, optical memories, etc.) containing computer - usable program codes.
Claims
1. A virtual button interaction method based on three-dimensional motion gestures, characterized in that Including: Performing multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit to obtain a standardized motion feature stream; Performing dynamic segmentation and gesture intention recognition processing on the standardized motion feature stream to obtain gesture recognition results; Performing visual perception and analysis on the screen display content of the current application to obtain screen button layout information, where the screen button layout information includes position information and category labels; Performing context association and function parsing processing on the screen button information based on a predefined application state model to obtain a set of valid virtual buttons containing function annotations; Performing context-aware mapping and validity confirmation on the gesture recognition results and the set of valid virtual buttons to obtain a determined target virtual button and an associated activation instruction; Performing event generation and execution processing on the target virtual button and activation instruction to obtain standard input events injected into the target platform and corresponding multimodal feedback signals.
2. The virtual key interaction method based on three-dimensional motion gestures according to claim 1, wherein Performing multi-level signal processing on the original three-dimensional motion data collected from the inertial measurement unit, including: Applying a digital low-pass filter to the original three-dimensional motion data for filtering processing to obtain filtered three-dimensional motion data; Applying a Kalman filter to the filtered three-dimensional motion data for sensor fusion processing to generate fusion motion data containing real-time attitude information; Applying a motion state detection algorithm to the fusion motion data for processing, by calculating the short-time energy, signal amplitude variance, and main frequency components of the fusion motion data within a sliding time window, and comparing with a motion activation threshold dynamically adjusted based on the user's recent activity level, to identify and screen out potential gesture signal segments representing the user's intention to execute a gesture; Mapping the data of the potential gesture motion segment to a preset range to obtain the standardized motion feature stream.
3. The virtual button interaction method based on three-dimensional motion gestures according to claim 1, wherein, Performing dynamic segmentation and gesture intention recognition processing on the standardized motion feature stream, including: Applying a sliding window to the standardized motion feature stream to calculate energy change rate metrics, direction consistency metrics, and motion complexity metrics to obtain a time series of gesture feature metrics; Applying weighted combination and adaptive threshold comparison processing to the time series of gesture feature metrics to mark the start point of the gesture; Applying a four-state transition model based on the statistical characteristics of the data to the standardized motion feature stream for state probability estimation processing to mark the end point of the gesture; Segmenting gesture segments from the standardized motion feature stream based on the gesture start point and the gesture end point; Applying a user intention prediction algorithm to the initial part of the data of the gesture segment to obtain a probability distribution of user intention gesture types covering predefined gesture categories, where the user intention prediction algorithm is based on the temporal pattern of the initial part of the data; Applying a dynamic time warping algorithm with Sakoe-Chiba bandwidth constraint and global path constraint to calculate the time series similarity score between the gesture segment and a predefined gesture template; Process using a fusion decision function by combining the time series similarity score and the probability distribution of the user's intended gesture types to obtain the gesture recognition result, where the fusion decision function synthesizes the prediction probability of the initial part of the data and the overall shape similarity of the complete gesture segment.
4. The virtual button interaction method based on three-dimensional motion gestures according to claim 1, characterized in that, Perform visual perception and analysis on the screen display content of the current application, including: Obtain a screen capture of the current application through an operating system interface or frame buffer reading method; Input the screen capture into a pre-trained object detection model to obtain a list of detected virtual buttons; Determine the position and size of each virtual button in the list of detected virtual buttons to obtain the bounding box position data of each virtual button; Perform category recognition on each virtual button in the list of detected virtual buttons to obtain the category label of each virtual button; Combine the bounding box position data and the category labels to obtain the screen button layout information.
5. The virtual button interaction method based on three-dimensional motion gestures according to claim 1, characterized in that, Perform context association and functional parsing processing on the screen button information based on a predefined application state model, including: Determine the current running state of the application based on the predefined application state model and the current interaction context information; Retrieve the expected virtual button categories and function annotation rules associated with the current running state from the application state model; Apply the expected virtual button categories and function annotation rules to perform matching, filtering, and functional attribute association processing on each virtual button element included in the screen button layout information to obtain the set of valid virtual buttons with function annotations.
6. The virtual key interaction method based on three-dimensional motion gestures according to claim 1, wherein Perform context-aware mapping and validity confirmation on the gesture recognition result and the set of valid virtual buttons, including: Query a preset gesture intent-button function mapping rule based on the gesture recognition result to determine the corresponding target interaction function; Search for a virtual button in the set of valid virtual buttons whose function annotation matches the target interaction function; When there is a unique matching virtual button, perform context execution validity confirmation processing based on the current running state and the state transition rules in the application state model; If the context execution validity confirmation processing passes, determine the unique matching virtual button as the target virtual button and generate the associated activation instruction.
7. The virtual button interaction method based on three-dimensional motion gestures according to claim 1, wherein Perform event generation and execution processing on the target virtual button and the activation instruction, specifically including: Determine the type and coordinate parameters of the target platform input event based on the position information of the target virtual button and the activation instruction; Call the event injection interface or auxiliary function interface of the target platform to generate and inject the standard input event according to the type and coordinate parameters; Query a preset reflection mode mapping table according to the execution status or type of the activation execution to generate a feedback instruction including a specified feedback channel, feedback mode, and feedback intensity level, where the feedback channel includes at least one of a tactile channel, an auditory channel, and a visual channel, and the feedback mode includes at least one of a preset vibration waveform, audio sample, and visual element style; Adjust the feedback intensity level or feedback channel priority in the feedback instruction according to the real-time environmental parameters obtained from the environmental sensor or the preference parameters set by the user, to obtain an adjusted feedback instruction; Output the adjusted feedback instruction through the feedback actuator corresponding to the feedback channel, to obtain the multi-modal feedback signal.
8. A virtual button interaction device based on three-dimensional motion gestures, characterized in that Comprising: An inertial measurement unit interface, configured to receive the original three-dimensional motion data collected from an inertial measurement unit; A signal processing module, connected to the inertial measurement unit interface, configured to perform multi-level signal processing including signal filtering and data normalization on the original three-dimensional motion data, to obtain a normalized motion feature stream; A gesture recognition and intention parsing module, connected to the signal processing module, configured to perform dynamic segmentation and gesture intention recognition processing on the normalized motion feature stream, to obtain a gesture recognition result; A screen vision analysis module, configured to perform visual perception and analysis on the screen display content of the current application program, to obtain screen key layout information, where the screen key layout information includes position information and category labels; A context association and function parsing module, configured to perform context association and function parsing processing on the screen key information based on a predefined application program state model, to obtain a set of valid virtual keys including function annotations; A mapping decision and verification module, connected to the gesture recognition and intention parsing module and the context association and function parsing module, configured to perform context-aware mapping and validity confirmation on the gesture recognition result and the set of valid virtual keys, to obtain a determined target virtual key and an associated activation instruction; An event injection and feedback module, connected to the mapping decision and verification module, configured to perform event generation and execution processing on the target virtual key and the activation instruction, to obtain a standard input event injected into the target platform and a corresponding multi-modal feedback signal.
9. A virtual key interaction device based on three-dimensional motion gestures, characterized in that Comprising a memory, a processor, and a virtual key interaction program based on three-dimensional motion gestures stored on the memory and executable on the processor, where when the processor executes the virtual key interaction program based on three-dimensional motion gestures, it implements the virtual key interaction method based on three-dimensional motion gestures according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A virtual key interaction program based on three-dimensional motion gestures is stored on the computer-readable storage medium, and when the virtual key interaction program based on three-dimensional motion gestures is executed by the processor, it implements the virtual key interaction method based on three-dimensional motion gestures according to any one of claims 1-7.
Citation Information
Cited By
Key event enhancement method and device based on Android system tactile feedback
CN120631184A
Quick option interaction method and device of electronic paper display equipment and electronic equipment
CN121807201A