Gesture-based interaction method, apparatus, storage medium, and electronic device

By acquiring hand movement information and using electromyography and inertial measurement units combined with a neural network model to identify hand interaction states, the problems of false triggering and response lag in existing gesture recognition technologies are solved, achieving efficient and accurate gesture interaction.

CN120848742BActive Publication Date: 2026-01-02HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511367201.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-01-02
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing gesture recognition technology cannot effectively identify users' interaction intentions, resulting in a high rate of false triggers and slow response, which affects user experience and interaction efficiency.

Method used

By acquiring motion information generated by hand gestures, electromyography (EMG) signals and acceleration signals are collected using a multi-channel EMG signal acquisition device and an inertial measurement unit. Combined with a feature classification model based on neural networks and a deep learning model, the interaction state of the hand is identified, and a target interaction command is generated when a clear interaction intention is identified.

Benefits of technology

It significantly improves the accuracy and response efficiency of gesture interaction, reduces false triggers, and enhances user experience and operational smoothness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848742B_ABST
    Figure CN120848742B_ABST
Patent Text Reader

Abstract

The application provides a gesture-based interaction method and device, a storage medium and an electronic device, and relates to the field of human-computer interaction. The electronic device acquires motion information generated by a gesture action of a hand; determines an interaction state in which the hand is located according to the motion information; and generates a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information. In this way, before the target interaction instruction is generated, the interaction state in which the hand of the user is currently located is first determined. Since the interaction state can represent whether the user has an explicit interaction intention.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of human-computer interaction, and in particular, to a gesture-based interaction method and device, a storage medium and an electronic device. BACKGROUND

[0002] Currently, human-computer interaction technology is evolving towards a more natural, seamless and portable direction. Under this trend, gesture-based input methods, especially gesture recognition technology relying on wearable devices, as a new type of human-computer interaction method that breaks away from the physical limitations of traditional mouse, keyboard and touch screen, have received widespread attention from the industry.

[0003] However, related gesture recognition technology still faces many challenges in actual application. Among them, the most prominent problem is that the user's interaction intention is not effectively recognized, resulting in a high false trigger rate (for example, the hand just touches the operation plane and is mistaken for a click operation) and response lag (for example, the user needs to make a gesture far beyond the normal click action amplitude to be recognized by the system). These problems seriously affect user experience and interaction efficiency, limiting the further popularization and application of gesture recognition technology. SUMMARY

[0004] In order to overcome at least one deficiency in the prior art, one of the purposes of the present application is to provide a gesture-based interaction method, device, storage medium and electronic device, comprising:

[0005] In a first aspect, the present application provides a gesture-based interaction method, the method comprising:

[0006] acquiring motion information generated by a gesture action of a hand;

[0007] determining an interaction state in which the hand is located according to the motion information;

[0008] generating a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information.

[0009] In a second aspect, the present application provides a gesture-based interaction device, the device comprising:

[0010] a motion information module for acquiring motion information generated by a gesture action of a hand;

[0011] an interaction state module for determining an interaction state in which the hand is located according to the motion information;

[0012] an instruction recognition module for generating a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information.

[0013] In a third aspect, the present application provides a storage medium, which stores a computer program. When the computer program is executed by a processor, the gesture-based interaction method is realized.

[0014] In a fourth aspect, the present application provides an electronic device, which comprises a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the gesture-based interaction method is realized.

[0015] Compared with the prior art, the present application has the following beneficial effects:

[0016] The present application provides a gesture-based interaction method, device, storage medium and electronic device. The electronic device acquires motion information generated by a gesture action of a hand; determines an interaction state in which the hand is located according to the motion information; and generates a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information. In this way, before the target interaction instruction is generated, the interaction state in which the user's hand is currently located is first determined. Since the interaction state can represent whether the user has an explicit interaction intention. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.

[0018] Figure 1 The flowchart of the gesture-based interaction method provided by the embodiments of the present application is shown in the figure;

[0019] Figure 2 The wearing interaction diagram of the acquisition device provided by the embodiments of the present application is shown in the figure;

[0020] Figure 3 The schematic diagram of the single-click operation provided by the embodiments of the present application is shown in the figure;

[0021] Figure 4 The schematic diagram of the right-click operation provided by the embodiments of the present application is shown in the figure;

[0022] Figure 5 The schematic diagram of the handwriting input provided by the embodiments of the present application is shown in the figure;

[0023] Figure 6 The structural schematic diagram of the gesture-based interaction device provided by the embodiments of the present application is shown in the figure;

[0024] Figure 7A structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to make the objectives, technical solutions and advantages of the embodiments of the present application (hereinafter referred to as the present embodiments) clearer, the technical solutions in the present embodiments will be described clearly and completely below with reference to the drawings in the present embodiments. Obviously, the described embodiments are some but not all of the embodiments of the present application. The components of the present embodiments described and shown in the drawings herein can be arranged and designed in various different configurations.

[0026] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by a person of ordinary skill in the art based on the embodiments in the present application without creative labor are within the scope of protection of the present application.

[0027] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0028] In the description of the present application, it should be noted that the terms "first", "second", "third" and the like are only used to distinguish descriptions, and cannot be understood as indicating or implying relative importance. In addition, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a…" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0029] Based on the above statement, as introduced in the background, the related technology for interaction based on gestures leads to high false triggering rate and response delay due to ineffective identification of user's interaction intention.

[0030] Exemplarily, taking the "click" operation as an example. In the related technology, the interaction system usually adopts a unified and non-context-specific recognition method when processing user input. This method mainly trains and processes the recognition model for the "click" action, but in actual use, it cannot accurately capture the real interaction intention implied by the click action in different scenarios.

[0031] For example, when a user's finger is gently placed on the operation plane from the air, or the hand has been placed on the desktop and is in a resting state, the subtle movements caused by slight muscle tremors or adjusting posture are often treated as the same input as the signal generated by an explicit click operation completed by the user's active force under the existing recognition mechanism. Due to the lack of accurate recognition ability for these subtle differences, unintentional contact is easily misjudged as an effective click, thereby triggering frequent false triggers. At the same time, in order to avoid triggering frequent false triggers, the sensitivity is often reduced, which makes the user have to make more exaggerated gestures than daily operations, which not only increases the operation burden, but also makes the interaction process appear sluggish and unnatural.

[0032] Therefore, due to the ineffective recognition of the user's interaction intention, the stability and reliability of the user experience are greatly affected, which has become a technical problem to be solved.

[0033] For example, continue to take the operation of "moving the cursor" as an example. Due to the lack of a special mechanism to determine whether the hand is in contact with the plane and ready to interact, it is difficult to accurately distinguish whether the user is unconsciously waving the arm in the air or intentionally moving the cursor on the plane. When the user's hand gradually approaches the desktop from the air and moves, the motion signal generated thereby can be incorrectly recognized as a cursor movement instruction, thereby causing the cursor to "drift" or "jitter" without clear operation intention, seriously affecting the accuracy of the operation and the fluency of the user during use. Similarly, when the user finishes the operation on the plane and is ready to lift the hand, the lifting action itself can also be misjudged as a drag or slide operation. Therefore, not only does it increase the complexity of the user's operation, but it can also cause unnecessary misoperation.

[0034] For example, the user may just want to end the current operation, but the lifting action is interpreted as a continuation of the operation, causing the cursor on the screen to be further dragged or slid. Therefore, due to the lack of an effective distinction mechanism between "in the air" and "on the plane" actions, there is a great uncertainty in processing user input. Not only does it affect the fluency and accuracy of the operation, but it also reduces the overall quality of the user experience.

[0035] For example, continue to take the scenario of "fine" operation as an example. For example, in document editing, graphics drawing or other application scenarios that require high precision control, users usually rely on a mouse or touchpad to achieve fine control of the cursor. In these scenarios, the gesture interaction system needs to have high sensitivity and high recognition accuracy for small movements to ensure that the cursor can move smoothly and accurately following the user's operation intention. However, in the related art, due to the lack of precise definition and recognition mechanism of the operation state, significant defects are shown in processing such fine operations.

[0036] Specifically, when the user's finger is slightly sliding on the plane, these subtle actions should be accurately captured and responded, but in actual use, often misjudge the unintentional signals (such as slight muscle tremor, posture adjustment, etc.) as valid input instructions, causing the cursor to move unexpectedly without clear operation intention. In addition, when the user's finger stops sliding or lifts from the plane, the cursor movement should be terminated immediately. But in the related art, often cannot identify this state change in time, causing the cursor to continue to "drift" or "jitter", further affecting the accuracy of the operation and the user's interaction experience. This results in the user needing to frequently adjust and correct the cursor position when processing documents, drawing or performing other tasks that require high precision, which not only increases the complexity of the operation, but also reduces the work efficiency.

[0037] It should be noted that the defects of the above prior art solutions are the result of careful research and practice, therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present application to solve the above problems should be considered as contributions to the present application in the process of invention, and should not be understood as technical content known to those skilled in the art.

[0038] Based on the discovery of the above technical problems, the present embodiment provides a gesture-based interaction method. As shown in the following Figure 1 The method comprises:

[0039] S1, acquiring motion information generated by a gesture action of a hand.

[0040] S2, determining an interaction state of the hand according to the motion information.

[0041] S3, generating a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information.

[0042] In this way, before generating the target interaction instruction, the current interaction state of the user's hand is first determined. Since the interaction state can represent whether the user has a clear interaction intention, and accurate interaction instructions are generated accordingly, the accuracy and response efficiency of the interaction are significantly improved.

[0043] It should be understood that the gesture-based interaction method provided by the present embodiment is applicable to various scenarios and devices that require non-contact or low-contact human-computer interaction. Exemplarily, the method can be widely applied to extended reality (Extended Reality, XR) devices, such as virtual reality (Virtual Reality, VR), augmented reality (Augmented Reality, AR) and mixed reality (Mixed Reality, MR) headsets, for realizing cursor control, menu selection, gesture operation and other interaction behaviors of the user in a three-dimensional virtual space.

[0044] In addition, the method can also be applied to intelligent wearable devices, such as smart watches and wristband health monitoring devices, to achieve quick control of device functions through gesture recognition, such as answering the phone, switching music, checking notifications, etc.

[0045] In the field of intelligent cockpits, the method can be integrated into a vehicle infotainment system or a driving assistance system to allow the driver to control functions such as navigation, air conditioning, and volume adjustment through gesture operation, thereby improving driving safety and operational convenience.

[0046] At the same time, the method can also be used in portable office equipment, such as tablets or two-in-one laptops, to allow users to achieve functions such as virtual mouse control, quick text input, or handwritten note input through gestures in a mobile environment.

[0047] To make the scheme provided in this embodiment clearer, the following describes each step of the method shown in the flowchart in detail, with a VR device as an electronic device implementing the gesture-based interaction method. Figure 1 It should be understood that the operations of the flowchart can not be implemented in sequence, and steps without logical context relationships can be reversed in order or implemented simultaneously. In addition, one or more other operations can be added to the flowchart or one or more operations can be removed from the flowchart under the guidance of the content of the present application. Continuing to refer to Figure 1 , the method comprises:

[0048] S1, acquiring motion information generated by a hand gesture action of a hand.

[0049] It should be understood that the gesture action refers to a physical action performed by the user through the hand to express a specific interaction intention, which can include but is not limited to local actions such as bending, stretching, sliding, tapping, and gripping of the fingers, and overall actions such as rotation, pitch, and swing of the wrist. Specifically, these gesture actions can be single-finger operations (such as index finger sliding), multi-finger cooperative operations (such as double-finger scrolling, triple-finger switching, and five-finger gripping), and displacement and posture changes of the hand as a whole, which are suitable for various interaction scenarios, such as cursor control, click operation, page scrolling, interface zooming, and execution of system-level shortcut instructions.

[0050] In this embodiment, the motion information can include electromyographic signals and acceleration signals generated by the gesture action. As Figure 2 shown, in this embodiment, the motion information can be collected by a signal collection device 11 worn on the wrist position of the user's hand. As an optional implementation, the signal collection device 11 can be a wearable device in the form of a bracelet or a watch, which is worn close to the skin surface above the wrist joint of the user's forearm, thereby achieving high-fidelity collection of electromyographic signals and acceleration signals.

[0051] From the structural composition, the signal acquisition device 11 includes a plurality of functional modules, mainly including a multi-channel electromyogram (EMG) signal acquisition device array and an inertial measurement unit (IMU). Among them, the multi-channel electromyogram signal acquisition device 11 is used to detect the electrical signal generated in the muscle contraction process, so as to obtain the electromyographic activity information of the user; the inertial measurement unit is used to collect the acceleration and angular velocity data of the hand in the three-dimensional space, and is further used to calculate the attitude, detect the impact and vibration and other motion characteristics.

[0052] In actual application process, the inertial measurement unit can use integrated six-axis or nine-axis inertial measurement unit with high sampling rate (for example, 200Hz), which is used to collect the acceleration signal of the user's hand in the three-dimensional space, including linear acceleration and angular velocity data, and is further used to calculate the attitude of the wrist (such as pitch angle, roll angle), detect the impact and micro-vibration characteristics when the hand contacts with the plane, and extract the feature parameters related to the attitude stability.

[0053] In addition, the signal acquisition device 11 also has the ability to communicate with the AR device 12 worn by the user, and can transmit the collected electromyographic signal and acceleration signal to the AR device 12 in real time, analyze the gesture action of the user, and display the response result of the gesture action in the operation interface 13.

[0054] Based on the above implementation of the motion information of the gesture action, continue to refer to Figure 1 , next continue to explain the step S2 in Figure 1 :

[0055] S2, according to the motion information, determine the interaction state of the hand.

[0056] As an optional implementation, a neural network-based model can be used to identify the motion information, so as to determine the interaction state of the hand. Specifically, the AR device processes the motion features extracted from the motion information through a pre-trained feature classification model, and obtains the interaction state of the hand.

[0057] Exemplarily, the AR device first acquires motion information generated by the user's hand during the execution of the gesture action, which includes electromyographic signals and acceleration signals collected by the multi-channel electromyographic signal acquisition device and the inertial measurement unit, and then extracts electromyographic features and acceleration features after preprocessing. Among them, the electromyographic features reflect the activation intensity and dynamic changes of the forearm muscles when performing the gesture action, specifically including time domain features (such as RMS, IEMG), transient features (such as WAMP, ZC), and frequency domain features (such as energy ratio of specific frequency band); the acceleration features represent the posture, impact and vibration characteristics, and posture stability of the hand in space, specifically including attitude quaternion, peak value and kurtosis of linear acceleration signal, angular velocity variance, etc. The above features are kept synchronous in time, and are spliced into a high-dimensional multi-modal feature sequence through feature fusion operation, which is used as the input of the feature classification model.

[0058] In actual application, the pre-trained feature classification model receives the fused multi-modal feature sequence as input and outputs the current interaction state of the user's hand. Among them, the interaction state is one of the hovering state, the flat state and the plane operation state.

[0059] The hovering state, that is, the hand is hovering and not in contact with any plane, is characterized by large attitude change of the inertial measurement unit, smooth acceleration without impact, and electromyographic signal at a low baseline level;

[0060] The flat state, that is, the hand has contacted the plane (for example, the desktop) but is in a resting or standby state, is characterized by stable attitude after an initial contact impact detected by the inertial measurement unit, only slight jitter of acceleration, and slightly high electromyographic signal intensity without transient burst;

[0061] The plane operation state, that is, the hand actively performs interactive action on the plane, is characterized by stable attitude but with in-plane displacement of the inertial measurement unit, continuous acceleration change, dynamic and high-intensity electromyographic signal, and high correlation with specific action.

[0062] In addition, for the feature classification model provided in the embodiment, a gradient boosting decision tree (GBDT) model with higher computational efficiency or a lightweight temporal convolutional network (TCN) can be used for training to obtain, which can effectively capture the classification boundary formed based on the impact, vibration and muscle tension difference, thereby realizing high-precision recognition of the three states.

[0063] In the process of constructing the training data set, a high-speed camera and a force-sensitive film can be used as external signal acquisition devices to provide accurate ground truth labels. The high-speed camera is used to determine whether the hand is suspended in the air, thereby distinguishing between "suspended state" and "non-suspended state"; the force-sensitive film is used to further distinguish between "flat state" (slight contact pressure) and "flat operation state" (significant and dynamic pressure change), ensuring that the training data has high physical interpretability and classification accuracy.

[0064] The training data set constructed in this way not only contains rich signal features, but also has strong physically interpretable label information.

[0065] After data acquisition and labeling, a training device (e.g., a high-performance server) uses a supervised learning method to train a gradient boosting decision tree model or a time series convolutional network. During model training, the training device takes the multi-modal feature sequence in the training data set as the input sample and the corresponding three-state labels (suspended state, flat state, and flat input state) as the target output, and optimizes the model parameters by minimizing the classification error. After training, the obtained feature classification model can receive real-time multi-modal feature sequences and output the current interactive state of the hand.

[0066] As another optional implementation, it is found that in the suspended state, the extensor muscle group maintains low-intensity contraction to resist gravity, and the flexor muscle group contracts synchronously to stabilize the joint, so the extensor muscle energy is continuously higher than the flexor muscle energy; in the flat state, both the extensor and flexor muscles are relaxed, and the electromyography signal tends to be silent; and in the operation state, the flexor muscle group produces a transient and intense contraction, and its signal energy is much higher than that of the extensor muscle group. However, since the electromyography signals of "suspended" and "flat" states are weak and similar, and the "operation" state signal has a very short duration, it is difficult for traditional methods to accurately distinguish these states, thereby affecting the stability and accuracy of the interaction.

[0067] In view of this, the present embodiment performs differential and relationship analysis on the electromyography signals of specific antagonistic muscles that control hand posture and movement to reveal unique physiological activation patterns in different interactive states. Then, based on this analysis, a hierarchical decision mixed state recognition system is constructed. The state recognition system first makes a quick and preliminary judgment, and then refers the ambiguous signal segments to a deep learning-based model for arbitration. Based on this concept, the motion information includes electromyography signals, which are divided into multiple signal segments, and the following implementation of step S2 is provided based on the multiple signal segments:

[0068] S2-1, sequentially selects a signal segment that has not been analyzed as a target signal segment from the multiple signal segments.

[0069] S2-2, obtaining energy information of the specified muscle according to the target signal segment.

[0070] During the execution of the above steps, the AR device selects an unanalyzed target signal segment from the divided multiple electromyography signal segments, which is derived from the electromyography signals collected by the multi-channel electromyography signal collection device array deployed at the specific position of the forearm. It should be noted that the multi-channel electromyography signal collection device array includes an extensor channel group and a flexor channel group. The electrodes of the extensor channel group are deployed on the dorsal side of the forearm, mainly covering the extensor carpi ulnaris, which is responsible for extending the fingers and wrist and is the main muscle for maintaining the suspended state of the hand. The flexor channel group is deployed on the palmar side of the forearm, mainly covering the flexor digitorum superficialis, which is responsible for flexing the fingers and is the main muscle for performing the point press and other downward actions. This electrode deployment method ensures the targeted collection of electromyography signals of the antagonistic muscle groups to obtain energy information of the specified muscle, including extensor energy, flexor energy, and flexor energy gradient.

[0071] In addition, it should be noted that the sampling frequency of the electromyography signal is not less than 1000 Hz to ensure the time resolution of the signal. After high-pass filtering (20 Hz) to remove low-frequency noise and notch filtering (50 / 60 Hz) to eliminate power frequency interference, the original signal is further divided into signal segments with overlapping regions by sliding window, for example, the window length is 50-200 milliseconds, and the sliding step is 20-50 milliseconds. In this way, continuous and fine time segment analysis of electromyography signals is ensured, while avoiding missing key state transition information due to window discontinuity.

[0072] On this basis, for each target signal segment, the AR device extracts the extensor energy (E ext) and the flexor energy (E flex). The extensor energy is defined as the average or sum of the root mean square (RMS) values of all signals of the extensor channel group, reflecting the overall activation intensity of the muscle group in the current time period. The flexor energy is the average or sum of the RMS values of the flexor channel group signals, representing the activation level of the flexor muscle group. In addition, the AR device further extracts the flexor energy gradient (Grad flex), which is the difference between the current window and the previous window flexor energy, used to capture the dynamic trend of muscle activity.

[0073] It should be understood that these energy information do not simply reflect the strength of muscle activity, but reveal the physiological differences in muscle contraction patterns under different interaction states based on the relative activation relationship between the antagonistic muscle groups. For example, in the suspended state, the extensor energy is usually higher than the flexor energy, and both are maintained at a certain baseline level; in the flat state, both energies decrease significantly and tend to be silent; and in the operating state, the flexor energy rises sharply and far exceeds the extensor energy.

[0074] Based on the above description of the energy information in the embodiments, step S2 further comprises:

[0075] S2-3, determining whether the interaction state of the hand can be determined through the energy information; if not, performing step S2-4, and if yes, obtaining the interaction state of the hand according to the energy information.

[0076] It should be understood that the interaction state is one of the hovering state, the flat state and the planar operation state. Therefore, as an optional embodiment, the AR device can determine whether the flexor energy gradient is greater than the burst threshold value and whether the ratio of the flexor energy to the extensor energy is greater than the flexion-extension ratio threshold value; if yes, it is determined that the hand is in the planar operation state; if not, it is determined whether the extensor energy and the flexor energy are less than the rest threshold value and maintained for a preset time length; if yes, it is determined that the hand is in the flat state; and if not, it is determined that the interaction state of the hand cannot be determined through the energy information.

[0077] It can be understood that, by setting a set of threshold values with clear physical meaning, the embodiment quickly identifies the interaction state with significant characteristics, thereby realizing efficient preliminary screening of the hand state. It should be noted that the determination logic depends on the characteristic parameters such as the extensor energy (E ext), the flexor energy (E flex) and the flexor energy gradient (Grad flex), and combines the preset burst threshold value (Th burst), the rest threshold value (Th rest) and the flexion-extension ratio threshold value (Th ratio) for state classification.

[0078] Based on the above parameters, during the execution of the above steps, the AR device calculates the flexor energy gradient and the flexor energy / extensor energy ratio based on the target signal segment, and determines whether the flexor energy gradient is greater than the burst threshold value and whether the ratio exceeds the flexion-extension ratio threshold value. If both conditions are met, the AR device determines that the current signal segment corresponds to the planar operation state. For this determination logic, it should be understood that when performing point press, sliding and other depression actions, the flexor muscle group will contract briefly and intensely, causing the flexor energy to rise rapidly and be significantly higher than the extensor energy, which is characterized by high energy gradient and high flexion-extension ratio. This feature is highly identifiable, and therefore can be quickly identified through simple threshold comparison.

[0079] If the above conditions are not met, the AR device further determines whether the extensor energy and the flexor energy are both lower than the resting threshold, and whether the low energy state lasts for a preset duration. If the conditions are met, the AR device determines that the hand is in a placed state. For this judgment logic, it should be understood that in the placed state, because the external support force replaces the active contraction of the muscle, the extensor and the flexor are both in a relaxed state, and the electromyographic energy is significantly reduced and tends to the baseline level. By setting the resting threshold and combining the duration requirement, the misjudgment caused by temporary signal fluctuations can be effectively excluded, and the stability of the judgment is improved.

[0080] If the above two conditions are not met, the AR device determines that the current energy information is not sufficient to determine the interaction state of the hand, and at this time the signal segment is marked as an uncertain state (PRE_UNCERTAIN), and further judgment needs to be made with the help of a signal classification model. In this way, when facing ambiguous signals, the AR device does not make an incorrect judgment directly, but arbitrates through the introduction of a deep model, thereby improving the accuracy of state recognition as a whole.

[0081] S2-4, processing the target signal segment by a pre-trained signal classification model to obtain a preliminary state of the hand.

[0082] During the execution of the above steps, the model input is the target signal segment from the information collected by the multi-channel electromyographic signal collection device array, which is in the form of a two-dimensional matrix with dimensions [channel number, sampling point number in window]. The channel number corresponds to the number of electrodes deployed in the forearm extensor and flexor muscle regions, and the sampling point number is determined by the window length (such as 50-200 milliseconds) and the sampling frequency (not less than 1000 Hz). In this way, the input data retains the timing characteristics and spatial distribution information between channels of the signal.

[0083] It should be understood that the signal classification model uses a lightweight one-dimensional convolutional neural network architecture, which includes two one-dimensional convolutional layers, one max-pooling layer, and one fully connected layer. Among them, the convolutional layer is used to extract local timing features in the target signal segment, the pooling layer is used to compress the feature dimension and enhance the translation invariance of the model, and the fully connected layer is responsible for mapping the extracted features to the output of state classification.

[0084] In addition, the signal classification model is trained offline on a large number of pre-collected and labeled electromyographic signal datasets, and the training samples cover the "hovering state", "placed state", and transition signals during state transition. After training, the model outputs the probability values of the current target signal segment belonging to the "hovering state" (PROB_AERIAL) and the "placed state" (PROB_PLACED). However, it should be noted that the probability value is only used as a preliminary judgment standard for the state of the hand, and needs to be further confirmed in combination with the recognition results of other signal segments.

[0085] S2-5, determine whether the preliminary state can determine the interaction state of the hand. If not, return to step S2-1 until the interaction state of the hand is determined. If yes, determine the interaction state of the hand according to the preliminary state.

[0086] It should be understood that if the hand of the user is in the plane operation state, obvious electromyographic signals will be generated. Therefore, in the embodiment, the preliminary state is one of the preliminary aerial state and the preliminary placed state, that is, only two states that are easily confused are distinguished.

[0087] Therefore, as an optional implementation, if the preliminary state is the preliminary aerial state, the AR device updates a first cumulative number of the plurality of signal segments determined as the preliminary aerial state. If the first cumulative number is greater than a first count threshold, it is determined that the hand is in the aerial state. If the first cumulative number is less than or equal to the first count threshold, it is determined that the preliminary state cannot represent the interaction state of the hand. If the preliminary state is the preliminary placed state, a second cumulative number of the plurality of signal segments determined as the preliminary placed state is updated. If the second cumulative number is greater than a second count threshold, it is determined that the hand is in the placed state. If the second cumulative number is less than or equal to the second count threshold, it is determined that the preliminary state cannot represent the interaction state of the hand.

[0088] During the execution of the above steps, the AR device receives the preliminary state output from the signal classification model, which includes two types of preliminary aerial state (PRE_AERIAL) and preliminary placed state (PRE_PLACED). When the AR device determines that the preliminary state of the current target signal segment is the preliminary aerial state, the first cumulative number (i.e., the cumulative number of times determined as the preliminary aerial state) is updated. If the cumulative number exceeds the first count threshold (for example, 2 times, corresponding to 100 ms), it is determined that the hand is currently in the aerial state. If the cumulative number does not exceed the threshold, it is considered that the current preliminary state is not sufficient to represent the true interaction state of the hand, and the state judgment will continue to be combined with subsequent signal segments.

[0089] Similarly, when the AR device determines that the preliminary state of the current target signal segment is the preliminary placed state, the second cumulative number (i.e., the cumulative number of times determined as the preliminary placed state) is updated. If the cumulative number exceeds the second count threshold (for example, 3 times, corresponding to 150 ms), it is determined that the hand is in the placed state. Otherwise, it is considered that the current state information is not sufficient to support the final judgment, and the AR device will continue to retain the current state and wait for the input of the subsequent signal segment.

[0090] It can be understood that the above embodiment is actually a state confirmation mechanism based on the cumulative number and the counting threshold. The AR device maintains a finite state machine for context awareness. The state machine is provided with two confidence counters for recording the number of segments continuously determined as the preliminary flat state and the preliminary suspension state. When the confidence of a certain state exceeds the set threshold, the state machine performs state switching and resets all counters to avoid false accumulation. For example, when the preliminary flat state is determined for three times in a row, the state machine switches the current state to the flat state; and when the preliminary suspension state is determined for two times in a row, the state machine switches to the suspension state.

[0091] Based on the interaction state obtained in the above embodiment, referring to Figure 1 , the step S3 in Figure 1 will be explained as follows:

[0092] S3, generating a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information.

[0093] It should be understood that in the present embodiment, not all motion information in all states generates an interaction instruction, but through the intention awareness mechanism of state classification, only when the user's explicit interaction intention is recognized (i.e., in the flat operation state), the instruction generation logic is executed, thereby effectively avoiding false triggering. The "suspension state" is defined as a state in which the user's hand is in the air and not in contact with any plane. From the perspective of interaction intention, this state usually does not correspond to an explicit interaction behavior, but is more likely to represent that the user is in a "non-interaction intention" or "hesitation state", i.e., has not decided whether to operate or is only performing free movement of the hand, adjusting posture, etc. non-interaction action.

[0094] Therefore, as an optional implementation, the present embodiment determines whether the interaction state is the suspension state before generating a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information.

[0095] In actual application, the AR device continuously monitors the state of the user's hand and actively suppresses the generation of interaction instructions when the "suspension state" is recognized. At this time, the AR device can only perform state tracking and gesture feature extraction, but will not convert these features into actual operation commands such as cursor movement, clicking, scrolling, etc. Only when the AR device detects that the user's hand enters the "flat operation state" and produces explicit sliding, clicking, gripping, etc. action, the gesture action and motion information are mapped to a specific target interaction instruction.

[0096] It is found in practice that the same action can also correspond to different operation intentions, especially in the case of fusion of multi-modal signals (such as electromyographic signals and inertial signals), and the manifestation of the action is more diverse. If the gesture type and motion trajectory are not explicitly identified, directly generating an interaction instruction will often result in the instruction not matching the user's true intention, for example, misjudging a swipe as a click, misjudging a two-finger scroll as a single-finger swipe, or producing a large cursor offset when the user only slightly moves the finger. In view of this, the following optional implementation of step S3 is also provided in the embodiment:

[0097] S3-1A, determining a gesture type and displacement information corresponding to the gesture action according to the interaction state and the motion information.

[0098] In the embodiment, the gesture translation model based on a neural network is used to convert the interaction state and the motion information into a specific gesture type and displacement information of the gesture. Therefore, the following optional implementation of step S3-1A is also provided in the embodiment:

[0099] S3-1A-1, extracting motion features of the hand from the motion information.

[0100] In the embodiment, the motion information includes electromyographic signals and acceleration signals; and the motion features include time domain features, transient features and frequency domain features extracted from the electromyographic signals, and hand posture features, impact features, vibration features and posture stability features extracted from the acceleration signals.

[0101] Exemplarily, the AR device performs filtering processing on the original electromyographic signals to remove environmental noise and power frequency interference, for example, by a band-pass filter (such as 20-450 Hz) to retain the effective frequency components related to muscle activity, and by a 50 / 60 Hz notch filter to suppress power grid interference, thereby obtaining clearer electromyographic signals that can be used for subsequent feature extraction.

[0102] Further, the AR device extracts time-domain features from the filtered electromyography signals, including root mean square (RMS) and integrated electromyography (IEMG), which reflect the overall activation intensity of the muscle and can be used as a reference for determining whether the user is actively exerting force. Next, to more finely distinguish between the "suspended state", "flat state" and "flat operation state" of the hand, the AR device also extracts transient features and frequency-domain features closely related to the dynamic changes of the signal. Among them, the transient features mainly include Willison amplitude (WAMP) and zero crossing rate (ZC), which can effectively capture the rapid contraction process of the muscle, and are particularly suitable for identifying actions such as clicking that have instantaneous force characteristics; while the frequency-domain features include specific frequency band energy ratios of power spectral density, which can be used to distinguish between continuous force states (such as finger sliding) and electromyography patterns in the resting state, and can enhance the recognition ability of the user's interactive intention.

[0103] Similarly, the AR device also obtains an acceleration signal from the inertial measurement unit, which contains linear acceleration information of the hand in three-dimensional space; then the acceleration signal and the gyroscope data are fused by a pose solving algorithm, and a Kalman filter or a complementary filter is used to calculate the hand pose features of the wrist in space, and the hand pose features are represented in the form of quaternions. It should be understood that the pose information is used to represent the current direction of the hand, which is used for further extracting the pose stability features.

[0104] Further, the AR device also performs high-pass filtering on the linear acceleration information after removing the gravity component to retain the high-frequency components related to the contact or tapping action of the hand; then, the peak value, kurtosis and energy of the signal are calculated in a short time window to obtain impact features and vibration features. Impact and vibration features can be used to identify whether the hand is in contact with the plane or produces a tapping action. Therefore, these features are the decisive basis for distinguishing between the "suspended state" and the "flat state" or the "flat operation state".

[0105] Based on the pose stability features obtained in the above embodiment, further pose stability features are extracted, especially the changes of the pitch angle and the roll angle. Specifically, the AR device judges whether the wrist is in a stable state by calculating the angular velocity variance corresponding to these angles. If the angular velocity variance is small, it indicates that the wrist is in a stable placement state; if the variance is large, it indicates that the hand may be in a free movement state in the air.

[0106] Based on the above description of the motion features of the embodiment, step S3-1A further includes:

[0107] S3-1A-2, the interaction state and motion features are processed by using a pre-trained gesture translation model to obtain gesture types and displacement information corresponding to the gesture actions.

[0108] Exemplarily, the input of the gesture translation model is a fused multi-modal feature sequence, which includes the feature information previously extracted from the electromyography signals and the acceleration signals, including electromyography features (such as RMS, IEMG, WAMP, ZC, etc.) and acceleration features (such as attitude quaternions, impact and vibration features, angular velocity variance, etc.), and maintains the time synchronization between each other. In addition, the multi-modal feature sequence also includes an interaction state, which represents an interaction intent as prior information for the gesture translation model to translate motion information, for enhancing the recognition accuracy of the gesture translation model on motion information.

[0109] In actual application, the gesture translation model can adopt a multi-task learning deep neural network architecture, for example, a model structure based on CNN-LSTM or Transformer, that is, containing a shared encoder and multiple parallel output heads, wherein the shared encoder is responsible for high-dimensional semantic modeling of the input multi-modal feature sequence, and the multiple output heads respectively undertake different task objectives to realize multi-dimensional analysis of gesture actions. The multiple output heads here include a trajectory regression head and a discrete action classification head.

[0110] The trajectory regression head is used to output two-dimensional displacement information , which includes the displacement direction and displacement information of the cursor or trackpad pointer. The model learns the mapping relationship from the changes of electromyography signals and acceleration signals during finger sliding to the actual plane displacement, thereby realizing the prediction of the moving trajectory of the user's hand on the plane. Therefore, the trajectory regression head ensures the continuity and stability of the cursor movement, avoiding the jitter or jumping phenomenon caused by misidentification or signal noise in traditional systems.

[0111] The discrete action classification head is used to output a multi-class action label, such as “index finger single click”, “index finger double click”, “middle finger single click”, “two-finger vertical sliding”, “three-finger horizontal sliding”, “five-finger grip”, etc. The output is used to identify the specific gesture type performed by the user, which can be used to realize advanced interaction functions (such as clicking, scrolling, zooming, system-level shortcut operation). In this way, by including multi-finger coordinated actions in the recognition range, it can support more rich interaction semantics than traditional gesture recognition systems, thereby meeting the needs of high-level human-computer interaction scenarios.

[0112] In terms of model training, the training data of the gesture translation model is only collected in the "flat operation state" to ensure that the training samples are highly consistent with the application scenarios of the model. Specifically, the input of the training sample is the fused multi-modal feature sequence, and the output includes two dimensions, namely the action category label and the displacement information label. The action category label is labeled synchronously by the actual gesture action of the user through an external signal collection device (such as a high-precision touch panel); the displacement information label can provide accurate ground truth through a synchronous digitizer or a high-resolution touch panel. Therefore, based on the joint labeling of real actions and physical displacements, the model can simultaneously learn the mapping relationship between gesture semantics and spatial trajectories, thereby realizing accurate analysis of gesture actions.

[0113] Based on the above description of gesture types and displacement information, step S3 further includes:

[0114] S3-2A, converting the gesture type and the displacement information into a target interaction instruction according to a first mapping rule.

[0115] In this embodiment, the target interaction instruction can be a keyboard and mouse operation instruction in a keyboard and mouse mode. Therefore, the AR device can map the gesture type to a keyboard and mouse operation instruction and take the displacement information as a command parameter of the keyboard and mouse operation instruction according to the first mapping rule.

[0116] For example, the AR device receives two key parameters output from the gesture translation model, namely the gesture type (Action_Label) and the displacement information . Among them, the gesture type is used to represent the specific action performed by the user, such as "single-finger sliding", "index finger single click", "double-finger vertical sliding", etc.; the displacement information is used to represent the continuous movement trajectory of the user's hand on the plane.

[0117] In actual application, the first mapping rule is a mapping table of gesture type to keyboard and mouse operation instruction. The AR device directly maps the recognized gesture type to the corresponding keyboard and mouse event according to the first mapping rule, and takes the displacement information as the parameter input of the event.

[0118] For example, as shown in Figure 3 , when the gesture type is "single-finger sliding", the AR device 12 generates a mouse movement instruction , where represents the two-dimensional displacement increment corresponding to the current gesture action, which is used to drive the movement of the cursor 14 in the operation interface 13.

[0119] As shown in Figure 4 , when the gesture type is "index finger single click", the AR device 12 generates a mouse left button single click instruction , for the position where the cursor 14 is located in the left-click operation interface 13; when the gesture type is "middle finger single click", the AR device 12 generates a mouse right-click instruction , for the position where the cursor 14 is located in the right-click operation interface 13.

[0120] Further, for actions involving multi-finger coordination, the AR device also translates them into more complex operation instructions according to the mapping rules. For example, the "double finger vertical sliding" gesture is mapped to a mouse wheel instruction , where represents the amplitude of the scrolling; the "three-finger horizontal sliding" is mapped to a keyboard shortcut instruction, such as , for switching virtual desktops; the "five-finger grip" gesture corresponds to the shortcut in the Windows system, for quickly displaying the desktop. Therefore, these mapping rules ensure that users can complete advanced function operations in the operating system through natural gesture actions, thereby improving interaction efficiency and convenience.

[0121] It is worth noting that the mapping rules are not static and fixed, but can be dynamically adjusted according to user preferences or application scenarios. For example, in some professional software, users can customize the mapping relationship between gesture types and keyboard / mouse instructions to adapt to specific workflows.

[0122] In this embodiment, in addition to mapping gesture actions to keyboard / mouse operation instructions in keyboard / mouse mode, related technologies can also map gesture actions to handwriting instructions in handwriting input mode. Specifically, related handwriting input methods usually rely on physical contact input devices (such as a stylus, a touch screen) to obtain handwriting trajectories, and their advantage is that they can directly determine the start and end of writing actions and stroke structures through information such as contact point pressure and position changes. However, for Chinese characters and other text systems with complex stroke structures, when users perform handwriting input through non-contact gesture interaction methods (such as based on wrist electromyography signals and inertial signals), there is a lack of clear "lift the pen, write the pen, pause the pen" state division mechanism, which makes it difficult to accurately distinguish the start, progress and end of writing actions, making it difficult for handwriting recognition engines to obtain complete writing dynamic information, thereby affecting the recognition accuracy of complex font structures, causing handwriting recognition errors, pen order logic confusion, and misjudgment of connected writing.

[0123] In view of this, the present embodiment also provides the following optional implementation of step S3:

[0124] S3-1B, determining the gesture type and displacement information corresponding to the gesture action according to the interaction state and motion information.

[0125] Similarly, the AR device can also utilize the pre-trained gesture translation model to process the interaction state and motion information, and obtain the gesture type and displacement information corresponding to the gesture action. For this, the present embodiment will not be described again, and the previous detailed description can be referred to.

[0126] S3-2B, according to the second mapping rule, converting the interaction state, gesture type and displacement information into handwriting instructions.

[0127] In the present embodiment, it has been introduced that the interaction state is one of the hovering state, the flat state and the plane operation state; and the handwriting instruction is one of the stroke instruction, the dwell instruction and the lift instruction. In the execution process of the above steps, if the interaction state is the plane operation state, the AR device generates the stroke instruction according to the gesture type and the displacement information; similarly, if the interaction state is the hovering state, the AR device generates the lift instruction; if the interaction state is the flat state, the AR device generates the dwell instruction.

[0128] It can be understood that the present embodiment realizes accurate recognition and expression of key action states such as “lift”, “stroke” and “dwell” in the writing process of Chinese characters by accurately mapping the interaction state and the handwriting instruction. The mapping mechanism not only improves the behavior recognition ability of the gesture interaction system in the handwriting input scene, but also provides higher quality input data for the handwriting recognition engine, thereby significantly improving the recognition accuracy of complex Chinese characters (such as running script and cursive script).

[0129] Exemplarily, as shown in Figure 5 when the interaction state classification result is the plane operation state, the AR device 12 generates the stroke instruction according to the gesture type (such as index finger sliding, middle finger clicking, etc.) and the displacement information The stroke instruction is used to indicate the actual drawing process of the stroke in the writing process, and the continuous displacement information output by the trajectory regression head is recorded as the handwriting trajectory, thereby forming the restoration of the stroke structure of Chinese characters.

[0130] If the interaction state is the hovering state, the AR device 12 generates the lift instruction, which is used to identify the action of lifting the hand away from the plane during the writing process to prepare for the next stroke. This state usually corresponds to the in-air movement between strokes and between characters in the writing of Chinese characters.

[0131] Further, when the interaction state is the flat state, the AR device 12 generates the dwell instruction, which represents the user's short pause during the writing process, or the preparation state when starting to write or putting down the pen.

[0132] For example, assume that the user inputs the Chinese character "write" by handwriting. The AR device 12 can sequentially recognize the states of lifting the pen, moving the pen, and pausing the pen according to the changes in the hand movements, and generate corresponding control instructions respectively. When the user lifts the finger from the desktop to prepare for writing, the AR device 12 recognizes it as a suspended state and generates a pen-lifting instruction. When the user starts to slide the finger on the desktop for stroke writing, the AR device 12 recognizes it as a planar operation state and generates a pen-moving instruction to record the trajectory points. During the writing process, when there is a short pause or when preparing for the next stroke, the AR device 12 recognizes it as a flat state and generates a pen-pausing instruction. These instructions are used in the subsequent handwriting recognition engine to reconstruct the complete writing process, including the stroke order, start and end points, pause time, etc. After actual testing, it is found that especially when recognizing writing forms with coherent strokes and complex structures such as running script and cursive script, the recognition accuracy can be significantly improved, showing obvious advantages.

[0133] Based on the same inventive concept as the gesture-based interaction method provided in this embodiment, this embodiment further provides a gesture-based interaction device, which includes at least one software function module that can be stored in a memory in software form or固化 in an electronic device. The processor in the electronic device is used to execute the executable module stored in the memory. For example, the software function modules and computer programs included in the device. Please refer to Figure 6 , functionally divided, the device may include:

[0134] A motion information module 21, configured to obtain the motion information generated by the gesture actions of the hand;

[0135] An interaction state module 22, configured to determine the interaction state of the hand according to the motion information;

[0136] An instruction recognition module 23, configured to generate a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information.

[0137] In this embodiment, the motion information module 21 is used to implement Figure 1 the step S1 in Figure 1 , the interaction state module 22 is used to implement Figure 1 the step S2 in

[0138] Therefore, for the detailed descriptions of the above modules, reference can be made to the specific implementation manners of the corresponding steps.

[0139] Optionally, the interaction state is one of a hovering state, a flat state, and a plane operation state; the instruction recognition module 23 is further configured to:

[0140] If the interaction state is the plane operation state, the target interaction instruction corresponding to the gesture action is generated according to the interaction state and the motion information.

[0141] Optionally, the instruction recognition module 23 is further configured to:

[0142] determine a gesture type and displacement information corresponding to the gesture action according to the interaction state and the motion information;

[0143] convert the gesture type and the displacement information into the target interaction instruction according to a first mapping rule.

[0144] Optionally, the instruction recognition module 23 is further configured to:

[0145] extract motion features of the hand from the motion information;

[0146] process the interaction state and the motion features by using a pre-trained gesture translation model to obtain the gesture type and the displacement information corresponding to the gesture action.

[0147] Optionally, the motion information includes electromyography signals and acceleration signals.

[0148] The motion features include time domain features, transient features, and frequency domain features extracted from the electromyography signals, and hand posture features, impact features, vibration features, and posture stability features extracted from the acceleration signals.

[0149] Optionally, a signal collection device is worn at a wrist position of the hand.

[0150] The electromyography signals and the acceleration signals are collected by the signal collection device.

[0151] Optionally, the target interaction instruction is a keyboard and mouse operation instruction in a keyboard and mouse mode, and the instruction recognition module 23 is further configured to:

[0152] map the gesture type into the keyboard and mouse operation instruction according to the first mapping rule, and take the displacement information as an instruction parameter of the keyboard and mouse operation instruction.

[0153] Optionally, the target interaction instruction is a handwriting instruction in a handwriting input mode, and the instruction recognition module 23 is further configured to:

[0154] determine a gesture type and displacement information corresponding to the gesture action according to the interaction state and the motion information;

[0155] According to the second mapping rule, the interaction state, the gesture type, and the displacement information are converted into the handwriting instruction.

[0156] Optionally, the interaction state is one of a hovering state, a flat state, and a plane operation state, and the handwriting instruction is one of a stroke instruction, a dwell instruction, and a lift instruction, and according to the second mapping rule, the instruction recognition module 23 is further configured to:

[0157] if the interaction state is the plane operation state, generating the stroke instruction according to the gesture type and the displacement information;

[0158] if the interaction state is the hovering state, generating the lift instruction;

[0159] if the interaction state is the flat state, generating the dwell instruction.

[0160] Optionally, the motion information includes an electromyography signal, the electromyography signal is divided into a plurality of signal segments, and the interaction state module is further configured to:

[0161] sequentially selecting an unanalyzed signal segment from the plurality of signal segments as a target signal segment;

[0162] obtaining energy information of a specified muscle according to the target signal segment;

[0163] determining whether the energy information can determine the interaction state in which the hand is located;

[0164] if not, processing the target signal segment by a pre-trained signal classification model to obtain a preliminary state of the hand;

[0165] determining whether the preliminary state can determine the interaction state in which the hand is located;

[0166] if not, returning to sequentially selecting an unanalyzed signal segment from the plurality of signal segments as a target signal segment until the interaction state in which the hand is located is determined.

[0167] Optionally, the interaction state is one of a hovering state, a flat state, and a plane operation state, and the energy information of the specified muscle includes an extensor energy, a flexor energy, and a flexor energy gradient; and the interaction state module is further configured to:

[0168] determining whether the flexor energy gradient is greater than a burst threshold value and whether a ratio of the flexor energy to the extensor energy is greater than a flexion-extension ratio threshold value;

[0169] if yes, determining that the hand is in the plane operation state;

[0170] if not, determining whether the extensor energy and the flexor energy are less than a resting threshold value and maintained for a preset time length;

[0171] If yes, it is determined that the hand is in a flat state;

[0172] If no, it is determined that the hand is in a state of interaction cannot be determined by the energy information.

[0173] Optionally, the preliminary state is one of a preliminary hovering state and a preliminary flat state, and the state of interaction is one of a hovering state and a flat state.

[0174] The state of interaction module is further configured to:

[0175] If the preliminary state is the preliminary hovering state, a first cumulative number of the plurality of signal segments determined as the preliminary hovering state is updated.

[0176] If the first cumulative number is greater than a first count threshold, it is determined that the hand is in a hovering state.

[0177] If the first cumulative number is less than or equal to the first count threshold, it is determined that the preliminary state cannot represent the state of interaction of the hand.

[0178] If the preliminary state is the preliminary flat state, a second cumulative number of the plurality of signal segments determined as the preliminary flat state is updated.

[0179] If the second cumulative number is greater than a second count threshold, it is determined that the hand is in a flat state.

[0180] If the second cumulative number is less than or equal to the second count threshold, it is determined that the preliminary state cannot represent the state of interaction of the hand.

[0181] In addition, each functional module in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0182] It should also be understood that the above embodiments, if implemented in the form of software functional modules and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.

[0183] Therefore, the embodiment also provides a storage medium, which is a computer readable storage medium. The storage medium stores a computer program. The computer program is executed by a processor to implement the gesture-based interaction method provided by the embodiment. The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or any medium that can store program codes.

[0184] The embodiment provides an electronic device for implementing the gesture-based interaction method. As shown in Figure 7 The electronic device can include a processor 32 and a memory 31. The memory 31 stores a computer program. The processor reads and executes the computer program corresponding to the above embodiments in the memory 31 to implement the gesture-based interaction method provided by the embodiment.

[0185] Continuing to refer to Figure 7 The electronic device further includes a communication unit 33. The memory 31, the processor 32 and the communication unit 33 are directly or indirectly electrically connected to each other through a system bus 34 to realize data transmission or interaction.

[0186] The memory 31 can be an information recording device based on any electronic, magnetic, optical or other physical principle for recording execution instructions, data, etc. In some embodiments, the memory 31 can be, but is not limited to, a volatile memory, a non-volatile memory, a storage drive, etc.

[0187] In some embodiments, the volatile memory can be a random access memory (RAM); in some embodiments, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, etc.; in some embodiments, the storage drive can be a magnetic disk drive, a solid state disk, any type of storage disk (such as an optical disk, a DVD, etc.), or a similar storage medium, or a combination thereof, etc.

[0188] The communication unit 33 is configured to transmit and receive data via a network. In some embodiments, the network can include a wired network, a wireless network, a fiber optic network, a telecommunications network, an intranet, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Network (WLAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a Public Switched Telephone Network (PSTN), a Bluetooth network, a ZigBee network, or a Near Field Communication (NFC) network, etc., or any combination thereof. In some embodiments, the network can include one or more network access points. For example, the network can include wired or wireless network access points, such as base stations and / or network switching nodes, through which one or more components of the service request processing system can connect to the network to exchange data and / or information.

[0189] The processor 32 can be an integrated circuit chip that has the ability to process signals and can include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the processor can include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC), or a microprocessor, etc., or any combination thereof.

[0190] It can be understood that, Figure 7The illustrated structure is merely schematic. The electronic device can also have more or less components than shown, or a different configuration or arrangement of the components. The various components shown in the figure can be implemented, individually and / or collectively, by a wide range of hardware structures or Figure 7 software structures. Figure 7 Figure 7 The size, shape and relative placement of the various components shown in the figure are not necessarily drawn to scale unless explicitly disclaimed.

[0191] It should be understood that the devices and methods disclosed in the above embodiments can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show possible architectural, functional and operational scenarios of devices, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that shown in the figure. For example, two consecutive blocks can actually be executed substantially in parallel, and they can also be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0192] The above description is merely illustrative of the various embodiments of the present application. It is not intended to limit the scope of the application. Any variations and modifications of the embodiments disclosed herein, which do not depart from the spirit and scope of the application, are intended to be within the scope of the application. Thus, to the maximum extent possible, the broadest scope of the application is to be determined by the following claims.​

Claims

1. A gesture-based interaction method, characterized in that, The method comprises: obtaining motion information generated by a gesture action of a hand; determining an interaction state in which the hand is located according to the motion information, wherein the interaction state is one of a hovering state, a flat state and a plane operation state; generating a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information, comprising: extracting motion features of the hand from the motion information; processing the interaction state and the motion features by using a pre-trained gesture translation model to obtain a gesture type and displacement information corresponding to the gesture action, wherein the interaction state represents an interaction intention as prior information for the gesture translation model to translate the motion information; converting the gesture type and the displacement information into the target interaction instruction according to a first mapping rule.

2. The gesture-based interaction method of claim 1, wherein, generating a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information, comprising: if the interaction state is a plane operation state, generating a target interaction instruction corresponding to the gesture action according to the interaction state and the motion information.

3. The gesture-based interaction method of claim 1, wherein, The motion information comprises electromyography signals and acceleration signals; The motion features comprise time domain features, transient features and frequency domain features extracted from the electromyography signals, and hand posture features, impact features, vibration features and posture stability features extracted from the acceleration signals.

4. The gesture-based interaction method of claim 3, wherein, A signal acquisition device is worn at a wrist position of the hand; The electromyography signals and the acceleration signals are acquired by the signal acquisition device.

5. The gesture-based interaction method of claim 1, wherein, The target interaction instruction is a keyboard and mouse operation instruction in a keyboard and mouse mode; Converting the gesture type and the displacement information into the target interaction instruction according to a first mapping rule, comprising: mapping the gesture type into the keyboard and mouse operation instruction according to the first mapping rule, and taking the displacement information as an instruction parameter of the keyboard and mouse operation instruction.

6. The gesture-based interaction method of claim 1, wherein, The motion information comprises electromyography signals, the electromyography signals are divided into a plurality of signal segments, and determining the interaction state in which the hand is located according to the motion information comprises: selecting a target signal segment that has not been analyzed from the plurality of signal segments in sequence; obtaining energy information of a specified muscle according to the target signal segment; determining whether the energy information can be used to determine the interaction state in which the hand is located; if not, processing the target signal segment by using a pre-trained signal classification model to obtain a preliminary state of the hand; determining whether the preliminary state can be used to determine the interaction state in which the hand is located; if not, returning to select a target signal segment that has not been analyzed from the plurality of signal segments in sequence until the interaction state in which the hand is located is determined.

7. The gesture-based interaction method of claim 6, wherein, The interaction state is one of a hovering state, a flat state and a plane operation state, and the energy information of the specified muscle comprises extensor muscle energy, flexor muscle energy and flexor muscle energy gradient; determining whether the energy information can be used to determine the interaction state in which the hand is located, comprising: determining whether the flexor energy gradient is greater than a burst threshold and whether a ratio of the flexor energy to the extensor energy is greater than a flexion-extension ratio threshold; if yes, determining that the hand is in the planar operation state; if no, determining whether the extensor energy and the flexor energy are less than a rest threshold and remain for a preset time length; if yes, determining that the hand is in the flat state; if no, determining that the hand gesture cannot be determined by the energy information.

8. The gesture-based interaction method of claim 6, wherein, the preliminary state is one of a preliminary hovering state and a preliminary flat state, and the interaction state is one of a hovering state and a flat state; determining whether the preliminary state can determine the interaction state of the hand, including: if the preliminary state is the preliminary hovering state, updating a first cumulative number of the plurality of signal segments determined as the preliminary hovering state; if the first cumulative number is greater than a first count threshold, determining that the hand is in the hovering state; if the first cumulative number is less than or equal to the first count threshold, determining that the preliminary state cannot represent the interaction state of the hand; if the preliminary state is the preliminary flat state, updating a second cumulative number of the plurality of signal segments determined as the preliminary flat state; if the second cumulative number is greater than a second count threshold, determining that the hand is in the flat state; if the second cumulative number is less than or equal to the second count threshold, determining that the preliminary state cannot represent the interaction state of the hand.

9. A gesture-based interaction device, characterized by The device includes: a motion information module configured to obtain motion information generated by a hand gesture action of a hand; an interaction state module configured to determine an interaction state of the hand according to the motion information, wherein the interaction state is one of a hovering state, a flat state, and a planar operation state; an instruction recognition module configured to generate a target interaction instruction corresponding to the hand gesture action according to the interaction state and the motion information, including: extracting motion features of the hand from the motion information; processing the interaction state and the motion features by using a pre-trained gesture translation model to obtain a gesture type and displacement information corresponding to the hand gesture action, wherein the interaction state represents an interaction intent as prior information for the gesture translation model to translate the motion information; converting the gesture type and the displacement information into the target interaction instruction according to a first mapping rule.

10. A storage medium, characterized by The storage medium stores a computer program, and the computer program, when executed by a processor, implements the gesture-based interaction method of any one of claims 1-8.

11. An electronic device, comprising: The electronic device includes a processor and a memory, and the memory stores a computer program, and the computer program, when executed by the processor, implements the gesture-based interaction method of any one of claims 1-8.

Citation Information

Patent Citations

  • Virtual reality robot interaction system and interaction method based on human body natural signals

    CN110119207A