Sonomyography-based interface control system and method
Patent Information
- Application Number
- PCT/CN2025/080837
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-05
- Filing Date
- 2025-03-05
- Publication Date
- 2025-10-02
AI Technical Summary
Traditional human-computer interface systems such as keyboards and mice have limitations. The electromyography method is highly invasive, the electrode position is unstable, and it is difficult to identify deep muscle movements. Surface electromyography is easily affected by artifacts and it is difficult to achieve fine gesture recognition.
An interface control system based on acoustic myography is adopted. Ultrasonic transducers arranged on the fingers and/or forearms collect muscle movement information, and artificial neural networks are combined to recognize gestures to achieve interface control.
It provides higher-precision muscle activity recognition, can distinguish deep muscle movements, supports fine gesture control, reduces the limitations of traditional methods, and improves the naturalness and accuracy of human-computer interaction.
Smart Images

Figure CN2025080837_02102025_PF_FP_ABST
Abstract
Description
Interface control system and method based on acoustic myogram Technical Field
[0001] The present disclosure relates to the technical field of human-machine interface (HMI), and in particular to an interface control system and method based on sonomyography (SMG). Background Art
[0002] Biomedical signals are being used to overcome the limitations of traditional HMI systems, such as keyboards and mice. Human arm muscle movements can be translated into natural computer control commands. Currently, much research is focused on wearable devices that can capture arm and finger muscle movements and recognize them as different gestures. Traditional gesture recognition methods include those based on electromyography (EMG). Because muscle contraction generates a weak electrical current, sensors attached to appropriate locations on the skin can measure this current in surface muscles. EMG is a graph of current intensity over time. Intramuscular electromyography (iEMG) is an invasive method that can be used to capture muscle activity in deeper muscles under various circumstances. However, it has drawbacks such as invasiveness, unstable electrode placement, and small recording volume. Surface electromyography (sEMG) captures electrical signals from muscles using sensors attached to the skin's surface. sEMG is non-invasive, but suffers from crosstalk with surrounding muscles and cannot effectively detect movement in deeper muscles. EMG is easily affected by artifacts and background activity, making it difficult to detect subtle finger movements and gestures.
[0003] Muscle contraction causes muscle movement, leading to changes in muscle structure and morphology. When different fingers flex or extend, a unique set of intermuscular septa are activated and mechanically deformed. Compared to EMG, ultrasound imaging (US), used to sense mechanical deformation of functional intermuscular septa, can distinguish deep, continuous intermuscular septa and can obtain robust, graded signals with low signal-to-noise ratios. Compared to EMG, US requires a much smaller probe size for image acquisition. Ultrasound is sound wave frequency exceeding 20 kHz. Ultrasound devices utilize the propagation and reflection of ultrasound pulses to non-invasively detect and measure objects, providing structural information. Ultrasound pulses can be reflected at various depths within the body. The delay of the reflected wave is proportional to the depth, while its intensity indicates the tissue type because different tissues have characteristic acoustic impedances. US can detect, for example, muscle thickness, muscle fiber length, pennation angle, and cross-sectional area. US not only provides static images of anatomical structures but also real-time, dynamic images of internal tissue movement associated with physical and physiological activity. The dynamic ultrasound images and the resulting dynamic signals obtained when ultrasound scans muscles are called SMG. SMG can better distinguish individual muscle activity and provide control signals proportional to muscle deformation. SMG can spatially resolve individual muscles with submillimeter accuracy.
[0004] This disclosure provides several novel SMG-based interface control systems. These systems compute features of SMG acquired from gloves and / or forearm wristbands via US and deploy artificial neural networks (ANNs) to recognize different gestures, finger combinations, and fine finger movements, thereby controlling interfaces on different media. Summary of the Invention
[0005] The present disclosure provides an interface control system based on sonomyogram (SMG), comprising: an SMG data acquisition device, the SMG data acquisition device comprising a plurality of ultrasonic transducers UT arranged on fingers and / or forearms and at least one data acquisition module; an SMG data processing device, the SMG data processing device processing the SMG data acquired by the SMG data acquisition device to obtain image data that is conducive to identifying predetermined target muscles; a feature extraction and selection device, the feature extraction and selection device extracting features from the morphology and / or movement pattern of the target muscles and selecting features that are conducive to identifying predefined target gestures; a classification device, the classification device classifying the selected features based on the target gesture to identify the correct gesture; and an interface control device, the interface control device interpreting the gestures recognized by the classification device as control commands for an interface on a medium based on a control algorithm deployed locally or online, thereby realizing human-computer interaction (HMI).
[0006] In some embodiments, the SMG data acquisition device is a wearable device, wherein the SMG data acquisition device is a wearable glove when arranged on a finger, the SMG data acquisition device is a wearable forearm band when arranged on a forearm, and the SMG data acquisition device is a combination of a glove and a forearm band when arranged on both the finger and the forearm, wherein the SMG data acquisition device is a left-hand glove, a right-hand glove, a left forearm band, a right forearm band, or any combination thereof.
[0007] In some embodiments, the SMG data acquisition device uses one or more of the following ultrasound modes: A-mode ultrasound, B-mode ultrasound, and M-mode ultrasound, and the acquired or reconstructed ultrasound dimensions include one or more of the following: one-dimensional, two-dimensional, or three-dimensional ultrasound.
[0008] In some embodiments, the SMG data processing device identifies the target muscle by segmenting the muscle aponeurosis in the SMG data and reconstructing the muscle bundles in the region.
[0009] In some embodiments, the SMG data processing device uses one or more of the following methods: time gain compensation (TGC), Gaussian filtering, Hilbert transform, logarithmic compression, and registration to generate the deformation field.
[0010] In some embodiments, when the SMG data acquisition device is a glove, the morphology and / or motion pattern of the target muscle is obtained from at least one or more of the following parameters: tissue displacement, length, width, area, angle, and grayscale gradient, as well as the magnitude and speed of changes in the above parameters over time.
[0011] In some embodiments, when the SMG data acquisition device is a forearm wristband, for the three-dimensional SMG reconstructed from multiple UTs in the forearm wristband, a three-dimensional model of the target muscle is generated by segmenting the muscle in each cross-sectional slice, and the extracted two-dimensional SMG features are spliced together according to the spatial distribution of the muscle to obtain an SMG feature map, wherein both the three-dimensional model of the target muscle and the SMG feature map can be used as inputs to the AI model.
[0012] In some embodiments, the feature extraction and selection device adopts one or more of the following methods: linear fitting, linear regression (LR), image edge detection, differential operator, high-pass filtering and discrete wavelet transform, wherein the classification device adopts one or more of the following methods: support vector machine SVM, back propagation BP neural network, linear discriminant analysis LDA, K nearest neighbor KNN algorithm, multi-layer perceptron MLP, stacked denoising autoencoder SDA, naive Bayes classifier and decision tree-based classifier.
[0013] In some embodiments, the control command is a custom sign language based on the target gesture, wherein the interface control device searches for the target sign language that matches the recognized gesture in a predefined sign language dictionary, and outputs a value corresponding to the target sign language when the match is successful. When the match is unsuccessful, the interface control device supports updating the sign language dictionary.
[0014] In some embodiments, the target gesture is a binary expression based on relaxation and bending of different fingers, wherein the finger states of the binary expression include: relaxation, bending thumb, bending index finger, bending middle finger, bending ring finger, bending little finger and making a fist.
[0015] In some embodiments, for the finger states expressed in binary, typing of all characters and function keys is achieved on ten keys by combining the bending states, bending times, and bending times of the ten fingers.
[0016] In some embodiments, for the finger state expressed in binary, the rows and columns on the custom key table are indexed respectively through the gestures of the left hand and the gestures of the right hand to find the corresponding keys at the intersection of the rows and columns, thereby realizing the typing of full characters and function keys on the custom key table.
[0017] In some embodiments, for the finger state of binary expression, one or more keyboard configurations are switched by making a fist with one hand while bending the fingers of the other hand. The keyboard configurations include one or more of the following: English keyboard, numeric keyboard, symbol keyboard, other language keyboards and user-defined shortcut keys.
[0018] In some embodiments, the interface control system also includes an interface display device, which presents an interface with a cursor, a finger pointer and / or a virtual keyboard to provide typing guidance, mouse clicks, cursor control and / or keyboard switching, and provides corresponding visual feedback when the user's gesture recognition is successful. The interface display device is one or more of the following: a display, a projection interface, augmented reality AR and mixed reality MR.
[0019] In some embodiments, the SMG data acquisition device includes a combination of gloves for both hands and a forearm wristband, which realizes mouse clicks, cursor movement and full keyboard input by collecting fine movements of the wrist and fingers, and provides real-time visual feedback on cursor movement and finger movement.
[0020] In some embodiments, the interface display device switches between the following three modes through gestures: virtual keyboard mode, touchpad mode, and mouse mode.
[0021] The present disclosure also discloses an interface control method based on sonomyogram (SMG), comprising: collecting SMG data through a plurality of ultrasonic transducers UT and at least one data acquisition module arranged on the fingers and / or forearms; processing the SMG data to obtain image data that is conducive to identifying predefined target muscles, wherein the image data is ultrasonic image data of the muscles, as well as muscle motion image data, muscle hardness image data and ultrasonic image data after segmentation of different muscles obtained by processing these ultrasonic image data, and the image data can be one-dimensional, two-dimensional, three-dimensional, or multi-dimensional; extracting features of the morphology and / or motion pattern of the target muscles, and selecting features that are conducive to identifying predefined target gestures; classifying the selected features based on the target gesture to identify the correct gesture; and interpreting the identified gestures as control commands of the interface on the media based on a control algorithm deployed locally or online, thereby realizing human-computer interaction HMI.
[0022] Other features and advantages will become apparent from the following detailed description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The foregoing and further features of the present disclosure will become apparent from the following description of preferred embodiments provided by way of example only and taken in conjunction with the accompanying drawings, in which:
[0024] FIG1 shows a schematic diagram of a forearm wristband according to an embodiment of the present disclosure;
[0025] FIG2 shows a schematic diagram of target muscles of the forearm according to an embodiment of the present disclosure;
[0026] FIG3 shows a schematic diagram of a glove according to an embodiment of the present disclosure;
[0027] FIG4 shows a schematic diagram of target muscles of the hand according to an embodiment of the present disclosure;
[0028] FIG5 shows a schematic diagram of SMG data processing according to an embodiment of the present disclosure;
[0029] FIG6 shows a flowchart of gesture recognition based on SMG according to an embodiment of the present disclosure;
[0030] FIG7A shows an interface control system for performing sign language recognition using both hands according to an embodiment of the present disclosure;
[0031] FIG7B shows an interface control system for performing sign language recognition using a single hand according to an embodiment of the present disclosure;
[0032] FIG8 shows a schematic flow chart of sign language matching and insertion according to an embodiment of the present disclosure;
[0033] FIG9 is a schematic diagram showing a finger state expressed by binary according to an embodiment of the present disclosure;
[0034] FIG10 shows a schematic diagram of ten-key typing steps according to an embodiment of the present disclosure;
[0035] FIG11 shows a schematic diagram of a method for switching keyboard configurations according to an embodiment of the present disclosure;
[0036] FIG12 is a schematic diagram showing a ten-key typing guide interface on a display or projection media according to an embodiment of the present disclosure;
[0037] FIG13 shows a schematic diagram of a ten-key typing guide interface on an AR or MR media according to an embodiment of the present disclosure;
[0038] FIG14 shows a schematic diagram of the steps of typing with a key table according to an embodiment of the present disclosure;
[0039] FIG15 is a schematic diagram showing a key table typing guide interface on a display or projection media according to an embodiment of the present disclosure;
[0040] FIG16 shows a schematic diagram of a key table typing guide interface on an AR or MR media according to an embodiment of the present disclosure;
[0041] 17A and 17B are schematic diagrams showing an interface control system using a cursor and a virtual keyboard according to an embodiment of the present disclosure;
[0042] FIG18 shows a schematic diagram of mode switching according to an embodiment of the present disclosure;
[0043] FIG19 is a schematic diagram showing a display and a control cursor on a projection medium according to an embodiment of the present disclosure;
[0044] FIG20 shows a schematic diagram of an SMG-based interface control system according to an embodiment of the present disclosure;
[0045] FIG21 shows a schematic diagram of an exemplary interface control system based on SMG according to an embodiment of the present disclosure;
[0046] FIG22 shows a flowchart of an SMG-based interface control method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0047] First, some terms used in the embodiments of the present disclosure are explained: ANN: Artificial Neural Network. AR: Augmented Reality. BP: Backpropagation. FCR: Flexor carpi radialis. FCU: Flexor carpi ulnaris. FDP: Flexor digitorum profundus. FDS: Flexor digitorum superficialis. FPL: Flexor pollicis longus. iEMG: Intramuscular electromyography. KNN: K-Nearest Neighbors. LDA: Linear Discriminant Analysis. LR: Linear Regression. MLP: Multilayer Perceptron. MR: Mixed Reality. RF: Radio Frequency. SDA: Stacked Denoising Autoencoder. EMG: Electromyography. sEMG: Surface Electromyography. SVM: Support Vector Machine. TGC: Time Gain Compensation. UT: Ultrasound Transducer. US: Ultrasound Imaging.
[0048] As used herein, the terms "first," "second," and "third" are used interchangeably to distinguish one component from another, rather than to indicate the position or importance of a single component. The singular expressions "a," "an," and "the" also include the plural, unless the context clearly dictates otherwise. Terms such as "coupled," "fixed," and "connected to" refer to direct coupling, fixing, or connection, as well as indirect coupling, fixing, or connection through one or more intermediate components or features, unless the context clearly dictates otherwise. The terms "comprise," "include," "compose," "have," or any other variations thereof herein are intended to encompass non-exclusive inclusion. For example, a process, method, article, or device that includes a set of features is not necessarily limited to those features, but may include features not explicitly listed or other features inherent to such a process, method, article, or device. In addition, unless expressly stated to the contrary, "or" is inclusive rather than exclusive. For example, any of the following satisfies condition A or B: A is true (or exists) and B is false (or does not exist), A is false (or does not exist) and B is true (or exists), and both A and B are true (or exist).
[0049] Terms indicating approximation, such as "about," "approximately," "approximately," or "substantially," include values within 10% greater or less than the stated value. When used in the context of an angle or direction, these terms include values within 10 degrees greater or less than the stated angle or direction. For example, "approximately perpendicular" includes directions within 10 degrees of perpendicular in any direction (e.g., clockwise or counterclockwise).
[0050] It will be appreciated that, if any prior art publication is cited herein, such reference does not constitute an admission that the publication forms part of the common general knowledge in the art in any country.
[0051] The hand and fingers, as the most flexible parts of the forearm, can achieve complex interface control through different movement patterns and combinations. Therefore, to recognize different gestures or finger states, it is necessary to obtain rich motion information of the forearm muscle groups, and even the movement patterns of smaller muscles on the back of the hand to identify more detailed finger movements. Wearable devices are deployed at different locations on the forearm. Ultrasonic transducers (UTs) can be flexibly placed on the forearm to monitor muscle activity.
[0052] US is typically acquired using a UT probe. UT is a device equipped with one or more piezoelectric elements. These elements convert electrical pulses into mechanical pulses, thereby generating ultrasound waves. These elements can also convert the mechanical energy of the reflected pulses into electrical signals in the radio frequency (RF) range, using the envelope of the RF signal to construct an image. These phenomena are known as the inverse piezoelectric effect and the direct piezoelectric effect. Rapid switching of electronic components allows the generation and reading of ultrasound waves using the same probe. Modern ultrasound imaging devices use algorithms to form images and apply beamforming and harmonic imaging to improve resolution.
[0053] US has multiple modes, such as A-mode, B-mode, and M-mode. A-mode, also known as amplitude mode, displays the received echo signal in the form of amplitude. Individual UTs can be fixed at a selected location during scanning. A-mode ultrasound echo signals are one-dimensional, offering advantages such as flexible placement, compact size, low cost, short computational latency, and minimal data processing. B-mode, known as brightness mode, displays information reflected from interfaces at different depths in grayscale. During scanning, a probe consisting of a linear array of UTs (e.g., 64 or 128 UTs) is used to capture multiple A-mode scan lines. B-mode ultrasound reflects tissue status in two-dimensional images. B-mode ultrasound offers superior signal robustness to A-mode ultrasound. M-mode, also known as motion mode, captures tissue motion over time by collecting continuous A-mode scan signals. During M-mode ultrasound, ultrasound pulses are emitted rapidly and continuously. When the boundary of the object generating the reflection moves relative to the probe over time, M-mode ultrasound can determine its velocity. Compared to B-mode ultrasound, M-mode ultrasound has a higher scanning frequency and can provide more detailed information about muscle status than A-mode ultrasound. The envelope of the A-mode scan can be used to construct 2D B-mode and M-mode images. 2D B-mode or M-mode images can also be used to construct real-time 3D images based on algorithms. This disclosure focuses on systems and methods for interface control based on sonomyography. Therefore, as long as the US sensing device can be miniaturized, this disclosure is not limited to the application of any one or combination of the aforementioned modes.
[0054] FIG1 shows a schematic diagram of a forearm wristband 100 according to an embodiment of the present disclosure. Multiple UTs 110 are arranged side by side on the forearm wristband 100, generally along the longitudinal extension direction (i.e., along the length of the forearm), to generate a three-dimensional model of the arm muscle groups in that area (the area covered by the forearm wristband). The forearm wristband is used to monitor simple flexion and extension movements of the wrist and fingers. FIG1 shows a schematic arrangement of multiple UTs 110 and a data acquisition module 120. Exemplarily, each of the multiple UTs on the forearm wristband 100 can be electrically connected to the data acquisition module 120. It should be understood that the location and number of the multiple UTs 110 and the data acquisition module 120 can be varied as needed.
[0055] Figure 2 shows a schematic diagram of a forearm target muscle 200 according to an embodiment of the present disclosure. Multiple UTs 110 can monitor the movement of the forearm target muscles 200. The forearm target muscles 200 may include the flexor carpi ulnaris (FCU) 210, the flexor carpi radialis (FCR) 220, the flexor digitorum superficialis (FDS) 230, the flexor digitorum profundus (FDP) 240, and the flexor pollicis longus (FPL) 250. The activity of the FCU and FCR can reflect wrist abduction and adduction. The FDP and FDS involve flexion of the index, middle, ring, and little fingers at the metacarpophalangeal, proximal, and distal interphalangeal joints. The FPL can reflect thumb flexion. It should be understood that the forearm target muscles 200 described above are merely preferred examples of the present disclosure and are not intended to limit the types of muscles that can be monitored. Certain gestures can trigger the movement of a range of muscles, and discrete gestures can be classified based on the changes in multiple muscles.
[0056] FIG3 shows a schematic diagram of a glove 300 according to an embodiment of the present disclosure. The glove 300 is capable of monitoring fine finger movements, including flexion and extension of the fingers, as well as left and right tilting of the fingers. The glove 300 has five UTs 310 arranged along the extension direction of the tendon 460 ( FIG4 ). For example, as shown in FIG4 , the five UTs 310 can be arranged along the UT direction to form an array. The UT direction refers to the arrangement direction of multiple UTs, and the UT arrangement direction is perpendicular to the UT arrangement direction. FIG3 shows a schematic arrangement of the five UTs 310 and the data acquisition module 320. For example, the five UTs 310 can all be electrically connected to the data acquisition module 320. It should be understood that the position and number of the UTs 310 and the data acquisition module 320 can be varied as needed.
[0057] FIG4 shows a schematic diagram of a hand target muscle 400 according to an embodiment of the present disclosure. The hand target muscles 400 include four lumbrical muscles 410, 420, 430, and 440 and the adductor pollicis 450. The lumbrical muscles 410-440 are involved in flexion of the metacarpophalangeal joints and extension of the interphalangeal joints. By measuring the first to fourth lumbrical muscles 410-440, fine motor movements of the little finger, ring finger, middle finger, and index finger can be monitored. The activity of the adductor pollicis 450 can reflect the adduction of the thumb.
[0058] The glove 300 and the forearm wristband 100 can be worn in different combinations on both arms to accommodate different application scenarios (e.g., single-handed and double-handed), usage modes (e.g., gesture control and keyboard and mouse control), or different types of users (e.g., healthy individuals and amputees). The data acquisition modules 120, 320 can be designed to be detachable. When the glove and wristband are used on the same side, one of the data acquisition modules can be removed to reduce weight. The image data collected by the UT is sent to the data acquisition module, and the data acquisition module 120, 320 can transmit the collected image data in real time to the terminal device or the cloud for subsequent processing.
[0059] Figure 5 illustrates a schematic diagram of SMG data processing according to an embodiment of the present disclosure. Muscle motion information conveyed by the SMG can be characterized in two ways. One approach is to directly use the SMG data acquired by the data acquisition modules 120 and 320 as input to the AI model. Another approach is to pre-process the SMG data from the data acquisition modules 120 and 320 before using it as input to the AI model. For example, the SMG data can first be segmented 510 to obtain a series of target muscle images 520, such as the forearm target muscles 200 and / or the hand target muscles 400. Subsequently, SMG feature calculation 530 is performed on one or more target muscle images 520 to obtain an SMG feature map 540. SMG data refers to real-time muscle ultrasound images. SMG features refer to quantitative features extracted from the SMG data. Since the SMG data acquired by the glove 300 is two-dimensional, two-dimensional SMG feature extraction can be performed on the SMG data. The extracted two-dimensional SMG features may include tissue displacement, length, width, area, angle, grayscale gradient, and / or the magnitude and / or velocity of changes in these parameters over time. These two-dimensional SMG features can be input into the AI model in the form of numerical values. For example, the two-dimensional SMG captured by the glove 300 can be used to identify the target muscle by segmenting the muscle aponeurosis and reconstructing the muscle bundles in the area. After muscle segmentation, the target muscle structure is quantified, i.e., feature extraction, to obtain the aforementioned two-dimensional SMG features. Based on the complete muscle structural information, the morphology and / or motion pattern of the target muscle can be calculated. For example, the following parameters can be used: tissue displacement, length, width, area, angle, grayscale gradient, and / or the magnitude and / or velocity of changes in these parameters over time. After segmenting the aponeurosis, the thickness between the aponeurosis is the muscle thickness. After obtaining the muscle bundles, the length, angle, and angle of the muscle bundle with the aponeurosis can be calculated. These parameters are related to the muscle's motion pattern. The two-dimensional SMG features include the morphology of the target muscle and reflect the muscle's motion pattern. The SMG data captured by the forearm wristband 100 is three-dimensional SMG. The morphology of the target muscle refers to muscle size, muscle bundle length, etc. The motion pattern is related to the parameters. The motion pattern is the movement pattern of the muscle, such as concentric contraction, eccentric contraction, and isometric contraction. For the forearm wristband 100, a longitudinal ultrasound image can be collected by multiple longitudinally positioned UTs 110 to reconstruct a three-dimensional SMG (three-dimensional ultrasound model). The three-dimensional ultrasound model can then be re-sliced in the transverse direction, and a three-dimensional model of the target muscle can be generated by segmenting the muscle in each transverse cross-sectional slice. Two-dimensional SMG features can also be calculated from the SMG data based on the spatial distribution of the muscle. The extracted two-dimensional SMG features can be concatenated to obtain an SMG feature map 540. For example, all features from the same transverse section can be combined into a matrix map, which serves as the SMG feature map 540.The three-dimensional model of the target muscle (or the three-dimensional ultrasound model obtained by the forearm wristband) and the SMG feature map 540 can both be used as inputs to the AI model. Among them, the SMG feature map obtained by the forearm wristband can also include parameters such as tissue displacement, length, width, area, angle, grayscale gradient, and / or the size and / or speed of the above parameters changing over time. That is, the two-dimensional SMG can be extracted on both the forearm wristband and the glove. The difference is that in the forearm wristband, it is obtained from the re-sliced ultrasound and input into the AI model in the form of a matrix feature map.
[0060] Figure 6 shows a flowchart for SMG-based gesture recognition according to an embodiment of the present disclosure. SMG-based gesture recognition generally includes the following steps: SMG data acquisition 610, SMG data processing 620, feature extraction and selection 630, and classification 640. Step 610 acquires SMG data, for example, including images acquired based on different modes. Based on the images acquired based on different modes, SMG data processing 620 can use various methods to reconstruct and segment the SMG data acquired in step 610 to obtain a segmented two-dimensional SMG or three-dimensional model, such as time gain compensation (TGC), Gaussian filtering, Hilbert transform, logarithmic compression, and registration (e.g., based on optical flow algorithms) to generate a deformation field. The deformation field is the movement, rotation, or deformation of a muscle detected between two adjacent images in a time series, representing the specific pattern expressed in the image during muscle contraction. For example, differences in the deformation field calculated from ultrasound images of the forearm can represent the movement of a stationary finger. Thus, the motion field can be used to detect which finger is moving, as well as the magnitude, force, and speed of the movement. Feature extraction and selection 630 can perform SMG feature extraction on the segmented two-dimensional SMG or three-dimensional model to obtain two-dimensional SMG features and feature maps. Feature extraction and selection 630 can use, for example, linear fitting, linear regression (LR), image edge detection, differential operators (such as Sobel, Prewitt, Laplacian, Roberts and Canny operators, etc.), high-pass filtering, discrete wavelet transform, etc. to process the deformation field generated in step 620 to perform feature extraction and selection. Classification 640 can use a variety of classifiers, such as support vector machine (SVM), back propagation (BP) neural network, linear discriminant analysis (LDA), K nearest neighbor (KNN) algorithm, multi-layer perceptron (MLP), stacked denoising autoencoder (SDA), naive Bayes classifier, decision tree-based classifier, etc. to process the features extracted and selected in step 630 to obtain classification results. Exemplarily, feature selection can be performed based on the classification results in step 630. Step 650 can be evaluated based on the classification results and classification indicators, such as accuracy, F1-SCORE, AUC and other indicators to obtain evaluation results. The performance of the classification system can be evaluated 650 based on various methods, for example, F score, etc. Step 660 can optimize and adjust one or more of the segmentation model, reconstruction algorithm, feature set, classifier hyperparameters, grid search parameter adjustment, etc. in the above steps based on the evaluation results. When using B-mode ultrasound or M-mode ultrasound, in order to reduce the amount of calculation, it can be verified whether the pre-defined target gesture can still ensure the accuracy of full resolution when reducing the lateral (arrangement direction of the UT array) spatial resolution. This can be based on the following fact: when the tissue depth range of interest is within the near field (Fresnel) region, the ultrasound beam size can be approximately the same as the UT size.This also means that it is feasible to successfully recognize the target gesture by deploying only a small number of UTs. When designing gestures, it should be considered that the recognition accuracy of discrete gestures is significantly higher than that of continuous gestures and has less computational complexity, especially to make full use of simplified gestures based on binary expressions.
[0061] FIG7A shows an interface control system for sign language recognition using both hands according to an embodiment of the present disclosure. Ultrasonic imaging is performed using a handband 100 on both or one side of the forearm, and SMG data of the forearm muscle groups is obtained. Simple muscle patterns and corresponding gestures are recognized in real time using AI algorithms and image processing. Sign language can be used as a direct gesture-based typing or command method. International sign languages are too complex and diverse, so custom sign languages are used to input words, characters, or commands.
[0062] Figure 7B shows an interface control system using single-handed sign language recognition according to an embodiment of the present disclosure. A single forearm wristband can facilitate simple interface control for people with disabilities. They can use similar custom sign language to type or send commands. The difference is that the sign language dictionary only stores the sign language of a single hand.
[0063] Custom sign language implements interface control, such as typing, by retrieving the value corresponding to the user's current sign language from the user-defined sign language dictionary. Figure 8 shows a schematic flow chart of sign language matching and insertion according to an embodiment of the present disclosure. When a user makes a custom sign language, the system attempts to match it in the sign language dictionary. If the corresponding custom sign language is found, the corresponding value is output. If the match is unsuccessful, the current sign language will be recognized as a new sign language. The user can use other inputs (for example, through voice) to customize the value of the new sign language and insert it into the sign language dictionary. The updated dictionary will be used to retrain the artificial intelligence model to ensure that the newly defined sign language can be retrieved in a timely manner.
[0064] In addition to sign language, the relaxation and bending of different fingers can also be used as binary expressions. Figure 9 shows a schematic diagram of finger states expressed through binary expressions according to an embodiment of the present disclosure. If each hand is only allowed to bend a single finger, there are 7 basic finger states, including: relaxation 910, thumb bending 920, index finger bending 930, middle finger bending 940, ring finger bending 950, little finger bending 960 and fist 970. These basic finger states can be combined between the two hands to convey more information. Compared with complex sign language, SMG-based gesture recognition is easier to achieve limited and fixed finger states. The way of combining fingers ensures the diversity of expression.
[0065] Figure 10 illustrates a schematic diagram of the steps involved in ten-key typing according to an embodiment of the present disclosure. Ten-key typing is a form of finger combination typing. For anyone who has used ten-key input on a phone's numeric keypad, this typing method is easy to learn. Each finger bend has a defined letter, symbol group, or function. For the left hand, bending the pinky, ring finger, middle finger, index finger, and thumb respectively represents ",.?!", "ABC," "DEF," "GHI," and "JKL." For the right hand, bending the thumb, index finger, middle finger, ring finger, and pinky respectively represents "MNO," "PQRS," "TUV," "WXYZ," and "Caps." The "Caps" command is used to lock caps, consistent with traditional keyboards. Bending a finger different times selects the corresponding character sequence within the symbol group. For example, if a user wants to type "POLYU," they should first bend their right pinky once to lock capital letters, bend their right index finger once to type "P," bend their right thumb three times to type "O," and then bend their left thumb three times to type "L." Similar to bending other fingers, output the remaining "Y", "U" and "!" in sequence.
[0066] Furthermore, the ten numbers "1234567890" correspond to the left pinky finger and the right pinky finger, respectively. Holding the fingers bent for more than one second will input the corresponding number. Bending the right pinky finger twice will backspace, and making a fist will input a space. It should be understood that the above input method is merely exemplary; other input methods can be defined for the ten keys by combining the bending state, bending times, and bending duration of the ten fingers.
[0067] Figure 11 shows a schematic diagram of a method for switching keyboard configurations according to an embodiment of the present disclosure. In order to achieve more interactive possibilities, the keyboard configuration can be switched by making a fist with one hand. When the right hand makes a fist, the left hand fingers are bent from left to right, representing keyboard configurations 1-5. When the left hand makes a fist, the right hand fingers are bent from left to right, representing keyboard configurations 6-10. Keyboard configuration 1 is the default English keyboard. Keyboard configurations can include numeric keyboards, symbol keyboards, other language keyboards, and even user-defined shortcut keys. It should be understood that the above-mentioned definitions of keyboard configurations are only exemplary, and other keyboard configurations can also be defined.
[0068] Figure 12 shows a schematic diagram of a ten-key typing guidance interface on a display or projection medium according to an embodiment of the present disclosure. Typing guidance interfaces are designed for different media, not only guiding users in inputting commands but also visualizing and providing feedback to users on their interactions, enhancing the sense of interaction. The media may include one or more of the following: display, projection, augmented reality (AR), and mixed reality (MR).
[0069] For traditional display and projection devices, ten virtual buttons will be displayed at the bottom of the interface, representing the fingers of both hands. When the muscle activity of finger flexion is recognized, the virtual button corresponding to the finger will be highlighted when pressed. The character group to be selected and the currently selected character will be displayed under the input cursor until another finger flexion is detected to complete the current character input. In the keyboard selection interface, the button on the side of the fist will be temporarily folded, and the keyboard options will be displayed on the other side.
[0070] Figure 13 shows a schematic diagram of a ten-key typing guide interface on an AR or MR media according to an embodiment of the present disclosure. AR and MR devices are highly interactive, and the displayed content can interact directly with reality. Therefore, the AR or MR device typing guide can interact with gestures and hand movements to provide more powerful visual feedback. When the entity of the hand does not appear in the field of view of AR or MR, the typing guide is similar to the interface of traditional display and projection, displayed at the bottom of the field of view. When the entity of the hand appears in the field of view of AR or MR, the typing guide will be directly adsorbed above each finger. When the finger is bent, the finger area and the character group above it will be highlighted to remind the user that the gesture is successfully recognized.
[0071] Figure 14 shows a schematic diagram of the steps of key table typing according to an embodiment of the present disclosure. Key table typing combines the bent fingers of both hands to represent characters. The fingers of the left hand are used to locate the rows in the key table. In addition to making a fist, the six basic finger positions: relaxed 910, bent thumb 920, bent index finger 930, bent middle finger 940, bent ring finger 950, and bent pinky 960 respectively locate the rows "AE", "FJ", "KO", "PT", "UY", and "Z-!" (see Figure 15). The fingers of the right hand are used to locate the columns in the key table. Therefore, combining the two hands can locate a specific character in the key table. Bending the right hand finger once represents lowercase, and bending it twice represents uppercase. For example, suppose the user wants to type "PolyU!". In this case, he / she should first bend his / her left middle finger once and right thumb twice to type "P", bend his / her left index finger once and right pinky finger once to type the lowercase "o", and then bend his / her left index finger once and right index finger once to type the lowercase "l". Similar to bending the other fingers, the remaining "y," "U," and "!" are output in sequence. It should be understood that other keyboards indexed by row and column can also be defined. For binary finger states, the left and right hand gestures are used to index the rows and columns of the custom key table, respectively, to find the corresponding key values, thereby enabling typing of all characters and function keys on the custom key table.
[0072] Additionally, making a fist with both hands simultaneously can input a space, while making a fist with one hand can enter keyboard configuration switching mode, similar to ten-key typing. Keyboard configuration 1 is the default English keyboard. Keyboard configurations can include a numeric keypad, a symbol keyboard, keyboards for other languages, and even user-defined shortcuts. It should be understood that the above keyboard configuration definitions are merely exemplary, and other keyboard configurations can also be defined.
[0073] Figure 15 shows a schematic diagram of a key table typing guide interface on a display or projection medium according to an embodiment of the present disclosure. For traditional display and projection devices, ten virtual buttons are displayed at the bottom of the interface, representing the fingers of two hands. The button on the left indicates the character corresponding to the row, while the button on the right indicates the key table above it. The rows and columns in the key table have indicator bars that can be slid to visually locate the characters. When the muscle activity of the finger flexion is recognized, the virtual button corresponding to the finger will be highlighted when it is pressed. At the same time, the indicator bar will also move to the corresponding position, and the character where the two indicator bars of the row and column intersect is the user's output. In the keyboard configuration switching interface, the button on the side of the fist will be temporarily folded, and the keyboard options will be displayed on the other side, which is the same as ten-key typing.
[0074] Figure 16 shows a schematic diagram of a key table typing guide interface on an AR or MR media according to an embodiment of the present disclosure. The typing guide in the AR or MR device can interact with gestures and hand movements to provide more powerful visual feedback. When the entity of the hand does not appear in the field of view of AR or MR, the typing guide is similar to the interface of traditional display and projection, displayed at the bottom of the field of view. When the entity of the hand appears in the field of view of AR or MR, the typing guide will be directly adsorbed above each finger. When the finger is bent, the finger area and the character group above it will be highlighted to remind the user that the action has been successfully recognized.
[0075] The aforementioned sign language and finger-based interface control system can monitor wrist muscle activity or simple finger movements using only a forearm strap or glove. To enable a more natural keyboard and mouse interface control system using SMG-based acquisition devices without requiring additional learning, a combination of gloves and forearm straps can be used to achieve gesture recognition down to the finger level.
[0076] Figures 17A and 17B illustrate schematic diagrams of an interface control system using a cursor and a virtual keyboard according to an embodiment of the present disclosure. In some embodiments, the SMG data acquisition device can be a wearable ultrasound acquisition device. The wearable ultrasound acquisition device can be configured in two ways, for example. One combination includes two gloves 300 and a forearm band 100 on only the dominant hand, as shown in Figure 17A . Another combination includes two gloves 300 and two forearm bands 100, as shown in Figures 20 and 20B . The gloves 300 and forearm bands 100 are equipped with UTs 110 and 310 to capture ultrasound images of target muscles. The acquired data is transmitted in real time to a terminal device or the cloud for subsequent processing via a data acquisition module 320 located inside the glove 300. The redundant data acquisition module 120 on the forearm band 100 is removed. The forearm band 100 and glove 300 combination can monitor fine movements of the wrist and fingers.
[0077] FIG18 shows a schematic diagram of mode switching according to an embodiment of the present disclosure. Mode switching gestures, such as a hand wave 1810, can be designed to switch between virtual keyboard mode 1820, touchpad mode 1830, and mouse mode 1840. As shown in FIG18, touchpad mode 1830 and mouse mode 1840 can be used to control the cursor, while virtual keyboard mode 1820 can be used for keyboard input. For cursor control, mouse mode 1840 is primarily implemented by monitoring wrist and finger movements using SMGs from the forearm muscles and lumbrical muscles. In the initial state, the middle finger, ring finger, and pinky finger of the dominant hand (i.e., the user's typically left or right hand) are bent, and the index finger and thumb are relaxed. When an action is detected, a classifier, such as an event classifier, determines the current action mode. An index finger up and down is a click event, a four-finger bend is a downward movement event, a four-finger extension is an upward movement event, a wrist twist to the left is a leftward movement event, and a wrist twist to the right is a rightward movement event. In touchpad mode 1830, the initial state is the same as in mouse mode. The index finger's flexion and extension, as well as left and right tilting movements, are monitored in real time to control cursor movement. For virtual keyboard mode 1820, the abduction and adduction of the two thumbs, as well as the fine movements of the other eight fingers, are monitored to reflect the finger positions on the virtual keyboard for each media in real time. Simultaneously, an event classifier identifies each finger press event.
[0078] Figure 19 shows a schematic diagram of controlling a cursor 1920 on a display and projection media according to an embodiment of the present disclosure. For display and projection media, the cursor 1920 is controlled to move on the interface through touchpad mode 1830 or mouse mode 1840. A virtual keyboard will be displayed at the bottom of the interface. Ten finger indicators 1910 will appear on the keyboard to indicate the current position of the finger in virtual keyboard mode 1820. To provide visual feedback to the user, the finger indicator 1910 performs real-time movement based on muscle activity. When the finger indicator 1910 hovers over a virtual button, the button will enlarge. When a click event on the finger is detected, the button below the finger indicator 1910 will be highlighted to indicate that the user has successfully entered.
[0079] Figure 20 shows a schematic diagram of an SMG-based interface control system 2000 according to an embodiment of the present disclosure. The interface control system 2000 includes an SMG data acquisition device 2010. The SMG data acquisition device 2010 comprises multiple ultrasonic transducers (UTs) deployed on fingers and / or forearms and at least one data acquisition module for collecting SMG data. When deployed on a finger, the SMG data acquisition device 2010 is a wearable glove; when deployed on a forearm, it is a wearable forearm wristband. When deployed on both a finger and forearm, the SMG data acquisition device 2010 is a combination of a glove and a forearm wristband. As previously mentioned, the SMG data acquisition device 2010 can include a left-hand glove, a right-hand glove, a left forearm wristband, a right forearm wristband, or any combination thereof. The multiple UTs can utilize multiple discrete UTs or UT array probes. The SMG data acquisition module 2010 can be deployed on one or more devices. The interface control system 2000 also includes an SMG data processing device 2020, a feature extraction and selection device 2030, and a classification device 2040. Devices 2020-2040 primarily perform SMG-based gesture recognition. The SMG data processing device 2020 processes the SMG data collected by the SMG data acquisition device 2010 to obtain image data that is useful for identifying predefined target muscles. For example, the image data that is useful for identifying predefined target muscles may be ultrasonic image data of the muscles, as well as muscle motion image data, muscle hardness image data, and ultrasonic image data of different muscles after segmentation obtained by processing these ultrasonic image data. This image data may be one-dimensional, two-dimensional, three-dimensional, or multi-dimensional. The feature extraction and selection device 2030 processes the image data of the target muscles to extract features of the morphology and / or motion pattern of the target muscles and select features that are useful for identifying predefined target gestures. The classification device 2040 classifies the selected features based on the target gesture to identify the correct gesture. One or more of the devices 2020-2040 may be deployed on one or more AI models. The explanatory portion of FIG. 6 lists the possible implementation methods of the algorithms, models, etc., which will not be repeated here.
[0080] The interface control system 2000 also includes an interface control device 2050. The interface control device 2050 can interpret the recognized gestures as control commands of the interface based on a control algorithm deployed locally or online, thereby realizing the HMI. Preferably, the interface control system 2000 also includes an interface display device 2060. The interface display device 2060 can be a display as a medium, a projection interface, an interface with a cursor, a finger pointer and / or a virtual keyboard presented in AR or MR, which can provide intuitive visual feedback. Optionally, the interface control system 2000 also includes a tactile feedback device. Feedback on the user's actions can be provided tactilely. Tactile feedback can be a group of devices arranged on the data acquisition device 2010 that can generate microcurrents or tiny mechanical vibrations. For example, when a finger "taps" a specific keyboard correctly or incorrectly, a tiny vibration feedback is sent to the finger.
[0081] In the present disclosure, to implement multimodal keyboard and mouse command output using the SMG, the classification device 2040 preferably uses an MLP-based event classifier to identify events such as "move," "click," "double-click," and "mode switch." When the current motion mode is classified as a "move" event, the SMG data is input into AI models of different modes for motion command output. In mouse mode 1840, the model outputs the direction of cursor movement based on the four hand motion patterns and the speed of cursor movement based on SMG features. In touchpad mode 1830, the model outputs the position of the cursor relative to the initial cursor position. In virtual keyboard mode 1820, the model outputs the position of the finger pointer 1910 relative to the initial position. Figure 21 shows a schematic diagram of an exemplary SMG-based interface control system according to an embodiment of the present disclosure. As shown in Figure 21, it is assumed that a user (indicated by "person" in the figure) wears a glove on their left hand and a forearm wristband on their right hand. The user's SMG data is collected via the glove and wristband and input into the SMG data processing device. The SMG data processing device performs muscle segmentation processing on the SMG data to identify the image of the target muscle, and the feature extraction and selection device performs feature extraction and selection processing on the image of the target muscle to obtain SMG features. The classification device includes an event classifier, and the event classifier processes the SMG features to identify events such as "movement", "click", "double-click" and "mode switching". It is assumed here that the identified event is "movement", then the classification device also includes AI models of different modes. For example, when the mode includes a mouse mode and a virtual keyboard mode, the AI models of different modes include an ANN model for pointer movement and an ANN model for indicator movement. The ANN model for pointer movement and the ANN model for indicator movement are respectively used for processing based on SMG features to perform mouse pointer control and finger indicator control.
[0082] FIG22 shows a flow chart of an exemplary SMG-based interface control method 2200 according to an embodiment of the present disclosure. The steps can be referred to the description of FIG20 and will not be repeated here.
[0083] This disclosure can be used to control robots (prosthetic limbs, exoskeletons, surgical robots, etc.), electronic devices (smartphones, keyboards, mice, etc.), game characters, sign language interpreters, etc.; it can also provide muscle activity information that can be integrated with musculoskeletal simulation software to provide useful information for rehabilitation and physical therapy; it is a novel system for detecting muscle activity during bodybuilding training, improving training output efficiency. Furthermore, this disclosure can help elderly people monitor their muscle activity to prevent falls.
[0084] The present disclosure is shown and described in detail above in conjunction with the drawings, but the above embodiments should be considered as illustrative rather than restrictive. The present disclosure also includes various combinations, modifications and variations of the exemplary embodiments without departing from the spirit and scope of the present disclosure.
Claims
1. An interface control system based on sonomyography (SMG), comprising: an SMG data acquisition device, the SMG data acquisition device comprising a plurality of ultrasonic transducers UT arranged on the fingers and / or forearms and at least one data acquisition module for acquiring SMG data; An SMG data processing device, the SMG data processing device is used to process the SMG data collected by the SMG data collection device to obtain image data for identifying a predetermined target muscle; a feature extraction and selection device for processing the image data to extract features of the morphology and / or movement pattern of the target muscle and select features for identifying a predefined target gesture; a classification device for classifying the selected features based on the target gesture to identify a correct gesture; An interface control device is used to interpret the gestures recognized by the classification device as control commands of an interface on a medium based on a control algorithm deployed locally or online, so as to realize human-computer interaction HMI.
2. The system according to claim 1, wherein: The SMG data acquisition device is a wearable device, wherein the SMG data acquisition device is a wearable glove when arranged on a finger, the SMG data acquisition device is a wearable forearm wristband when arranged on a forearm, and the SMG data acquisition device is a combination of a glove and a forearm wristband when arranged on both the finger and the forearm, wherein the SMG data acquisition device is a left-hand glove, a right-hand glove, a left forearm wristband, a right forearm wristband, or any combination thereof.
3. The system according to claim 1, wherein: The SMG data acquisition device adopts one or more of the following ultrasound modes: A-mode ultrasound, B-mode ultrasound and M-mode ultrasound, and the ultrasound dimensions of the collected or reconstructed SMG data include one or more of the following: one-dimensional, two-dimensional or three-dimensional ultrasound.
4. The system according to claim 1, wherein: The SMG data processing device identifies the target muscle by segmenting the muscle aponeurosis in the SMG data and reconstructing the muscle bundle in the area where the muscle aponeurosis is located.
5. The system according to claim 1, wherein The SMG data processing device adopts one or more of the following methods: time gain compensation TGC, Gaussian filtering, Hilbert transform, logarithmic compression, and registration to generate a deformation field, and the image data includes the deformation field.
6. The system according to claim 2, wherein: When the SMG data acquisition device is a glove, the morphology and / or movement pattern of the target muscle is obtained from at least one or more of the following parameters: tissue displacement, length, width, area, angle and grayscale gradient, as well as the magnitude and speed of changes of the above parameters over time.
7. The system according to claim 2, wherein: When the SMG data acquisition device is a forearm wristband, the SMG data includes three-dimensional SMG reconstructed from multiple UTs in the forearm wristband; For the three-dimensional SMGs reconstructed from multiple UTs in the forearm wristband, a three-dimensional model of the target muscle is generated by segmenting the muscle in each transverse cross-sectional slice, and the extracted two-dimensional SMG features are spliced together according to the spatial distribution of the muscle to obtain an SMG feature map, wherein the three-dimensional model of the target muscle and the SMG feature map are both used as inputs of an AI model, the classification device includes the AI model, and the features include the three-dimensional model of the target muscle and the SMG feature map.
8. The system according to claim 1, wherein: The feature extraction and selection device adopts one or more of the following methods: linear fitting, linear regression, image edge detection, differential operator, high-pass filtering and discrete wavelet transform, wherein the classification device adopts one or more of the following methods: support vector machine SVM, back propagation BP neural network, linear discriminant analysis LDA, K nearest neighbor KNN algorithm, multi-layer perceptron MLP, stacked denoising autoencoder SDA, naive Bayes classifier and decision tree-based classifier.
9. The system according to claim 1, wherein: The control command is a custom sign language based on the target gesture, wherein the interface control device searches for a target sign language that matches the recognized gesture in a predefined sign language dictionary, and outputs a value corresponding to the target sign language when the match is successful. When the match is unsuccessful, the interface control device supports updating the sign language dictionary.
10. The system according to claim 1, wherein: The target gesture is a binary expression based on relaxation and bending of different fingers, wherein the finger states of the binary expression include: relaxation, bending of thumb, bending of index finger, bending of middle finger, bending of ring finger, bending of little finger and fist.
11. The system according to claim 10, wherein: For the finger states expressed in binary, typing of all characters and function keys is achieved on ten keys by combining the bending states, bending times and bending times of the ten fingers.
12. The system according to claim 10, wherein: For the finger state of the binary expression, the rows and columns on the custom key table are indexed respectively through the left hand gesture and the right hand gesture to find the corresponding key at the intersection of the row and column, thereby realizing the typing of full characters and function keys on the custom key table.
13. The system according to claim 10, wherein: For the finger state of the binary expression, one or more keyboard configurations are switched by making a fist with one hand and bending the fingers of the other hand. The keyboard configurations include one or more of the following: English keyboard, numeric keyboard, symbol keyboard, other language keyboards and user-defined shortcut keys.
14. The system according to claim 1, wherein: The interface control system also includes an interface display device, which presents an interface with a cursor, a finger pointer and / or a virtual keyboard to provide typing guidance, mouse clicks, cursor control and / or keyboard switching, and provides corresponding visual feedback when the user's gesture recognition is successful. The interface display device is one or more of the following: a display, a projection interface, augmented reality AR and mixed reality MR.
15. The system according to claim 14, wherein: The SMG data acquisition device includes a combination of gloves for both hands and a forearm wristband. It realizes mouse clicks, cursor movement and full keyboard input by collecting wrist and finger movements, and provides real-time visual feedback on cursor movement and finger movement.
16. The system of claim 14, wherein: The interface display device switches between the following three modes through gestures: virtual keyboard mode, touchpad mode and mouse mode.
17. A method for controlling an interface based on a sonomyogram (SMG), comprising: Collecting SMG data by means of a plurality of ultrasonic transducers UT and at least one data acquisition module arranged on the fingers and / or forearms; processing the SMG data to obtain image data for identifying a predefined target muscle; Processing the image data to extract features of the morphology and / or movement pattern of the target muscle and selecting features for recognizing a predefined target gesture; classifying the selected features based on the target gesture to identify a correct gesture; and The control algorithm based on local or online deployment interprets the recognized gestures as control commands of the interface on the media to realize human-computer interaction HMI.