Non-contact medical picture operation terminal gesture decoding algorithm based on electromyographic signal recognition
By employing a multi-channel flexible electromyography (EMG) sensor array, anti-interference preprocessing, multi-dimensional feature fusion, and CNN-LSTM hybrid network decoding, combined with online incremental learning and scene-adaptive optimization, the shortcomings of existing EMG signal decoding technologies in complex clinical environments have been addressed. This has resulted in high-precision, robust, and continuous gesture decoding, enhancing the practicality and safety of the contactless medical drawing operation terminal.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 绍兴职业技术学院
- Filing Date
- 2025-12-25
- Publication Date
- 2026-05-05
AI Technical Summary
Existing electromyography (EMG) signal decoding technologies suffer from problems such as strong model dependence, insufficient generalization ability, over-reliance on static gestures for action recognition, single feature dimension, limited anti-interference ability, and limited clinical application in complex clinical environments, making it difficult to meet the requirements of high precision, adaptability, and smooth operation.
By employing a multi-channel flexible electromyography (EMG) sensor array with dynamic adjustment of sampling frequency and gain, combined with anti-interference preprocessing, multi-dimensional feature fusion, and CNN-LSTM hybrid network decoding, and introducing online incremental learning and scene adaptive optimization, high-precision, robust, and continuous gesture decoding of EMG signals can be achieved.
It improves the practicality and stability of the contactless medical image operation terminal, ensuring high-precision recognition and smooth operation in complex environments, adapting to the needs of different users and scenarios, reducing the risk of misoperation, and improving the efficiency and safety of clinical operations.
Smart Images

Figure CN121979384A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioelectric signal processing and intelligent control, and in particular to a gesture decoding algorithm for a contactless medical drawing operation terminal based on electromyography signal recognition. Background Technology
[0002] In recent years, with the development of intelligent medical devices and contactless operating terminals, gesture recognition and motion decoding technologies based on surface electromyography (sEMG) signals have been gradually applied to clinical surgical control, robot-assisted surgery, and neurorehabilitation training. SEMG signals are non-invasive, have a fast response time, and are easy to acquire, allowing for the issuance of operating commands to medical terminals without direct contact with the device. This helps reduce the risk of cross-contamination during surgery and improves the efficiency of continuous operation. However, existing SEMG signal decoding technologies still have several shortcomings in complex clinical environments and cannot fully meet the needs of high-precision medical control.
[0003] (I) Current Status of Industry Technology Current contactless electromyography (EMG) control technology mainly consists of key components such as EMG sensor arrays, signal preprocessing algorithms, feature extraction models, and gesture decoding networks.
[0004] The current state of sensor and hardware integration technology is as follows: Existing technologies generally employ multi-channel surface electromyography (EMG) sensor arrays on flexible substrates, using PDMS material as the support structure and conductive silver paste to construct electrodes, enabling attached acquisition of data from key muscle groups such as the biceps brachii, flexor carpi radialis, pronator, and supinator. Typical products are configured with an 8-channel array, typically measuring 80mm × 40mm × 0.3mm, covering six major muscle groups involved in fine hand movements and forearm rotation.
[0005] The hardware acquisition unit typically integrates the ADS1298 bioelectrical signal acquisition chip, which provides a sampling frequency range of 1000-2000Hz and supports gain adjustment from ×1000 to ×5000 to accommodate electromyographic signals of varying intensities. Under FPGA control, the sampling rate can be switched according to the specific surgical procedures required. Simultaneously, a magnetic shielding structure (permeability μ>5000) is added, and 50Hz power frequency interference is suppressed to improve signal quality and system stability.
[0006] Current motion recognition and scene adaptation technologies, in sterile operating room settings, typically rely on improved motion segment detection algorithms. These algorithms use a combination of energy thresholds, length constraints, and waveform features to determine valid motion segments within electromyographic (EMG) signals. For example, a resting energy threshold of 3 times the average resting energy, a rotational motion duration of ≥0.8s, and a scaling motion duration of ≥0.5s are set, along with a zero crossover rate (ZCR) > 0.8 to indicate motion occurrence. Rotation control largely depends on the coordinated activation of the pronator and supinator muscles; a rotation command is triggered when the MPF (mean power frequency) difference exceeds a certain threshold (e.g., 30Hz). Scaling motions are triggered by the rate of change of the integrated EMG value between the biceps and triceps (>50% / s). Some systems support multimodal motion combinations, such as "clenching a fist + rotating," enabling complex control actions like 360° field of view operation, reducing instrument switching frequency, and improving operational efficiency.
[0007] In postoperative rehabilitation, existing systems typically use the antagonist ratio (CR) to monitor muscle coordination. A CR > 1.8 indicates abnormal muscle compensation, helping patients adjust training intensity. Simultaneously, some systems incorporate near-infrared spectroscopy (NIRS) to monitor muscle oxygen saturation (StO2). Training is automatically terminated when StO2 falls below 60% to prevent excessive muscle fatigue. Furthermore, based on the patient's muscle strength score (e.g., MRC classification), the control system adjusts the sensitivity ratio to adapt to different rehabilitation stages. For example, at MRC level 3, the sensitivity is set to 1.5 times the normal level to compensate for low-amplitude signals from weak muscles.
[0008] The shortcomings of existing technologies are as follows: Although existing technologies have made some progress in electromyography (EMG) acquisition and gesture recognition, there are still several significant technical limitations in practical applications, mainly including the following aspects: 1. Strong model dependence and insufficient generalization ability: Existing gesture decoding models typically employ deep neural network structures, which are highly sensitive to the number of training samples, category distribution, and individual differences. When the sample size is insufficient or the dataset lacks diversity, the model is prone to overfitting, resulting in insufficient generalization ability. Differences in muscle strength, degree of nerve damage, degree of muscle fatigue, and training habits among different patients can significantly affect electromyographic signal characteristics, making it difficult for uniformly trained models to transfer stably among different users, resulting in large fluctuations in recognition accuracy. 2. Action recognition relies too heavily on static gestures and lacks adaptability to continuous actions: Existing methods mainly rely on predefined gesture libraries for recognition, and the recognition process is discrete, which cannot effectively support continuous and smooth operation control. In clinical surgery, such as continuous rotation of instruments and continuous adjustment of the camera, the system needs to be able to decode continuous electromyography changes in real time; however, traditional discrimination methods based on static gestures cannot meet this requirement, resulting in large operation delays, discontinuous actions, and affecting surgical efficiency. 3. Limited feature dimensions and insufficient robustness: Most existing systems rely primarily on time-domain features such as RMS for identification, without effectively integrating frequency-domain features (such as MPF and MF), time-frequency features (such as wavelet packet energy), and muscle group coordination features (such as phase difference IPD). When the amplitude of the electromyographic signal decreases (e.g., the signal attenuation is 60%-75% in patients with brachial plexus injury) or noise contamination occurs, the discriminative ability of most time-domain features decreases significantly, making it difficult to maintain the recognition accuracy above the clinical requirement of 90%, and unable to adapt to special populations or weak signal scenarios. 4. Limited anti-interference capability and lack of real-time scene adaptation: Although existing systems employ shielding structures to suppress electromagnetic interference, it is still difficult to ensure a sufficient signal-to-noise ratio in complex clinical environments. For example, in cardiovascular interventional surgery, electromagnetic interference often exceeds 10μT, causing distortion of electromyographic signal waveforms and leading to an increased failure rate in motion segment detection. Furthermore, most existing decoding models are static models, lacking real-time adaptive capability. They cannot dynamically adjust gain, filtering parameters, and model weights based on the current signal quality, resulting in a rapid decline in recognition performance in complex environments. 5. Limited clinical application and large-scale promotion: The complex hardware design and high cost of existing systems limit their popularization in primary healthcare institutions or large-scale rehabilitation centers. Some rehabilitation systems lack the ability to track long-term changes in patients' muscle strength in real time, resulting in the inability to dynamically update control sensitivity. This can easily lead to risks such as mismatch between operation difficulty and training effect, or even secondary muscle injury during training.
[0009] In summary, existing electromyography (EMG) signal decoding technologies still have shortcomings in terms of data dependence, model generalization ability, feature robustness, scenario adaptability, and clinical applicability. It is necessary to propose a novel real-time EMG signal decoding system that is stable under multiple scenarios and population conditions in order to achieve higher accuracy, stronger adaptability, and a better clinical operation experience. Summary of the Invention
[0010] To address the aforementioned technical problems, this application provides a gesture decoding algorithm for a contactless medical drawing operation terminal based on electromyography signal recognition.
[0011] The gesture decoding algorithm for a contactless medical drawing terminal based on electromyography signal recognition provided in this application adopts the following technical solution: A gesture decoding algorithm for a contactless medical imaging terminal based on electromyography (EMG) signal recognition includes the following steps: Step 1: Multi-channel electromyography (EMG) signal acquisition: EMG signals are acquired by an 8-channel flexible EMG sensor array attached to the surface of muscle groups such as the biceps brachii, flexor carpi radialis, pronator, and supinator. The EMG signals are digitized by the ADS1298 bioelectric chip, and the FPGA control module dynamically adjusts the sampling frequency and gain coefficient of ×1000 to ×5000 within the range of 500 to 2000 Hz according to the needs of the scenario. Step 2: Anti-interference preprocessing: Bandpass filtering (20-500Hz), 50Hz notch filtering, adaptive sliding window smoothing, and Z-score normalization are performed on the digitized electromyography signal. The start and end points of the action segment are detected based on energy threshold, length constraint, waveform features, and slope threshold to obtain the effective action signal segment. Step 3: Multi-dimensional feature fusion and extraction: Extract time-domain features, frequency-domain features, time-frequency features and muscle group coordination features from the effective action signal segment, and obtain 8 core features including IEMG, RMS, MPF, MF, WPE3, CR, IPD1 and IPD2 from 32 initial features through a joint screening method of genetic algorithm and LASSO regression. Step 4: CNN-LSTM hybrid network decoding: The 8-channel core features are constructed into a time-spectrum map with a size of 128×128×8. Spatial features are extracted through a CNN network including three convolutional layers and two pooling layers, and time-series features are extracted through a two-layer bidirectional LSTM network. Finally, the 6-DOF control commands corresponding to the medical imaging terminal are output. Step 5: Scene Adaptive Optimization: The model parameters are updated based on new feature data during user use using an online incremental learning method. The dynamic response speed, signal-to-noise ratio threshold, and control sensitivity are adjusted in the sterile operating room mode and postoperative rehabilitation mode, respectively. When the muscle oxygen saturation is lower than the preset threshold or the signal quality is substandard, control pause and prompt are triggered.
[0012] By adopting the above technical solutions, high-precision, robust, and continuous gesture decoding control of electromyographic signals in medical interaction scenarios is achieved, improving the practicality and stability of the contactless medical drawing operation terminal. The multi-channel flexible electromyographic array, combined with dynamically adjustable sampling rate and gain control, enables the system to maintain high-fidelity signal acquisition for different muscle groups, different movement intensities, and different clinical scenarios, laying the foundation for subsequent decoding. The anti-interference preprocessing module, through multi-level processing such as bandpass filtering, notch filtering, and adaptive smoothing, combined with comprehensive constraints of energy, duration, waveform, and slope, can accurately identify effective action segments even in strong electromagnetic interference environments, improving the stability of action segment extraction. Through multi-dimensional fusion feature extraction and joint feature screening methods, the system extracts the most discriminative core features from the time domain, frequency domain, time-frequency, and muscle group coordination relationships, ensuring that the features maintain high discriminative power even under weak signal or abnormal muscle strength conditions. The CNN-LSTM hybrid network utilizes convolutional structures to extract spatial correlations and bidirectional LSTM to capture temporal dependencies, achieving accurate encoding of complex continuous action patterns and outputting 6... The freedom of control commands significantly improves the smoothness of operation. The introduction of online incremental learning and scenario-adaptive strategies enables the system to dynamically adjust response speed, control sensitivity, and signal quality thresholds in different application modes such as the operating room and rehabilitation training. It can also automatically pause and prompt when insufficient muscle oxygen saturation or decreased signal quality is detected, ensuring operational safety and training effectiveness.
[0013] Optionally, in the detection of the start and end points of the action segment, the slope threshold is set to a signal rise or fall slope greater than 0.5 mV / s.
[0014] By adopting the above technical solution, the slope threshold is set to a signal rise or fall slope greater than 0.5 mV / s, which enables the detection of the start and end points of the action segment to more accurately distinguish the real action changes and background fluctuations in the electromyographic signal. This threshold, combined with energy features, waveform features and length constraints, can effectively improve the robustness of action segment recognition and reduce false triggers or missed detections. The system can not only quickly lock the effective start and end positions of the action, but also ensure the reliability and stability of action decoding under complex interference conditions, thereby enhancing the real-time performance and accuracy of the entire gesture decoding algorithm.
[0015] Optionally, the sliding window smoothing process uses a moving average method with a window size of 0.2s.
[0016] By adopting the above technical solution, a sliding window smoothing process is performed using a moving average method with a window size of 0.2s. This effectively suppresses random noise and transient spike interference in electromyographic signals without significantly delaying the system response. The window length can achieve a balance between smoothness and agility, making the overall waveform of the signal more stable and feature extraction more reliable. This improves the accuracy and robustness of subsequent action segment detection and feature recognition, providing a higher quality input foundation for gesture decoding.
[0017] Optionally, the genetic algorithm has a population size of 50, an iteration count of 30, a crossover probability of 0.6, and a mutation probability of 0.05.
[0018] By adopting the above technical solution, the genetic algorithm uses a population size of 50, 30 iterations, a crossover probability of 0.6 and a mutation probability of 0.05 to select core features. This achieves a good balance between sufficient search range and computational efficiency. The parameter settings ensure that the algorithm has sufficient population diversity and evolution speed, effectively escapes local optima and finds more discriminative feature combinations. The core features selected are more stable and more representative, providing high-quality input for subsequent neural network decoding and improving the overall accuracy and generalization ability of gesture recognition.
[0019] Optionally, the time and frequency dimensions of the constructed time spectrum diagram are both set to a resolution of 128 points.
[0020] By adopting the above technical solution, the time and frequency dimensions of the time spectrogram are both set to a resolution of 128 points. While ensuring the details of the timing and frequency information, the computational efficiency is also taken into account. This resolution can fully express the dynamic characteristics of electromyographic signals in the time and frequency domains, enabling the convolutional neural network to effectively capture spatial correlations and the bidirectional LSTM to extract temporal dependencies. This improves the accuracy and stability of continuous motion decoding, helps to maintain recognition robustness under different users and different muscle strength conditions, and provides reliable gesture control input for medical imaging terminals.
[0021] Optionally, the CNN has a 3×3 convolution kernel size, a stride of 1, and a padding of 1, and the pooling layer uses a 2×2 pooling kernel.
[0022] By adopting the above technical solution, the CNN network has a convolution kernel size of 3×3, a stride of 1, and padding of 1. The pooling layer uses a 2×2 pooling kernel, which allows the temporal spectrogram to retain fine-grained information and effectively reduce feature dimensionality during spatial feature extraction. Convolution operation enhances local pattern recognition capability, while pooling operation enhances translation invariance and reduces computational load, thereby improving the extraction accuracy of key features in electromyography signals. This significantly enhances the network's ability to discriminate continuous hand gestures and improves the accuracy and real-time response performance of gesture decoding on medical imaging terminals.
[0023] Optionally, the LSTM network has a bidirectional structure and each layer contains 128 hidden units.
[0024] By adopting the above technical solution, the LSTM network adopts a bidirectional structure, with each layer containing 128 hidden units. It simultaneously captures the forward and backward dependencies of electromyographic signals in the time series, enhancing the network's ability to model the dynamic changes of continuous movements. This allows subtle temporal features in the action sequence to be preserved and utilized, improving the accuracy and stability of continuous gesture decoding. The combination of spatial features extracted by bidirectional LSTM and CNN provides high-precision, low-latency real-time control commands for the medical imaging operation terminal, improving the smoothness and safety of clinical operations.
[0025] Optionally, the online incremental learning performs a model parameter update once after every 10 valid operations, and the update uses 1 / 10 of the pre-training learning rate.
[0026] By adopting the above technical solution, online incremental learning updates the model parameters after every 10 valid operations, and sets the update learning rate to 1 / 10 of the pre-training stage. While ensuring model stability, it achieves personalized adaptation, enabling the model to capture changes in user operating habits and the dynamic characteristics of electromyographic signals in a timely manner. It can optimize the decoding effect without retraining, improve the accuracy and robustness of continuous gesture recognition, and provide a smoother, safer and more personalized control experience for medical drawing operation terminals.
[0027] Optionally, in sterile operating room mode, the gain factor is automatically increased to enhance signal quality when the signal-to-noise ratio is below 30dB.
[0028] By adopting the above technical solution, in the sterile operating room mode, when the signal-to-noise ratio of electromyography signals is lower than 30dB in real time, the system automatically increases the gain coefficient to enhance signal quality, effectively resists the influence of electromagnetic interference and environmental noise on signal acquisition, and ensures that key gestures can still be accurately recognized in a high-interference environment. The medical imaging terminal maintains high precision and stability during the operation, improves the operation response speed and safety, reduces the risk of misoperation, and ensures the reliability and smoothness of clinical sterile operation.
[0029] Optionally, in the postoperative rehabilitation mode, the sensitivity can be adjusted according to the MRC muscle strength grade, where MRC grade 1-2 is set to ×2.0, MRC grade 3 is set to ×1.5, and MRC grade 4 is set to ×1.0. When the signal-to-noise ratio is below 25dB for 5 consecutive seconds or the failure rate of motion segment recognition exceeds 30%, the system triggers a hardware self-test and prompts the user to reattach the sensor or adjust its position.
[0030] By adopting the above technical solution, the sensitivity of the control is dynamically adjusted according to the MRC muscle strength level in the postoperative rehabilitation mode (MRC 1-2 × 2.0, MRC 3 × 1.5, MRC 4 × 1.0), which can adapt to the operational ability of patients with different muscle strength, improve the accuracy and safety of training. When the signal-to-noise ratio is lower than 25dB for 5 consecutive seconds or the failure rate of action segment recognition exceeds 30%, the system automatically triggers hardware self-test and prompts the user to reattach the sensor or adjust its position, effectively ensuring the signal acquisition quality, avoiding the risk of misoperation or secondary injury, and improving the effect and reliability of rehabilitation training.
[0031] In summary, this application includes at least one of the following beneficial technical effects: By combining a multi-channel flexible electromyography (EMG) sensor array with dynamically adjustable sampling frequency and gain coefficient, high-fidelity acquisition of EMG signals from different muscle groups, different movement intensities, and different clinical scenarios is achieved, providing reliable input for subsequent decoding and improving the accuracy and precision of gesture recognition.
[0032] A comprehensive action segment detection method is adopted, which combines bandpass filtering, notch filtering, adaptive sliding window smoothing, slope threshold, energy threshold, waveform features and length constraints. This method effectively suppresses random noise and transient interference, improves the stability of action segment recognition, and reduces false triggering and missed detection.
[0033] By fusing time-domain, frequency-domain, time-frequency, and muscle group synergy features, and combining genetic algorithms with LASSO regression for joint screening, core features with strong discriminative power are obtained, improving the recognition ability and generalization performance under weak signal, abnormal muscle strength, or small sample conditions.
[0034] The CNN-LSTM hybrid network structure combines convolutional layers to extract spatial features and bidirectional LSTM to capture temporal dependencies, enabling accurate decoding of complex continuous action patterns. It can output 6-DOF control commands adapted to medical imaging terminals, ensuring smooth operation and real-time response.
[0035] The online incremental learning mechanism can dynamically update the model based on user operating habits and new feature data, achieving personalized adaptation without retraining; dynamic parameter adjustment (gain, sensitivity, response speed, signal quality threshold, etc.) in sterile operating room mode and postoperative rehabilitation mode ensures operational accuracy, safety and training effect in different scenarios.
[0036] When muscle oxygen saturation falls below the threshold or signal quality fails to meet standards, the system automatically pauses operation and prompts the user to make adjustments to avoid misoperation and secondary damage, ensuring the safety of surgical procedures and rehabilitation training.
[0037] The system can maintain the stability of signal acquisition and motion recognition even in high electromagnetic interference and complex clinical environments, improving the reliability and practicality of the contactless medical image operation terminal.
[0038] Multimodal gesture recognition and continuous control capabilities make medical imaging terminal operation more intuitive, faster and smoother, reduce surgical instrument switching time, improve rehabilitation training experience and increase clinical operation efficiency. Attached Figure Description
[0039] Figure 1 This is a system architecture diagram of an embodiment of this application.
[0040] Figure 2 This is a schematic diagram of the sensor attachment position according to an embodiment of this application.
[0041] Figure 3 This is a diagram of the CNN-LSTM hybrid network structure of an embodiment of this application.
[0042] Figure 4 This is a flowchart illustrating an embodiment of this application.
[0043] Figure 5 This is the algorithm code of an embodiment of this application. Detailed Implementation
[0044] The following is in conjunction with the appendix Figure 1-5 This application will be described in further detail.
[0045] This application discloses a gesture decoding algorithm for a contactless medical drawing terminal based on electromyography (EMG) signal recognition. It includes the following steps: Step 1: Multi-channel electromyography (EMG) signal acquisition: EMG signals are acquired by an 8-channel flexible EMG sensor array attached to the surface of muscle groups such as the biceps brachii, flexor carpi radialis, pronator, and supinator. The EMG signals are digitized by the ADS1298 bioelectric chip, and the FPGA control module dynamically adjusts the sampling frequency and gain coefficient of ×1000 to ×5000 within the range of 500 to 2000 Hz according to the needs of the scenario. Step 2: Anti-interference preprocessing: Bandpass filtering (20-500Hz), 50Hz notch filtering, adaptive sliding window smoothing, and Z-score normalization are performed on the digitized electromyography signal. The start and end points of the action segment are detected based on energy threshold, length constraint, waveform features, and slope threshold to obtain the effective action signal segment. Step 3: Multi-dimensional feature fusion and extraction: Extract time-domain features, frequency-domain features, time-frequency features and muscle group coordination features from the effective action signal segment, and obtain 8 core features including IEMG, RMS, MPF, MF, WPE3, CR, IPD1 and IPD2 from 32 initial features through a joint screening method of genetic algorithm and LASSO regression. Step 4: CNN-LSTM hybrid network decoding: The 8-channel core features are constructed into a time-spectrum map with a size of 128×128×8. Spatial features are extracted through a CNN network including three convolutional layers and two pooling layers, and time-series features are extracted through a two-layer bidirectional LSTM network. Finally, the 6-DOF control commands corresponding to the medical imaging terminal are output. Step 5: Scene Adaptive Optimization: The model parameters are updated based on new feature data during user use using an online incremental learning method. The dynamic response speed, signal-to-noise ratio threshold, and control sensitivity are adjusted in the sterile operating room mode and postoperative rehabilitation mode, respectively. When the muscle oxygen saturation is lower than the preset threshold or the signal quality is substandard, control pause and prompt are triggered.
[0046] Reference Figures 1-4 The overall structure of the medical drawing operation terminal that supports gesture decoding algorithms based on electromyography (EMG) signal recognition includes five modules: a multi-channel EMG signal acquisition module, an anti-interference preprocessing module, a multi-dimensional feature fusion extraction module, a CNN-LSTM hybrid network decoding module, and a scene adaptive optimization module. These modules work together to achieve high-precision and robust gesture decoding and control. Sensor configuration for multi-channel electromyography signal acquisition: The multi-channel flexible electromyography sensor array (size: 80mm×40mm×0.3mm) is used and attached to the surface of 6 key muscle groups, including the biceps brachii, flexor carpi radialis, pronator, and supinator, to ensure comprehensive signal acquisition. Data source: A dual-source acquisition strategy of "self-collected data + public database" was adopted. The self-collected data came from 200 healthy subjects and 50 postoperative rehabilitation patients (including brachial plexus injury, muscle weakness and other conditions). The public database used authoritative datasets such as NinaPro and EMGDB to ensure the diversity and representativeness of the training data. Hardware parameters: The ADS1298 bioelectric chip and FPGA control unit are retained, and the sampling frequency dynamic adjustment range is optimized to 500-2000Hz (adaptively switching according to scenario requirements, such as 1500-2000Hz for surgical scenarios and 500-1000Hz for rehabilitation training scenarios), and the gain coefficient is maintained at ×1000-×5000. Anti-interference preprocessing includes filtering: a combination of "bandpass filtering + notch filtering + adaptive smoothing" is adopted. The bandpass filtering frequency range is set to 20-500Hz (to filter low-frequency noise and high-frequency interference other than electromyography signals), and the notch filtering is used to specifically remove 50Hz power frequency interference and harmonics. Signal optimization: Signal jitter is reduced by using a sliding window smoothing algorithm (window size 0.2s), and Z-score normalization is used (the signal mean is normalized to 0 and the standard deviation is normalized to 1) to eliminate the influence of differences in signal amplitude between different individuals. Signal segmentation: Based on the improved multi-threshold continuous start and end point detection algorithm, a slope threshold (signal rise / fall slope > 0.5mV / s) is added on the basis of the original energy threshold, length constraint and waveform characteristics, which further improves the accuracy of action segment recognition and obtains a single-channel effective action signal segment after segmentation.
[0047] Multi-dimensional feature fusion extraction includes feature dimension design: fusing four major categories of features: time domain, frequency domain, time-frequency, and muscle group coordination, covering a total of 32 initial features. Temporal characteristics: integrated electromyography (IEMG) value, root mean square (RMS) value, peak value, waveform length (WL), etc. Frequency domain characteristics: average power frequency (MPF), median frequency (MF), peak power spectral density (PSD), etc. Time-frequency features: wavelet packet energy (WPE, decomposition level 3, covering 8 sub-features from WPE1 to WPE8). Synergistic features: antagonistic muscle ratio (CR), intermuscular phase difference (IPD, including IPD1-IPD4, a total of 4 intermuscular phase difference features); Feature selection optimization: The 32 initial features were selected using a joint optimization strategy of "Genetic Algorithm (GA) + LASSO regression". First, the genetic algorithm (population size 50, number of iterations 30, crossover probability 0.6, mutation probability 0.05) was used to initially select 15 effective features. Then, LASSO regression (regularization parameter λ=0.01) was used to further eliminate redundant features, and finally 8 core features (IEMG, RMS, MPF, MF, WPE3, CR, IPD1, IPD2) were retained to reduce the computational complexity of the model.
[0048] The CNN-LSTM hybrid network decoding includes input processing: converting the core features of the 8-channel electromyography signal into a time-spectrum graph (size: 128×128×8), with 128 points in the time dimension, 128 points in the frequency dimension, and 8 points in the channel dimension, taking into account both spatiotemporal feature representation; The network architecture includes a hybrid CNN+LSTM network. The CNN part uses 3 convolutional layers (3×3 kernel size, stride 1, padding=1) + 2 pooling layers (2×2 kernel size, stride 2) to extract spatial features from the temporal spectrogram. The LSTM part uses 2 bidirectional LSTM layers (128 hidden units) to capture the temporal series features of the signal. The output layer is a fully connected layer that outputs 6 degrees of freedom control commands (3-axis rotation + 3-axis translation) to meet the requirements of continuous operation. The model training process includes: first, pre-training on a dataset of 200 healthy subjects (50 iterations, batch size 32, learning rate 0.001, and Adam optimizer) to obtain the basic model; then, fine-tuning on a dedicated dataset for postoperative patients (such as brachial plexus injury, MRC grade 1-4) (20 iterations, batch size 16, and learning rate 0.0001) to improve the model's adaptability to specific populations.
[0049] Scene adaptive optimization includes online incremental learning (OIL): an online incremental learning mechanism is introduced, which automatically extracts the feature data of the new operation after every 10 effective operations and incrementally updates the model parameters (the learning rate decays to 1 / 10 of the pre-training stage), so that it can adapt to changes in user operation habits without retraining. Scene parameter adaptation: Different model parameter configurations are preset for two major scenarios: sterile operating room and postoperative rehabilitation. Sterile operating room mode: Optimize dynamic response speed, control motion recognition delay within 50ms, and enhance anti-electromagnetic interference capability (by real-time monitoring of signal-to-noise ratio, automatically increase gain coefficient when signal-to-noise ratio < 30dB). Postoperative rehabilitation mode: The sensitivity of the control is dynamically adjusted according to the MRC muscle strength grade (MRC 1-2: sensitivity × 2.0; MRC 3: sensitivity × 1.5; MRC 4: sensitivity × 1.0), and the muscle oxygen saturation (StO2) monitoring function is retained. When StO2 < 60%, training is automatically paused to ensure rehabilitation safety. Anomaly handling mechanism: When the signal quality fails to meet the standard for 5 consecutive seconds (e.g., signal-to-noise ratio <25dB, action segment recognition failure rate >30%), hardware self-test and user prompts are triggered, suggesting that the sensor be re-attached or the acquisition position adjusted.
[0050] The core components and connections are achieved through a multi-channel flexible electromyography sensor array, an ADS1298 bioelectrical acquisition chip, an FPGA control module, a CNN-LSTM algorithm processing module, and a wireless communication module (used to transmit control commands to the medical operating terminal). Connection: The sensor array is attached to the surface of the muscle group through conductive silver paste. The collected electromyographic signals are converted into digital signals by the ADS1298 chip and then transmitted to the FPGA control module for sampling frequency and gain adjustment. The preprocessed signal is transmitted to the CNN-LSTM algorithm processing module through the data bus to complete feature extraction and decoding. The decoded control commands are sent to the medical operation terminal through the wireless communication module (Bluetooth 5.0 / BLE) to realize contactless control.
[0051] Referring to Figures 1-4, a gesture decoding algorithm for a contactless medical drawing operation terminal based on electromyography (EMG) signal recognition achieves continuous, precise, and robust gesture control of the medical operation terminal through multi-channel EMG signal acquisition, anti-interference preprocessing, multi-dimensional feature fusion, CNN-LSTM hybrid network decoding, and scene adaptive optimization. I. Multi-channel electromyography signal acquisition module Referring to Figures 1 and 2, this embodiment employs a multi-channel flexible electromyography (EMG) sensor array (size: 80mm × 40mm × 0.3mm) attached to the surface of six key muscle groups, including the biceps brachii, flexor carpi radialis, pronator, and supinator, ensuring comprehensive signal acquisition from these key muscle groups. The sensor array adheres to the skin via conductive silver paste to ensure stable signal transmission. The acquired EMG signals are digitized using an ADS1298 bioelectric chip and transmitted to an FPGA control module. The FPGA control module can dynamically adjust the sampling frequency (500–2000Hz) and gain coefficient (×1000–×5000) according to different application scenarios. In surgical scenarios, the sampling frequency can be increased to 1500–2000Hz to ensure high-precision real-time operation; in rehabilitation training scenarios, the sampling frequency can be reduced to 500–1000Hz to reduce power consumption and meet training requirements. This multi-channel acquisition strategy can simultaneously acquire signals from different muscle groups, enhancing the system's ability to perceive complex movements and providing high-fidelity data for subsequent gesture decoding. The data sources adopted a dual-source strategy of "self-collected data + public databases". The self-collected data included 200 healthy subjects and 50 postoperative rehabilitation patients, covering brachial plexus injury, muscle weakness and patients of different ages. The public databases included NinaPro, EMGDB and other databases to ensure the diversity and representativeness of the training data and support the generalization ability of the model.
[0052] II. Anti-interference preprocessing module The acquired digital electromyography (EMG) signals are first subjected to filtering and signal optimization processing, specifically including: Filtering: Bandpass filtering (20-500Hz) is used to remove low-frequency baseline drift and high-frequency noise; a 50Hz notch filter is used to suppress power frequency interference and its harmonics to ensure signal quality. Sliding window smoothing: A moving average algorithm with a window size of 0.2s is used for smoothing to suppress random noise and transient spike interference while ensuring the real-time performance of the signal response; Normalization: Z-score normalization is performed on the signals of each channel, adjusting the mean to 0 and the standard deviation to 1 to eliminate the differences in electromyography amplitude among different individuals and improve the consistency of feature extraction. Action segment start and end point detection: Based on an improved multi-threshold continuous start and end point detection algorithm, this method combines energy threshold, length constraint, waveform features, and slope threshold (signal rise or fall slope > 0.5 mV / s) to accurately identify the start and end positions of action segments and obtain valid action signal segments. This method ensures robustness while reducing false triggers and missed detections. III. Multi-dimensional Feature Fusion and Extraction Module Based on the obtained effective action signal segments, this embodiment performs multi-dimensional feature extraction, covering four types of features: Time-domain characteristics: integrated electromyography (IEMG), root mean square (RMS), peak value, and waveform length (WL). Frequency domain characteristics: average power frequency (MPF), median frequency (MF), peak power spectral density (PSD); Time-frequency characteristics: wavelet packet energy (WPE1-WPE8, decomposition level 3). Muscle group synergy characteristics: antagonistic muscle ratio (CR), intermuscular phase difference (IPD1-IPD4); Initially, 32 features were extracted. To reduce computational complexity and improve discriminative ability, this embodiment employs a joint optimization strategy of genetic algorithm + LASSO regression for feature selection. The genetic algorithm was configured with a population size of 50, 30 iterations, a crossover probability of 0.6, and a mutation probability of 0.05, initially selecting 15 effective features. Then, redundant features were eliminated using LASSO regression (λ=0.01), ultimately retaining eight core features: IEMG, RMS, MPF, MF, WPE3, CR, IPD1, and IPD2. These features fully express the discriminative information of electromyography signals in different dimensions, improving the accuracy of continuous motion decoding. IV. CNN-LSTM Hybrid Network Decoding Module For the selected 8 core features, a time-frequency spectrogram (size 128×128×8, 128 points in time dimension, 128 points in frequency dimension, and 8 points in channel dimension) is constructed to represent both time and frequency domain features. Network structure: CNN part: 3 convolutional layers (3×3 kernels, stride 1, padding 1) + 2 pooling layers (2×2 kernels, stride 2), used to extract spatial correlation features; LSTM section: 2-layer bidirectional LSTM, 128 hidden units per layer, used to capture the time-series dependencies of the signal and enhance the ability to decode continuous actions; Output layer: The fully connected layer outputs 6-DOF control commands (3-axis rotation + 3-axis translation) to the medical imaging terminal, adapting to continuous operation requirements; Model training: Pre-training: The model was trained on a dataset of 200 healthy subjects for 50 iterations, with a batch size of 32, a learning rate of 0.001, and the Adam optimizer was used to obtain the base model. Fine-tuning: For the postoperative rehabilitation patient dataset (including brachial plexus injury, MRC 1-4), 20 iterations were performed with a batch size of 16 and a learning rate of 0.0001 to improve the model's adaptability to special populations; V. Scene Adaptive Optimization Module Online incremental learning: Introducing the OIL mechanism, after every 10 valid operations, the feature data of the new operation is automatically extracted and the model parameters are incrementally updated. The learning rate is 1 / 10 of that in the pre-training stage, which enables personalized adaptation and captures user operation habits and changes in electromyography signals without retraining. Scene parameter adaptation: Aseptic operating room mode: Dynamic response speed optimization, motion recognition delay <50ms; real-time monitoring of signal-to-noise ratio, automatically increasing the gain coefficient when the signal-to-noise ratio is <30dB, enhancing anti-interference ability, and ensuring the accuracy and stability of surgical operations; Postoperative rehabilitation mode: Sensitivity is dynamically adjusted according to MRC muscle strength level (MRC 1-2 ×2.0, MRC 3 ×1.5, MRC 4 ×1.0); training is automatically paused when muscle oxygen saturation monitoring (StO2<60%) to ensure rehabilitation safety; Anomaly handling mechanism: If the signal quality is substandard for 5 consecutive seconds (signal-to-noise ratio <25dB or action segment recognition failure rate >30%), the hardware self-test is triggered and the user is prompted to reattach the sensor or adjust its position, effectively avoiding the risk of misoperation and secondary damage. VI. System Connection and Control Process The overall system consists of a multi-channel flexible electromyography (EMG) sensor array, an ADS1298 chip, an FPGA control module, a CNN-LSTM algorithm processing module, and a wireless communication module (Bluetooth 5.0 / BLE). The sensor array adheres to the surface of the muscle group to collect EMG signals. After being digitized by the ADS1298 chip, the signals are transmitted to the FPGA module for sampling rate and gain control. The preprocessed signals are transmitted via a data bus to the CNN-LSTM module for feature extraction and gesture decoding. The decoded control commands are sent to the medical imaging terminal via the wireless communication module, enabling continuous, contactless operation.
[0053] Figure 5 This is the algorithm code of an embodiment of this application.
[0054] The code is explained as follows: Signal acquisition and preprocessing: Bandpass filtering is performed on the signal of each channel using bandpass_filter to remove low-frequency noise and high-frequency interference.
[0055] Feature extraction: Root mean square (RMS) was extracted as a feature. In practice, other time-domain or frequency-domain features can be extracted as needed.
[0056] Classification and Recognition: The extracted features are classified using a Support Vector Machine (SVM) model to distinguish different gestures.
[0057] Gesture decoding: Decode each sample based on the classification results and map it to the corresponding operation (e.g., execute a command).
[0058] Key technical points of this technical solution: 1. Multi-dimensional feature fusion and selection strategy: Integrate time domain, frequency domain, time-frequency and muscle group coordination features, and select 8 core features through joint optimization of "genetic algorithm + LASSO regression" to balance recognition accuracy and model efficiency; 2. CNN-LSTM Hybrid Network Decoding Architecture: Innovatively adopts a combination of spatial feature extraction (CNN) and temporal series feature capture (LSTM) to achieve continuous control command output with 6 degrees of freedom; 3. Scenario-adaptive optimization mechanism: including dynamic sensitivity adjustment based on MRC classification, individual adaptation through online incremental learning, and parameter presets for different scenarios, to improve the clinical adaptability of the technology; 4. Anti-interference preprocessing scheme: The combination strategy of "bandpass filtering + notch filtering + adaptive smoothing" ensures signal quality in complex electromagnetic environments.
[0059] Technical protection points of this technical solution 1. Electromyography signal acquisition and anti-interference design: configuration of multi-channel flexible sensor, preprocessing flow and parameter range of "bandpass filtering + notch filtering + adaptive smoothing"; 2. Feature Fusion and Deep Learning Decoding: Selection of multi-dimensional features and feature screening method of "genetic algorithm + LASSO regression", architecture design of CNN-LSTM hybrid network (including convolutional layer and LSTM layer parameters). 3. Adaptive control for clinical scenarios: sensitivity adjustment rules based on MRC grading, online incremental learning model update mechanism, and parameter adaptation schemes for different scenarios (surgery / rehabilitation); 4. Full-process technical solution: The complete technical chain from signal acquisition, preprocessing, feature extraction, model decoding to control command output, as well as the training method of dual-source data acquisition (self-acquisition + public database).
[0060] Technical alternatives to this technical solution: Based on technical research and analysis, there is currently no completely equivalent alternative that can achieve all the technical effects of this patent. The following is a comparative analysis of two types of partial functional alternatives: 1. Visual recognition-based gesture decoding solution: This solution captures hand gestures through a camera to achieve contactless control. However, it is susceptible to environmental factors such as surgical lighting and obstructions, and it cannot be adapted to patients with postoperative muscle weakness and limited range of motion. In sterile operating room settings, its robustness is far lower than that of electromyography signal recognition solutions. 2. Decoding scheme based on single feature + traditional machine learning (SVM / random forest): This scheme has low model complexity, but it only relies on a single-dimensional feature, resulting in insufficient recognition accuracy and dynamic adaptability. It cannot meet clinical needs in scenarios with signal attenuation or large individual differences, and it does not have online adaptive optimization capabilities. In summary, the above-mentioned alternative solutions are inferior to the patented technology solution in terms of core indicators such as clinical suitability, robustness of identification, and individualized adaptation, and cannot achieve the same technical effect.
[0061] The implementation principle of the gesture decoding algorithm for a contactless medical drawing terminal based on electromyography (EMG) signal recognition in this application embodiment is as follows: This algorithm is based on the principle that the surface electromyography (sEMG) signals generated by the human upper limb muscle groups when performing different gestures and movement commands have characteristic changes. Through multi-dimensional signal processing and deep neural network structure, the mapping of gestures to medical drawing terminal control commands is realized. The potential activity generated by the muscle during contraction propagates along the muscle fibers and forms an 8-channel EMG signal after being detected by a flexible electrode array. The amplitude, spectral distribution, time-frequency texture and intermuscular coordination pattern of the electrical activity of different muscle groups in flexion, extension, pronation, supination and micro-manipulation movements are significantly different, which constitute the basic information source for action recognition. To extract stable and distinguishable motion features from the raw electromyography (EMG) signal, bandpass filtering and notch filtering are first used to remove low-frequency baseline drift, power frequency interference and high-frequency noise. Then, sliding window smoothing and normalization are used to improve the temporal consistency and cross-user comparability of the signal. After multi-threshold motion segment detection, EMG segments corresponding to a single effective motion are obtained. Based on the multi-scale characteristics of muscle electromyography, this application constructs four categories of features: time domain, frequency domain, time-frequency, and coordination. A two-stage screening is performed using a genetic algorithm and LASSO regression to optimize the 8 core features that best represent the differences in movement from 32 initial features. Subsequently, the multi-channel features are converted into time-spectrum maps to preserve the complete structural information of electromyographic signals in the time, frequency, and muscle group dimensions. In the decoding stage, a CNN-LSTM hybrid deep network is constructed: the convolutional neural network is responsible for extracting local spatial texture features in the time-spectral graph, the bidirectional LSTM network captures the time dependence of electromyographic signals and the muscle activation sequence pattern, and finally outputs 6-DOF control commands corresponding to the medical image operation terminal through the fully connected layer to realize continuous and real-time decoding from action to spatial operation. Furthermore, based on the dynamic changes in electromyographic signals due to fatigue, adhesion, and individual differences, an online incremental learning mechanism is introduced. This mechanism updates model parameters based on real-time user data without requiring retraining, improving the long-term stability of the system. Additionally, based on the different needs of sterile operating rooms and postoperative rehabilitation scenarios, the gain, sensitivity, and signal-to-noise ratio thresholds are dynamically adjusted to ensure the algorithm maintains high robustness and safety in various clinical environments. This enables high-precision, contactless mapping from multi-muscle group electrophysiological signals to multi-degree-of-freedom medical drawing terminal control commands, providing reliable human-computer interaction capabilities for surgical drawing, intraoperative annotation, and rehabilitation assistance.
[0062] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A gesture decoding algorithm for a contactless medical drawing terminal based on electromyography signal recognition, characterized in that: Includes the following steps, Step 1: Multi-channel electromyography (EMG) signal acquisition: EMG signals are acquired by an 8-channel flexible EMG sensor array attached to the surface of muscle groups such as the biceps brachii, flexor carpi radialis, pronator, and supinator. The EMG signals are digitized by the ADS1298 bioelectric chip, and the FPGA control module dynamically adjusts the sampling frequency and gain coefficient of ×1000 to ×5000 within the range of 500 to 2000 Hz according to the needs of the scenario. Step 2: Anti-interference preprocessing: Bandpass filtering (20-500Hz), 50Hz notch filtering, adaptive sliding window smoothing, and Z-score normalization are performed on the digitized electromyography signal. The start and end points of the action segment are detected based on energy threshold, length constraint, waveform features, and slope threshold to obtain the effective action signal segment. Step 3: Multi-dimensional feature fusion and extraction: Extract time-domain features, frequency-domain features, time-frequency features and muscle group coordination features from the effective action signal segment, and obtain 8 core features including IEMG, RMS, MPF, MF, WPE3, CR, IPD1 and IPD2 from 32 initial features through a joint screening method of genetic algorithm and LASSO regression. Step 4: CNN-LSTM hybrid network decoding: The 8-channel core features are constructed into a time-spectrum map with a size of 128×128×8. Spatial features are extracted through a CNN network including three convolutional layers and two pooling layers, and time-series features are extracted through a two-layer bidirectional LSTM network. Finally, the 6-DOF control commands corresponding to the medical imaging terminal are output. Step 5: Scene Adaptive Optimization: The model parameters are updated based on new feature data during user use using an online incremental learning method. The dynamic response speed, signal-to-noise ratio threshold, and control sensitivity are adjusted in the sterile operating room mode and postoperative rehabilitation mode, respectively. When the muscle oxygen saturation is lower than the preset threshold or the signal quality is substandard, control pause and prompt are triggered.
2. The algorithm according to claim 1, characterized in that: In the detection of the start and end points of the action segment, the slope threshold is set to a signal rise or fall slope greater than 0.5 mV / s.
3. The algorithm according to claim 1, characterized in that: The sliding window smoothing process uses a moving average method with a window size of 0.2s.
4. The algorithm according to claim 1, characterized in that: The genetic algorithm has a population size of 50, an iteration count of 30, a crossover probability of 0.6, and a mutation probability of 0.
05.
5. The algorithm according to claim 1, characterized in that: The time and frequency dimensions of the constructed time spectrum diagram both use a resolution of 128 points.
6. The algorithm according to claim 1, characterized in that: The CNN has a 3×3 convolution kernel size, a stride of 1, and a padding of 1. The pooling layer uses a 2×2 pooling kernel.
7. The algorithm according to claim 1, characterized in that: The LSTM network has a bidirectional structure and each layer contains 128 hidden units.
8. The algorithm according to claim 1, characterized in that: The online incremental learning performs a model parameter update once after every 10 valid operations, and the update uses 1 / 10 of the pre-training learning rate.
9. The algorithm according to claim 1, characterized in that: In sterile operating room mode, the gain factor is automatically increased to enhance signal quality when the signal-to-noise ratio is below 30dB.
10. The algorithm according to claim 1, characterized in that: In the postoperative rehabilitation mode, the sensitivity is adjusted according to the MRC muscle strength grade, with MRC grade 1-2 set to ×2.0, MRC grade 3 set to ×1.5, and MRC grade 4 set to ×1.
0. When the signal-to-noise ratio is below 25dB for 5 consecutive seconds or the failure rate of motion segment recognition exceeds 30%, the system triggers a hardware self-test and prompts the user to reattach the sensor or adjust its position.