Muscle stiffness symptom quantitative evaluation method based on machine learning and related equipment
By using a machine learning-based method and constructing a model using arm joint resistance, electromyography, and tremor signals, the subjective problem of myotonia diagnosis and assessment was solved, objective quantitative assessment of myotonia symptoms and early risk identification were achieved, and the consistency and repeatability of the assessment were improved.
Patent Information
- Application Number
- CN202510767399.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-10
AI Technical Summary
The diagnosis and symptom assessment of myotonia mainly rely on the subjective judgment of clinicians and lack objective quantitative standards, which leads to inconsistencies in diagnosis and assessment, affecting the formulation of treatment plans and the evaluation of efficacy.
A machine learning-based method was used to collect patients' physiological signal feature data, including arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals. XGBOOST, LightGBM, and random forest models were constructed for training and validation to obtain a quantitative assessment of the target patients' myotonia symptoms.
It achieves objective quantitative assessment of myotonia symptoms, avoids physician experience and subjective bias, improves the consistency and repeatability of the assessment, enables early identification of myotonia risks, and supports dynamic monitoring of treatment effects.
Smart Images

Figure CN120753590A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of smart medical care. More specifically, the present invention relates to a method for quantitatively evaluating myotonia symptoms based on machine learning and related equipment. Background Art
[0002] Myotonia is a chronic neuromuscular disease that primarily affects middle-aged and elderly individuals. Typical symptoms include increased muscle tone, stiffness, and twitching. These symptoms not only severely impact patients' quality of life but also place a heavy burden on their families and society. Currently, the diagnosis and symptom assessment of myotonia rely primarily on subjective judgment by clinicians, lacking objective quantitative criteria. This subjectivity can lead to inconsistencies in diagnosis and assessment, which in turn impacts treatment planning and efficacy assessment. Summary of the Invention
[0003] The Summary of the Invention introduces a series of simplified concepts that will be further described in the Detailed Description of the Invention. The Summary of the Invention is not intended to limit the key features and essential features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.
[0004] To address the current problem that the diagnosis and symptom assessment of myotonia mainly rely on the subjective judgment of clinicians and lack objective quantitative standards. This subjectivity may lead to inconsistencies in diagnosis and assessment, which in turn affects the formulation of treatment plans and the evaluation of efficacy. In a first aspect, the present invention proposes a method for quantitative assessment of myotonia symptoms based on machine learning, comprising: Collecting a preliminary data set of the patient, the preliminary data set includes physiological signal feature data and label data, the feature data includes arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals, and the label data is the scoring data obtained by the doctor based on the patient's condition analysis; The data set is divided into a training set and a test set, so as to train an evaluation model with the training set, and to verify and optimize the accuracy of the trained evaluation model with the test set, wherein the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyography signals, and a random forest model constructed based on arm joint tremor signals; Acquire physiological signal characteristic data of a target patient, and quantitatively evaluate the myotonia symptom of the target patient through an optimized evaluation model based on the physiological signal characteristic data.
[0005] Optionally, before the step of dividing the data set into a training set and a test set, the method further includes: The preliminary data set is analyzed by principal component analysis, and the linear combination principal components with the largest amount of information are retained to obtain key feature data and reduce feature dimensions.
[0006] Optionally, the evaluation model is trained using a training set, including The sample distribution is adjusted through the oversampling strategy, and three types of models are trained separately to extract the discriminant rules of the arm joint resistance data, arm joint electromyographic signals and arm joint tremor signal features.
[0007] Optionally, also include: A weighted fusion mechanism is used to integrate the prediction results of the three models, where the weights are set according to the AUC values of the three models in the test set.
[0008] Optionally, also include: After the evaluation model training is completed, the trained evaluation model is applied to independent test sets that have not been used in the training process and do not overlap with the training data. The mean absolute error and recall rate are used as performance evaluation indicators. The mean absolute error is used to measure the average deviation between the model prediction score and the doctor's annotation score, and the recall rate is used to measure the model's sensitivity to identifying patients with high-level symptoms.
[0009] Optionally, before the step of dividing the data set into a training set and a test set, the method further includes: Filtering and denoising the arm joint electromyographic signal data to obtain clean electromyographic signal data; The arm joint resistance data and arm joint tremor signal data were standardized to eliminate data deviations between different patients.
[0010] Optionally, before the step of dividing the data set into a training set and a test set, the method further includes: The characteristic information of the arm joint resistance data, the arm joint electromyographic signal data and the arm joint tremor signal data is extracted respectively using time domain and frequency domain analysis methods, wherein the characteristic information includes at least one of the mean, variance and frequency component.
[0011] In a second aspect, the present invention further proposes a device for quantitatively evaluating myotonia symptoms based on machine learning, comprising: a preprocessing unit, configured to collect a preliminary data set of the patient, the preliminary data set including physiological signal feature data and label data, the feature data including arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals, and the label data being scoring data obtained by a doctor based on an analysis of the patient's symptoms; a model optimization unit, configured to divide the data set into a training set and a test set, so as to train an evaluation model using the training set, and to verify and optimize the accuracy of the trained evaluation model using the test set, wherein the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyography signals, and a random forest model constructed based on arm joint tremor signals; The evaluation unit is used to obtain physiological signal characteristic data of a target patient, and to quantitatively evaluate the myotonia symptom of the target patient through an optimized evaluation model based on the physiological signal characteristic data.
[0012] In a third aspect, an electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the steps of the method for quantitatively assessing myotonia symptoms based on machine learning as described in any one of the first aspects above when executing the computer program stored in the memory.
[0013] In a fourth aspect, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for quantitatively assessing myotonia symptoms based on machine learning according to any one of the above-mentioned items in the first aspect.
[0014] In summary, the machine learning-based quantitative assessment method for myotonia symptoms proposed in this application collects a preliminary patient data set, the preliminary data set including physiological signal feature data and label data. The feature data includes arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals. The label data is the score data obtained by the doctor based on the patient's symptom analysis; the data set is divided into a training set and a test set, and the evaluation model is trained with the training set, and the accuracy of the trained evaluation model is verified and optimized with the test set. The evaluation model includes an XGBOOST model built based on arm joint resistance data, a LightGBM model built based on arm joint electromyographic signals, and a random forest model built based on arm joint tremor signals; the physiological signal feature data of the target patient is obtained, and the target patient's myotonia symptoms are quantitatively assessed using the optimized evaluation model based on the physiological signal feature data. Automatic quantification through machine learning avoids subjective bias caused by physician experience, mood, and attention fluctuations. The high correlation between the physiological features learned in the training set and the clinical labels makes the assessment more repeatable. Because it uses feature-level detection, even minor abnormalities in electromyographic discharges can be captured by the LightGBM model, indicating the risk of early myotonia. Quantitative changes in scores can be directly used to assess the effectiveness of rehabilitation or drug interventions.
[0015] The machine learning-based quantitative assessment method for myotonia symptoms of the present invention, and other advantages, objectives, and features of the present invention will be partially reflected in the following description and partially understood by those skilled in the art through research and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present description. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 A flowchart of a method for quantitatively assessing myotonia symptoms based on machine learning provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a device for quantitatively assessing myotonia symptoms based on machine learning provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of an electronic device for quantitatively evaluating myotonia symptoms based on machine learning provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the embodiments described are only some embodiments of the present application, not all embodiments.
[0018] Currently, the diagnosis and symptom assessment of myotonia primarily rely on the subjective judgment of clinicians, lacking objective quantitative standards. This subjectivity can lead to inconsistencies in diagnosis and assessment, which in turn affects the formulation of treatment plans and the evaluation of efficacy. With advances in medical technology, a growing number of studies are focusing on how to use objective physiological signals to quantify myotonia symptoms. For example, physiological data such as electromyography (EMG), joint resistance data, and tremor signals can reflect the degree of a patient's motor dysfunction. However, traditional signal processing methods often rely on manual feature extraction and simple statistical analysis, making it difficult to fully capture the complex physiological signal characteristics and susceptible to noise and individual differences.
[0019] In order to solve the above problems, a quantitative assessment method of myotonia symptoms based on machine learning is provided. Figure 1 , is a flow chart of a method for quantitatively evaluating myotonia symptoms based on machine learning provided in an embodiment of the present application, which may specifically include: steps S110 to S130.
[0020] S110, collecting a preliminary data set of the patient, the preliminary data set includes physiological signal feature data and label data, the feature data includes arm joint resistance data, arm joint electromyographic signals and arm joint tremor signals, and the label data is the scoring data obtained by the doctor based on the analysis of the patient's symptoms.
[0021] S120, dividing the data set into a training set and a test set, so as to train the evaluation model through the training set, and to verify and optimize the accuracy of the trained evaluation model through the test set, wherein the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyography signals, and a random forest model constructed based on arm joint tremor signals.
[0022] S130 , acquiring physiological signal characteristic data of a target patient, and quantitatively evaluating the myotonia symptom of the target patient through an optimized evaluation model based on the physiological signal characteristic data.
[0023] For example, to achieve objective quantitative assessment of myotonia symptoms, in a specific embodiment, a preliminary patient data set is first collected through a standardized process. The data collection phase focuses on the patient's arm joints, sequentially collecting three types of physiological signals: joint motion resistance data, electromyographic (EMG) signals, and tremor signals. Resistance data collection utilizes an electronic joint resistance measurement device equipped with a high-precision force sensor. Flexion and extension movements are performed under set angle and speed control conditions, recording the applied torque and angle curves throughout the joint movement, thereby capturing abnormally high resistance due to myotonia. EMG signals are collected by non-invasively attaching surface electromyography electrodes to target muscle groups (such as the biceps and triceps) during joint movement. The electrical discharge activity generated by the muscles at different stages is synchronously recorded, and the amplitude and frequency characteristics are extracted to reflect the physiological mechanisms of muscle stiffness and involuntary electrical discharge. Tremor signal collection utilizes EMG to monitor subtle changes in arm vibrations, capturing involuntary tremor. After these multi-source signals are collected, senior expert physicians (for example, at least three) use a unified scale (such as the modified Ashworth scale) to independently evaluate the degree of muscle rigidity of each patient, and take the average as the label data to ensure the objectivity and consistency of the label and form a high-quality training data set.
[0024] For example, after completing preliminary data collection, the dataset was partitioned into a training set and a test set in an 8:2 ratio. During the partitioning process, the distribution ratio of samples of different myotonia severity levels (mild, moderate, and severe) was maintained in both groups to avoid model bias. Separate machine learning sub-models were constructed for each signal type, tailored to the characteristics of each signal type. For resistance data, the XGBoost model was selected due to its excellent performance in handling nonlinear feature interactions and ranking feature importance, making it suitable for capturing the complex rise / fall rate variations found in myotonia resistance curves. When inputting resistance data, preprocessing was performed, including smoothing and denoising, derivative extraction (for velocity and acceleration), and feature engineering (for maximum resistance, average resistance, rate of change of resistance, and area integral under the torque curve) to enrich the information available to the model. For electromyographic signals, a LightGBM model was constructed. Because electromyographic signals have high-dimensional and sparse features in the time and frequency domains (such as zero-crossing rate variations and short-term energy spikes), LightGBM can quickly learn important features through its efficient splitting strategy, reducing the risk of overfitting. EMG signal feature extraction includes time-domain statistical features (RMS, mean, standard deviation) and frequency-domain features (dominant frequency, spectral entropy, and power spectral density) to more comprehensively reflect abnormal muscle discharge patterns. For tremor signals, a random forest model is constructed. Given the strong volatility and local irregularities of tremor signals, the random forest multi-decision tree ensemble approach can effectively handle nonlinear relationships between features and improve robustness. Input features include mean tremor amplitude, dominant tremor frequency, tremor energy distribution, and integrated acceleration.
[0025] For example, during the training process of each sub-model, training and performance evaluation are performed using a five-fold cross-validation approach to minimize performance fluctuations caused by accidental data partitioning. Furthermore, hyperparameter tuning is employed for each model (e.g., learning rate, maximum tree depth, and sub-sample ratio for XGBoost; number of leaves and maximum depth for LightGBM; and number of decision trees and feature limits for Random Forest) to achieve optimal predictive performance. Model performance is comprehensively evaluated using mean absolute error (MAE), mean squared error (MSE), and coefficient of determination (R²). Based on the optimized model, a fusion inference mechanism is constructed to perform a weighted average of the outputs of each sub-model. Weights are adjusted based on the performance of each model on the training set. For example, the EMG and resistance data models are each given a 40% weight due to their superior performance, while the tremor signal model is given a 20% weight. The final output is a unified quantitative score for myotonia symptoms, accompanied by a confidence interval to indicate the reliability of the score.
[0026] It is understandable that the nonlinear mapping learning of high-dimensional complex physiological signal features and clinical score labels based on machine learning can systematically extract, summarize and generalize the physiological manifestation patterns of myotonia in different patients, and when applied to new patients, quickly and objectively infer the severity of myotonia based on their physiological signal characteristics. Compared with traditional evaluation methods based on experience and observation, this solution greatly reduces human subjectivity and improves the consistency and repeatability of the evaluation; at the same time, due to the continuity of quantitative output, it can also dynamically monitor small changes in the patient's treatment process, support early detection of symptom progression or improvement, and thus guide more accurate treatment decisions. For example, during the patient's rehabilitation process, continuous monitoring of the maximum resistance value decreased, the electromyography RMS decreased, and the tremor amplitude decreased. The comprehensive score dropped from 3.5 in the initial stage to 2.4 points, objectively and quantitatively reflecting the positive effect of the treatment.
[0027] It is understandable that when the data is complete and of reliable quality, the other steps of feature extraction can be directly performed to extract key features, which will be used to build the model.
[0028] For example, the resistance data is modeled using the XGBoost model. XGBoost excels at handling nonlinear feature interactions and complex time series patterns. Resistance data (such as torque curves and resistance rate of change) often contain dynamic nonlinear relationships (such as gear-like resistance characteristics). The model's tree structure automatically captures these high-order interactions. Furthermore, XGBoost supports regularization (L1 / L2), which effectively prevents overfitting and improves model generalization. This is particularly useful in scenarios where clinical data may be noisy or contain small sample sizes.
[0029] The feature engineering of resistance data involves time-domain derivatives (velocity, acceleration) and frequency-domain integrals (area under the curve), and the XGBoost feature importance ranking (such as SHAP values) can intuitively explain the influence of key features (such as the maximum resistance value) on the score, helping doctors understand the basis for model decision-making.
[0030] The input features include time-domain features (mean, variance, RMS, ZCR) and frequency-domain features (FFT energy distribution) of resistance data.
[0031] Feature standardization:
[0032] wherein, and is the mean and standard deviation of the training set, and the principal components that retain the top 90% of the variance contribution rate are retained by PCA, and x is the original sample feature.
[0033] The model training uses mean square error (MSE) as the loss function:
[0034] wherein, is the resistance data, is the weighted coefficient, λ is the regularization coefficient, N is the number of samples.
[0035] Model output: The patient's symptom score (0-4 points), and the feature importance is analyzed by SHAP value:
[0036] wherein, L is the number of tree structures of the model, is the feature weight.
[0037] It can be understood that the early manifestation of muscle stiffness is that the muscle produces significantly increased resistance when passively pulled, and the resistance rises nonlinearly, often accompanied by a gear-like phenomenon. Through RMS and variance, the intensity of resistance signal fluctuations within a unit of time can be reflected, thereby quantifying the muscle counter-tension trend. ZCR can reflect whether the muscle repeatedly produces reactive contraction phenomenon during passive movement, such as frequent stuttering. The FFT energy is concentrated in a specific low frequency (<5 Hz), which can reveal the increase of viscous resistance component, which is a characteristic of chronic muscle tension disorder. XGBoost can automatically learn the nonlinear combination between features and assign importance weights, which is suitable for identifying resistance change patterns under individual differences. SHAP values reveal the relationship between specific features and predicted scores, which helps to explain the basis for model decision-making and assist doctors in developing differentiated intervention plans.
[0038] For example, the electromyographic signal uses a LightGBM model.
[0039] LightGBM's histogram algorithm and gradient-based single-edge sampling (GOSS) efficiently process high-dimensional sparse features. The frequency-domain binning features of electromyographic signals are highly sparse, allowing LightGBM to quickly filter out important frequency bands (such as the 20-250Hz abnormal discharge frequency band, the primary frequency distribution of electromyographic signals). LightGBM supports parallel training and GPU acceleration, making it suitable for large-scale feature extraction of long-duration, high-frequency signals.
[0040] The sparsity of EMG signal features (such as short-term energy spikes) requires efficient feature selection. LightGBM's leaf-wise growth strategy prioritizes splitting nodes with the highest gain, accurately capturing abnormal discharge patterns in EMG signals (such as RMS spikes), improving sensitivity for early symptom identification.
[0041] Input features: Frequency domain energy proportion (20-500Hz), dominant frequency component, and time domain zero-crossing rate of the EMG signal. Feature encoding: Binning statistics of the frequency domain energy distribution (e.g., the mean energy value of 20 frequency bands).
[0042] Model training uses log loss (Log Loss) combined with AUC optimization:
[0043] in, is the electromyography data processing value, is the predicted probability.
[0044] Histogram algorithm: discretizes continuous features into histogram buckets to speed up split calculations:
[0045] The model output is a symptom score, sorted by feature importance (based on split gain):
[0046] in, is the gradient statistic of node j in tree t, , is the node weight.
[0047] It is understandable that myotonia often causes abnormal sustained discharges (especially when the EMG is continuously highly active when the posture is maintained), with the main frequency concentrated in the 100-250Hz band. An increase in the proportion of high-frequency energy indicates abnormal discharge synchronization and can quantify muscle overactivation. A shift in the proportion of low-frequency energy or a decrease in the main frequency indicates a trend of weak muscle strength and denervation. ZCR can measure the "strayness" and waveform irregularity of electrical signals, reflecting the loss of control of myoelectric patterns. LightGBM is suitable for processing high-dimensional sparse features (such as frequency band binning) and can stably mine subtle differences in electrical activity patterns. Using the feature importance ranking based on gain sorting, it can be determined which frequency bands are more sensitive to different severities, supporting electrophysiological targeted treatment.
[0048] Exemplarily, a random forest model is used for the tremor signal.
[0049] Random forests reduce variance by integrating multiple decision trees, making them robust to local fluctuations and noise (such as environmental vibration interference) in tremor signals. Random forest models naturally support mixed feature types (such as time domain mean, frequency domain dominant frequency, and wavelet packet energy), eliminating the need for strict data distribution assumptions and making them suitable for multi-scale analysis of tremor signals (e.g., short-duration strong tremor versus chronic oscillatory background).
[0050] The asymmetry (e.g., differences in left and right muscle contraction) and randomness (e.g., sudden acceleration) of tremor signals require models to consider both global and local features. The bagging strategy and random feature subset selection of random forests can effectively discover discriminative patterns (e.g., dominant frequency shift) in tremor signals, preventing overfitting of a single decision tree.
[0051] Input features include: amplitude (RMS), dominant frequency, and time domain asymmetry index of the tremor signal:
[0052] in, , is the characteristic quantity of right tremor.
[0053] Feature enhancement: Extract multi-scale energy features through wavelet packet decomposition.
[0054] Feature selection: Screening features with high correlation with data labels based on ANOVA F value:
[0055] in, is the between-class variance, is the intra-class variance.
[0056] The model training adopts an ensemble strategy: 200 decision trees are constructed, and each tree uses Bootstrap sampling (with replacement sampling).
[0057] Splitting criterion: based on Gini impurity
[0058] Among them, p k is the proportion of category k in the node, and K is the number of categories.
[0059] The model output is the score prediction value, and the generalization performance is evaluated by OOB (Out-of-Bag) error:
[0060] Among them, L is the loss function, N oob is the number of OOB samples.
[0061] Understandably, patients with myotonia often experience postural tremor, whose dominant frequency distribution generally ranges from 4–8 Hz, and whose RMS amplitude can be significantly increased. A high asymmetry index indicates asynchronous contraction of the left and right muscle groups, a key manifestation of abnormal neural control pathways. Wavelet packet analysis is used to extract energy features from different frequency bands, effectively distinguishing short, intense tremor episodes from a chronic oscillatory background. F-value feature selection ensures a significant statistical correlation between input and score labels, eliminates the influence of invalid variables, and improves model stability. Random forests naturally support multiple feature scales and types (time-frequency mixing, wavelet coefficients), making them suitable for modeling complex electromyographic motion features. Out-of-band (OOB) evaluation reflects the model's performance on samples outside the training set, enhancing its ability to generalize tremor patterns.
[0062] In summary, the embodiment of the present application provides a method for quantitatively evaluating myotonia symptoms based on machine learning, which collects a preliminary data set of patients, wherein the preliminary data set includes physiological signal feature data and label data, wherein the feature data includes arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals, and the label data is the scoring data obtained by the doctor based on the analysis of the patient's symptoms; the data set is divided into a training set and a test set, so as to train the evaluation model through the training set, and to verify and optimize the accuracy of the trained evaluation model through the test set, wherein the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyographic signals, and a random forest model constructed based on arm joint tremor signals; the physiological signal feature data of the target patient is obtained, and the myotonia symptoms of the target patient are quantitatively evaluated based on the physiological signal feature data through the optimized evaluation model. Automatic quantification through machine learning avoids subjective bias caused by doctor experience, mood, and attention fluctuations. The high correlation between the physiological features learned in the training set and the clinical labels makes the evaluation more repeatable. Because it uses feature-level detection, even minor abnormalities in electromyographic discharges can be captured by the LightGBM model, indicating the risk of early myotonia. Quantitative changes in scores can be directly used to assess the effectiveness of rehabilitation or drug interventions.
[0063] According to some embodiments, before the step of dividing the data set into a training set and a test set, the method further includes: The preliminary data set is analyzed by principal component analysis, and the linear combination principal components with the largest amount of information are retained to obtain key feature data and reduce feature dimensions.
[0064] For example, principal component analysis (PCA) can be performed on the preliminary dataset. As a linear dimensionality reduction technique, feature data from diverse sources, multiple dimensions, and potentially containing redundancy or noise are first standardized (e.g., by returning the mean to zero and normalizing the variance). The covariance matrix is then calculated and its eigenvalues and eigenvectors are solved. The principal components that explain the highest proportion of the overall variance are then selected as new key features. This approach maximizes the information contained in the original high-dimensional features in a lower dimensional manner, eliminating multicollinearity between features and improving the training efficiency and generalization of subsequent machine learning models. For example, raw resistance data may contain torque values at hundreds of time points. PCA can extract 3-5 comprehensive eigenvectors that represent the main trends of variation, significantly reducing feature dimensionality while retaining the key information. After PCA, the reduced dataset is then partitioned into training and test sets in an 8:2 ratio, ensuring a balanced distribution of samples across different myotonia severities to avoid training bias.
[0065] It is understandable that the principal component analysis step significantly improves the sensitivity of the machine learning model to key physiological change patterns by retaining the most representative linear combination features, suppresses noise interference, reduces the risk of overfitting, and at the same time shortens training time and improves the speed of model convergence. Each sub-model is specifically modeled for different signals, which not only fully utilizes the advantages of a specific model in fitting specific data characteristics, but also makes up for the limitations of single signal evaluation through a fusion mechanism, ultimately making the myotonia quantitative score evaluated by the system more objective, stable, and clinically valuable. In terms of technical effect, this method can significantly improve the consistency of myotonia symptom assessment, achieve early identification of disease progression, and provide quantitative support for personalized rehabilitation plan adjustment and dynamic tracking of treatment efficacy, thereby improving the overall chances of improving patient prognosis and reducing the burden of family and social care.
[0066] In some examples, the training of the evaluation model using the training set includes The sample distribution is adjusted through the oversampling strategy, and three types of models are trained separately to extract the discriminant rules of the arm joint resistance data, arm joint electromyographic signals and arm joint tremor signal features.
[0067] For example, after the dataset is partitioned and before formal training, an oversampling strategy can be added to adjust the sample distribution. Because data distribution across different myotonia severities (e.g., mild, moderate, and severe) within a patient population is often uneven, with more mild samples and fewer severe samples, direct training can lead to overfitting the model to the majority class and insufficient recognition of the minority class. To address this, oversampling methods, particularly strategies like SMOTE (Synthetic Minority Over-sampling Technique), are employed to generate new synthetic samples in the feature space based on existing minority class samples. This increases the number of minority class samples without simply copying the original samples, achieving a balanced sample size across all classes. This strategy effectively alleviates the class imbalance issue, enabling the model to more fully learn the characteristic patterns of mild, moderate, and severe myotonia during training, improving discrimination capabilities and reducing the risk of errors caused by class bias. After achieving sample balance, during training, a sub-model is further trained on each physiological signal type to extract specific discriminative patterns. Specifically, the XGBoost model was trained on joint resistance data to extract discriminant features that reflect muscle stiffness and the degree of joint mobility restriction, such as the maximum value, rate of rise, and area integral of the resistance curve. The LightGBM model was trained on electromyographic signals to identify discriminant features related to muscle stiffness and abnormal discharge, such as RMS value, dominant frequency in the frequency domain, and spectral energy distribution. The random forest model was trained on tremor signals to extract discriminant features closely related to tremor, such as dominant frequency, amplitude change rate, and acceleration change energy. Each sub-model establishes the most adaptable discriminant model based on the mapping between the deep learning essential features of its specific physiological signal type and clinical scores.
[0068] Understandably, oversampling adjusts the sample distribution to improve the class balance of the training data, allowing the model obtained from subsequent training to have balanced discrimination capabilities across different levels of stiffness, rather than being biased towards classes with larger sample sizes. The strategy of extracting discriminative rules separately maximizes the advantages of each physiological signal in characterizing disease characteristics. This allows the overall system to integrate the system, rather than relying on a single signal for one-sided judgment, based on comprehensive reasoning from multiple signals and perspectives, thereby improving the accuracy, stability, and early sensitivity of quantitative assessments.
[0069] In some examples, this also includes: A weighted fusion mechanism is used to integrate the prediction results of the three models, where the weights are set according to the AUC values of the three models in the test set.
[0070] To further enhance the overall accuracy and stability of the final quantitative assessment of myotonia symptoms, a weighted fusion mechanism was added to integrate the predictions of the three sub-models (i.e., XGBoost, LightGBM, and Random Forest models trained on arm joint resistance data, electromyographic signals, and tremor signals, respectively) after they completed their training and prediction outputs. Specifically, rather than using simple averaging, the weights of each sub-model in the fusion process were set based on their Area Under the Curve (AUC) performance on the test set.
[0071] For example, in practice, each sub-model is independently evaluated by calculating its AUC value on the test set. The AUC value reflects the overall predictive ability of the model for different categories (i.e., different myotonia severity score intervals), with values closer to 1 indicating more accurate model discrimination. Therefore, the AUC values of each sub-model are normalized and used as weighting coefficients for their respective prediction results, leading to a linear weighted fusion. This fusion method, based on dynamic AUC weighting, better reflects the true differences in predictive ability of each sub-model on the actual test data, compared to traditional simple averaging or fixed-weight fusion. It prioritizes the contribution of higher-performing models and suppresses error propagation from weaker models, resulting in a more accurate and robust overall combined output. Furthermore, because the AUC considers the model's overall discriminative performance at all possible thresholds, rather than just a single decision point, its use as a weighting factor can cover the prediction needs of the model under different symptom cutoffs in actual use, improving the model's adaptability in clinical applications.
[0072] For example, a weighted voting mechanism is used to assign weights to each model based on its AUC value in the validation set:
[0073] in, is the AUC value of the ith model.
[0074] Final Rating:
[0075] in, is the XGBoost model weight, Score the predictions output by the XGBoost model on the test set, is the LightGBM model weight, Score the predictions output by the LightGBM model on the test set. is the random forest model weight, Score the predictions output by the random forest model on the test set, It is a rounding function that outputs the final discrete integer score symptom level score.
[0076] The actual severity of the patient's symptoms is not necessarily linearly related to the three parameter characteristics, and the degree of symptom manifestation of different patients has certain specificity. Therefore, static weights (based on AUC) cannot dynamically adapt to the differences in characteristic distributions of different patients (for example, some patients have more significant tremor signals, while other patients have more critical resistance data).
[0077] Model fusion uses stacking generalization to meet the problem of patient symptom specificity and further improve diagnostic accuracy. The specific method is as follows: For primary model training, XGBoost, LightGBM, and Random Forest can be used on the training set to generate predictions. Meta-feature construction can then be performed, combining the primary model's predictions with the original features to construct a meta-training set. Meta-model training can then be performed using a neural network as a meta-model to learn the combination patterns of the primary model's predictions. For the final prediction, the primary model's predictions for the target patient can be input into the meta-model to output the final symptom score.
[0078] Compared to direct weighting, the stacking fusion model can learn the contribution weights of different models in different feature scenarios (for example, for patients with tremor-dominant syndrome, the weight of random forest predictions is automatically increased). The neural network meta-model can also capture the complex interactions between primary models, surpassing the limitations of linear weighted fusion.
[0079] Primary model training and meta-feature construction can include: The training set Divide into K-folds to perform data partitioning Cross-validation analysis: For each base model , conduct K-fold training in sequence, and at the k-th fold, use Training the model , and in Generate predicted values .
[0080] Finally, each sample The predicted values of M primary models will be obtained { }, forming the meta-feature matrix .
[0081] The mth model pairs sample Predicted value:
[0082] in, Represents the prediction result of the m-th model on the k-fold.
[0083] During meta-model training: The meta-feature matrix Z can be combined with the true label Y to obtain ,in ] Construct a meta-training set.
[0084] Train the metamodel, select metamodel F, The goal of training is to minimize the prediction error.
[0085] The objective function of the meta-model adopts ridge regression:
[0086] in, is the weight of the mth primary model, is the regularization coefficient.
[0087] The final prediction process may include: All training data can be used Retrain each base model , and get the final model , and conduct full training of the primary model.
[0088] Patient target data ,pass Generate predicted values , forming the element feature vector , generate meta-test features.
[0089] Will Input metamodel F and perform metamodel inference to output the final symptom score:
[0090] Stacking significantly outperforms static weighted fusion by dynamically learning the combined weights of primary models and introducing nonlinear mapping.
[0091] In some examples, this also includes: After the evaluation model training is completed, the trained evaluation model is applied to independent test sets that have not been used in the training process and do not overlap with the training data. The mean absolute error and recall rate are used as performance evaluation indicators. The mean absolute error is used to measure the average deviation between the model prediction score and the doctor's annotation score, and the recall rate is used to measure the model's sensitivity to identifying patients with high-level symptoms.
[0092] Understandably, to further validate the practical application of the evaluation model, after training, a separate dataset, independent of the training set, was set up for testing. Specifically, the trained evaluation model was applied to an independent test set that had never been used during training and had no overlap with the training data. This step ensured objectivity in the model evaluation and verified its generalization ability, avoiding overestimation of performance due to information leakage or overfitting during training.
[0093] For example, during this independent testing phase, model performance can be comprehensively evaluated using the mean absolute error (MAE) and recall metrics. The mean absolute error (MAE) primarily measures the average deviation between the model's predicted myotonia score and the standard score manually annotated by physicians. It is defined as the average of the absolute differences between the predicted scores and the true labeled scores for all test samples. A smaller MAE indicates lower overall numerical deviation in the model's predictions, greater consistency between the scores and the physician annotations, and greater credibility in the evaluation system. Meanwhile, recall, in this scenario, is specifically used to measure the model's sensitivity in identifying patients with high-severity symptoms (such as moderate to severe myotonia). Specifically, given a set score threshold (e.g., a physician score of 3 or higher is considered severe), recall represents the proportion of all actual high-severity patients that the model successfully identified. A higher recall indicates a stronger model's ability to detect high-risk and severe patients, enabling more effective clinical decision-making for early intervention and treatment.
[0094] Understandably, mean absolute error primarily focuses on the model's overall predictive accuracy, reflecting the system's numerical fit in continuous variable (symptom score) regression tasks. Recall, on the other hand, focuses more on the model's coverage of critical cases in real-world use. Especially in medical applications, a high recall means fewer missed critically ill patients, thereby reducing potential risk. Therefore, combining these two metrics not only measures overall predictive accuracy but also highlights the system's performance in detecting critical cases, providing a dual guarantee. Because the MAE of an independent test set is typically slightly higher than that of the training set, it more accurately reflects application effectiveness, making the overall system evaluation more objective and generalizable. It can clearly quantify the system's recognition performance in a clinically important population (high-risk patients), ensuring that sensitivity for detecting critically ill patients is prioritized in treatment decisions. It also supports the subsequent dynamic optimization of fusion weights or adjustment of scoring cutoffs based on different evaluation metrics, further enhancing the overall practicality and robustness of the system.
[0095] In some examples, before the step of dividing the dataset into a training set and a test set, the method further includes: Filtering and denoising the arm joint electromyographic signal data to obtain clean electromyographic signal data; The arm joint resistance data and arm joint tremor signal data were standardized to eliminate data deviations between different patients.
[0096] For example, to further enhance the accuracy and stability of subsequent machine learning modeling, specialized preprocessing steps for different types of physiological signals were added before the dataset was divided into training and test sets. Specifically, arm joint electromyographic (EMG) data was first filtered and denoised to extract clean EMG signals. Bandpass filtering (e.g., 20Hz-450Hz) was typically used to remove low-frequency drift and high-frequency electrical noise, while notch filtering (50Hz or 60Hz power-frequency noise suppression) was combined to further eliminate environmental electromagnetic interference. After filtering, methods such as wavelet denoising or adaptive noise removal were applied to further smooth the EMG signal curve and remove non-physiological interference such as motion artifacts and contact noise. This series of steps enabled the obtained EMG signals to more accurately reflect the firing state of the muscles themselves, facilitating more accurate feature extraction (e.g., RMS value, dominant frequency characteristics), significantly improving the model's ability to identify muscle stiffness and spasticity. Furthermore, after acquisition, the arm joint resistance data and arm joint tremor signal data were standardized. The standardization step typically uses Z-score normalization, which involves subtracting the mean of each feature and then dividing it by the standard deviation. This ensures that the processed data has a standard normal distribution with a mean of 0 and a variance of 1 in all dimensions. Standardization is introduced primarily to address differences in data dimensions and numerical scales between patients due to individual differences (such as age, muscle mass, and joint flexibility). By unifying the scale, the model is protected from bias during training due to differences in the absolute size of input features, thereby improving the comparability and learning consistency of data from different patients in the feature space.
[0097] It is understandable that after filtering and denoising the electromyographic signal, the signal-to-noise ratio of the bioelectric signal is essentially improved, allowing the feature extraction process to focus on changes in muscle activity that are truly diagnostically valuable and avoid being misled by environmental noise or motion artifacts; while the standardization of resistance data and tremor signals improves the scale consistency and distribution stability of the input features, allowing the machine learning model to converge more quickly during training, and a more reasonable weight distribution relationship between different features in the modeling process, so that the learning effect of small-amplitude features will not be suppressed by large-amplitude features. As a result, for example, the reduction in RMS and spectral main frequency extraction deviations leads to improved electromyographic feature extraction accuracy; the input data distribution of the resistance curve and tremor curve is more standardized, and the convergence speed of the model training process is accelerated; in the subsequent training and testing stages, the performance fluctuation of the model on different patient samples has significantly decreased, and the evaluation stability has been enhanced.
[0098] In some examples, before the step of dividing the dataset into a training set and a test set, the method further includes: The characteristic information of the arm joint resistance data, the arm joint electromyographic signal data and the arm joint tremor signal data is extracted respectively using time domain and frequency domain analysis methods, wherein the characteristic information includes at least one of the mean, variance and frequency component.
[0099] It is understandable that in order to further improve the effectiveness of subsequent model training, a feature extraction process based on time domain and frequency domain analysis was added before the step of dividing the dataset into training and test sets. Specifically, for each type of physiological signal data, namely arm joint resistance data, arm joint electromyographic signal data, and arm joint tremor signal data, an adaptive analysis method was used to extract feature information that can represent the core change pattern of the signal. The feature information includes at least one or more of the time domain mean, time domain variance, and frequency component characteristics.
[0100] For example, for arm joint resistance data, time domain analysis is first performed to extract the mean (representing the average resistance level of the entire muscle to movement) and variance (reflecting the degree of resistance fluctuation and revealing the stability of muscle tension) during the resistance change process. Simultaneously, in the frequency domain, the frequency of torque signal changes is analyzed using fast Fourier transform (FFT) to identify high-frequency jitter or abnormal tremor components, further enriching the description of muscle rigidity and movement coordination. For arm joint electromyographic signal data, time domain feature extraction includes root mean square (RMS), mean potential value, and variance, which are used to characterize muscle discharge intensity, muscle baseline activity, and potential fluctuation, respectively. In the frequency domain, features such as peak frequency and spectral centroid are extracted to reflect the energy distribution characteristics of the electromyographic signal. In particular, an increase in high-frequency components is often closely related to muscle rigidity and abnormal discharge activity. For arm joint tremor signal data, the mean (measure of static deviation trend) and variance (measure of tremor amplitude fluctuation range) of the tremor acceleration signal are calculated in the time domain. In the frequency domain, the main frequency component (Tremor Frequency) of the tremor is analyzed through FFT to capture the typical tremor frequency band, such as 5Hz-8Hz physiological tremor or pathological tremor characteristics in higher frequency bands.
[0101] It can be understood that the time domain features (such as mean, variance, RMS, etc.) provide information about signal strength and stability, and are important indicators for describing the overall level changes of the signal; while the frequency domain features (such as dominant frequency, spectral centroid) reveal the internal oscillation, repetition mode and energy distribution law of the signal, which are particularly important for identifying abnormal oscillation caused by rigidity, discharge frequency change, tremor frequency anomaly and other pathological states. Joint extraction of time domain and frequency domain features can maximize the retention of diagnostic information in the original signal, and convert the original complex time series into a structured feature vector, providing a high-quality, low-dimensional, and strongly representative input basis for machine learning models. Thus, the information density of the model input features is significantly improved, enabling the model to learn the mapping relationship between physiological signals and muscle rigidity scores more quickly and accurately; the convergence speed during model training is improved; the mean absolute error (MAE) of the model prediction on the independent test set is further reduced; the system's ability to distinguish different symptom severity (mild, moderate, severe) is enhanced; and the most discriminative features can be identified in subsequent feature importance analysis, such as the dominant frequency of the electromyographic signal as a sensitive indicator of early muscle rigidity exacerbation.
[0102] In some examples, to achieve effective landing of the evaluation model in the clinic and convenience for doctors to use, after completing model inference and symptom scoring, the system further supports generating a muscle rigidity quantification evaluation report to display the results and analysis indicators output by the model in a graphical and structured form, including predicted score value, signal feature trend chart, key feature weight prompt, etc. The report can be automatically generated through a visual interface, allowing clinicians to quickly understand the evaluation results without deep involvement in computational details, and thus develop personalized intervention plans based on their own experience.
[0103] In some examples, to cooperate with data storage and subsequent tracking analysis, the software end is designed with a complete database system to save all collected original physiological signal data and model evaluation output results. The contents saved in the database include not only structured model score results, but also original signal sequences, signal feature vectors, patient basic information (such as name, age), evaluation dates, and key information fields such as the identifier of the attending physician, forming a complete record structure. The system supports establishing a follow-up record for each patient, and doctors can compare the quantitative scores, trend curves, and feature index changes at different time nodes to assist in judging treatment response or disease progression.
[0104] In some examples, to protect patient privacy and ensure data security, the database has login and permission control functions, allowing system administrators to configure viewing / modification permissions for different roles (such as doctors, nurses, and researchers), preventing unauthorized data access or changes. All historical data will be persistently stored, supporting instant retrieval for long-term longitudinal disease tracking and re-evaluation analysis, and providing valuable real-world data support for subsequent model updates or retraining.
[0105] In some examples, doctors can access the interface through a personal computer, tablet, or other compatible device as a terminal, log in to the system, call the database, view historical reports, initiate new assessments, download reports, or export structured data, etc. The entire process is designed to closely match the actual clinical application, taking into account both computational performance and deployment flexibility, while ensuring that doctors can efficiently obtain the information they need during the diagnosis and treatment process. The database deployment scheme is flexible and scalable, supporting both local server deployment within the hospital to meet strict data control requirements, and deployment on an encrypted and authenticated cloud platform to improve cross-department access efficiency or support remote medical system construction, meeting the differentiated needs of hospitals or clinics of different sizes.
[0106] Please refer to Figure 2 An embodiment of the dystonia symptom quantification evaluation device based on machine learning in the embodiments of the present application can include: A preprocessing unit 21 is configured to collect a preliminary data set of a patient, wherein the preliminary data set includes physiological signal feature data and label data, the feature data includes arm joint resistance data, arm joint electromyography signal, and arm joint tremor signal, and the label data is score data obtained by a doctor according to patient condition analysis; A model optimization unit 22 is configured to divide the data set into a training set and a test set, train an evaluation model through the training set, verify and optimize the trained evaluation model through the test set, and the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyography signal, and a random forest model constructed based on arm joint tremor signal; An evaluation unit 23 is configured to obtain physiological signal feature data of a target patient, and quantitatively evaluate the muscle stiffness symptom of the target patient based on the physiological signal feature data through the optimized evaluation model.
[0107] In summary, the embodiment of the present application provides a quantitative assessment device for myotonia symptoms based on machine learning, which collects a preliminary data set of patients, wherein the preliminary data set includes physiological signal feature data and label data, wherein the feature data includes arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals, and the label data is the scoring data obtained by the doctor based on the analysis of the patient's symptoms; the data set is divided into a training set and a test set, so as to train the evaluation model through the training set, and to verify and optimize the accuracy of the trained evaluation model through the test set, wherein the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyographic signals, and a random forest model constructed based on arm joint tremor signals; the physiological signal feature data of the target patient is obtained, and the target patient's myotonia symptoms are quantitatively assessed based on the physiological signal feature data through the optimized evaluation model. Automatic quantification through machine learning avoids subjective bias caused by doctor's experience, mood, and attention fluctuations. The high correlation between the physiological features learned in the training set and the clinical labels makes the evaluation more repeatable. Because it uses feature-level detection, even minor abnormalities in electromyographic discharges can be captured by the LightGBM model, indicating the risk of early myotonia. Quantitative changes in scores can be directly used to assess the effectiveness of rehabilitation or drug interventions.
[0108] like Figure 3 As shown, an embodiment of the present application further provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 320 and executable on the processor. When the processor 320 executes the computer program 311, the steps of any of the above-mentioned methods for quantitatively assessing myotonia symptoms based on machine learning are implemented: Collecting a preliminary data set of the patient, the preliminary data set includes physiological signal feature data and label data, the feature data includes arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals, and the label data is the scoring data obtained by the doctor based on the patient's condition analysis; The data set is divided into a training set and a test set, so as to train an evaluation model with the training set, and to verify and optimize the accuracy of the trained evaluation model with the test set, wherein the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyography signals, and a random forest model constructed based on arm joint tremor signals; Acquire physiological signal characteristic data of a target patient, and quantitatively evaluate the myotonia symptom of the target patient through an optimized evaluation model based on the physiological signal characteristic data.
[0109] Since the electronic device introduced in this embodiment is a device used to implement a machine learning-based quantitative assessment device for myotonia symptoms in the embodiment of the present application, based on the method introduced in the embodiment of the present application, those skilled in the art can understand the specific implementation of the electronic device of the present embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of the present application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in the embodiment of the present application falls within the scope of protection to be provided by this application.
[0110] In the specific implementation process, the computer program 311 can be implemented when executed by the processor Figure 1 Any implementation manner in the corresponding embodiment: Collecting a preliminary data set of the patient, the preliminary data set includes physiological signal feature data and label data, the feature data includes arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals, and the label data is the scoring data obtained by the doctor based on the patient's condition analysis; The data set is divided into a training set and a test set, so as to train an evaluation model with the training set, and to verify and optimize the accuracy of the trained evaluation model with the test set, wherein the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyography signals, and a random forest model constructed based on arm joint tremor signals; Acquire physiological signal characteristic data of a target patient, and quantitatively evaluate the myotonia symptom of the target patient through an optimized evaluation model based on the physiological signal characteristic data.
[0111] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0112] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0113] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0114] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0116] The present application also provides a computer program product, which includes computer software instructions. When the computer software instructions are executed on a processing device, the processing device is caused to execute the following Figure 1 The process of quantitative assessment of myotonia symptoms based on machine learning in the corresponding embodiment.
[0117] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, they fully or partially produce the processes or functions according to the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be stored by a computer, or a data storage device such as a server or data center that integrates one or more available media. Available media can include magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0118] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0119] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0120] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0121] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0122] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0123] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for quantitatively evaluating myotonia symptoms based on machine learning, characterized in that: include: Collecting a preliminary data set of the patient, the preliminary data set includes physiological signal feature data and label data, the feature data includes arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals, and the label data is the scoring data obtained by the doctor based on the patient's condition analysis; The data set is divided into a training set and a test set, so as to train an evaluation model with the training set, and to verify and optimize the accuracy of the trained evaluation model with the test set, wherein the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyography signals, and a random forest model constructed based on arm joint tremor signals; Acquire physiological signal characteristic data of a target patient, and quantitatively evaluate the myotonia symptom of the target patient through an optimized evaluation model based on the physiological signal characteristic data.
2. The method according to claim 1, wherein Before the step of dividing the data set into a training set and a test set, the method further includes: The preliminary data set is analyzed by principal component analysis, and the linear combination principal components with the largest amount of information are retained to obtain key feature data and reduce feature dimensions.
3. The method according to claim 1, wherein The evaluation model is trained by the training set, including The sample distribution is adjusted through the oversampling strategy, and three types of models are trained separately to extract the discriminant rules of the arm joint resistance data, arm joint electromyographic signals and arm joint tremor signal features.
4. The method according to claim 1, wherein Also includes: A weighted fusion mechanism is used to integrate the prediction results of the three models, where the weights are set according to the AUC values of the three models in the test set.
5. The method according to claim 1, wherein Also includes: After the evaluation model training is completed, the trained evaluation model is applied to independent test sets that have not been used in the training process and do not overlap with the training data. The mean absolute error and recall rate are used as performance evaluation indicators. The mean absolute error is used to measure the average deviation between the model prediction score and the doctor's annotation score, and the recall rate is used to measure the model's sensitivity to identifying patients with high-level symptoms.
6. The method according to claim 1, wherein Before the step of dividing the data set into a training set and a test set, the method further includes: Filtering and denoising the arm joint electromyographic signal data to obtain clean electromyographic signal data; The arm joint resistance data and arm joint tremor signal data were standardized to eliminate data deviations between different patients.
7. The method according to claim 1, wherein Before the step of dividing the data set into a training set and a test set, the method further includes: The characteristic information of the arm joint resistance data, the arm joint electromyographic signal data and the arm joint tremor signal data is extracted respectively using time domain and frequency domain analysis methods, wherein the characteristic information includes at least one of the mean, variance and frequency component.
8. A device for quantitatively evaluating myotonia symptoms based on machine learning, characterized in that: include: a pre-processing unit, configured to collect a preliminary data set of the patient, the preliminary data set including physiological signal feature data and label data, the feature data including arm joint resistance data, arm joint electromyographic signals, and arm joint tremor signals, and the label data being scoring data obtained by a doctor based on an analysis of the patient's symptoms; a model optimization unit, configured to divide the data set into a training set and a test set, so as to train an evaluation model using the training set, and to verify and optimize the accuracy of the trained evaluation model using the test set, wherein the evaluation model includes an XGBOOST model constructed based on arm joint resistance data, a LightGBM model constructed based on arm joint electromyography signals, and a random forest model constructed based on arm joint tremor signals; The evaluation unit is used to obtain physiological signal characteristic data of a target patient, and to quantitatively evaluate the myotonia symptom of the target patient through an optimized evaluation model based on the physiological signal characteristic data.
9. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the steps of the method for quantitatively assessing myotonia symptoms based on machine learning as described in any one of claims 1 to 7 when executing the computer program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for quantitatively evaluating myotonia symptoms based on machine learning according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Method and device for determining state of pet, electronic equipment and storage medium
CN121502385A