A stroke dysarthria speech feature selection method based on dynamic weight compensation

By employing dynamic weight compensation and temporal verification methods, the challenge of screening dysarthria features in stroke from high-dimensional redundant features was solved, achieving high-precision and robust feature extraction and improving the accuracy of stroke auxiliary diagnosis.

CN122163146APending Publication Date: 2026-06-09DALIAN NEUSOFT UNIV OF INFORMATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN NEUSOFT UNIV OF INFORMATION
Filing Date
2026-02-28
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing technologies struggle to stably and accurately screen out specific and robust speech features from high-dimensional mixed features, making it difficult to extract pathological features from stroke speech signals.

Method used

A dynamic weight compensation-based approach is adopted, which combines a dual-threshold outlier detection algorithm and a dynamic weight compensation strategy with a time-series dependency verification model to select a high-weight feature candidate set and perform anomaly detection and compensation. Finally, an adaptive truncation strategy is used to construct the final feature subset.

Benefits of technology

It significantly improves the discriminative power and interpretability of features, enhances the recognition accuracy and robustness of speech signals in stroke patients with dysarthria, and strengthens the accuracy and generalization ability of auxiliary diagnostic models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122163146A_ABST
    Figure CN122163146A_ABST
Patent Text Reader

Abstract

This invention discloses a method for selecting speech features in stroke-related dysarthria based on dynamic weight compensation. First, the patient's acoustic features are acquired and initial feature weights are calculated, dividing the data into high-weight and low-weight feature candidate sets. Second, outliers in high-weight features are identified using a dual-threshold outlier detection algorithm, and a dynamic weight compensation strategy is employed to correct weight biases while preserving potential pathological information, resulting in an optimized compensated feature dataset. Then, a training feature dataset is constructed using the low-weight candidate set, input into a pre-trained temporal dependency validation model for analysis, and weights are updated based on the feature's contribution to the classification results to obtain the final feature weights. Finally, features are classified according to these weights, and an adaptive truncation strategy is used to select the final feature subset. This invention constructs a two-layer selection mechanism from coarse to fine screening, effectively extracting the core acoustic features most relevant to pathology, improving feature discriminative power, interpretability, and model generalization ability, and providing high-quality feature input for assisted diagnosis and assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of stroke diagnosis technology, and in particular to a method for selecting speech features of stroke-related articulation disorders based on dynamic weight compensation. Background Technology

[0002] Stroke, also known as cerebrovascular accident or apoplexy, is a common clinical disease caused by the sudden rupture or blockage of blood vessels in the brain, leading to impaired blood circulation and subsequent brain tissue damage. It is characterized by high incidence, high disability rate, and high mortality rate. The clinical manifestations of stroke are diverse, including altered consciousness, slurred speech, and limb weakness, severely impacting patients' quality of life and social functioning.

[0003] Among them, the speech signals of stroke patients with dysarthria are a direct acoustic manifestation of damage to the neuromuscular control of the speech organs. Its pathological characteristics are complex and varied, often manifesting as unclear pronunciation, slow speech rate, and prosodic disorders. These abnormalities are specifically reflected in acoustic features such as increased fundamental frequency variability, reduced vowel space area, and decreased formant dynamic range. These pathological features are often submerged in a high-dimensional set of general acoustic features and intertwined with individual differences in healthy individuals, making it extremely difficult to directly and accurately extract highly discriminative pathological features from raw speech.

[0004] Therefore, there is an urgent need for a method that can automatically and accurately select a subset of features with strong stroke specificity and robustness from high-dimensional mixed features. Summary of the Invention

[0005] This invention provides a method for selecting speech features in stroke-induced dysarthria based on dynamic weight compensation, in order to overcome the above-mentioned technical problems.

[0006] To achieve the above objectives, the technical solution of the present invention is as follows: A method for selecting speech features in stroke-related articulation disorders based on dynamic weight compensation, comprising the following steps: S1. Obtain acoustic feature data of patients in the stroke patient dataset; S2. Calculate the initial feature weights of the acoustic feature data, and then obtain a high-weight feature candidate set and a low-weight feature candidate set based on the acoustic feature data and the initial feature weights; S3. Based on the designed dual-threshold outlier detection algorithm, perform weight anomaly detection on the high-weight feature candidate set, and compensate the anomaly detection results based on the designed dynamic weight compensation strategy to obtain the compensated feature dataset. Optimize the compensated feature dataset to obtain the optimized compensated feature dataset. Construct the training feature dataset based on the optimized compensated feature dataset and the low-weight feature candidate set. S4. Obtain the pre-trained temporal dependency verification model, input the training feature dataset into the pre-trained temporal dependency verification model to obtain the classification results of each optimized compensation feature, analyze the contribution of each acoustic feature in the training feature dataset to the classification results based on the classification results, and update the weight of each acoustic feature to obtain the final feature weight. S5. Classify the acoustic features in the training feature dataset according to the final feature weights, and use an adaptive truncation strategy to truncate the classified acoustic features to obtain the final feature subset.

[0007] Further, the specific steps for calculating the initial feature weights of the acoustic feature data, and then obtaining the high-weight feature candidate set and the low-weight feature candidate set based on the acoustic feature data and the initial feature weights, include: Suppose that the dataset of stroke patients contains a total of N There are samples, where the acoustic feature set is . F ={ f 1, f 2,… f m …, f M The tag set is C ={ c 1, c 2,…, c K}, then the first m Acoustic features f m Initial feature weights w m The calculation formula is: (1) in, and Features f m and categories c k The mean; M The total number of acoustic features; K Number of categories; This represents the m-th feature value of the i-th sample; This indicates that the i-th sample belongs to category [i]. c k The tag value; Acoustic features are assigned according to initial feature weights. w m The weights are sorted in descending order to obtain the initial acoustic feature sequence. S init =[ f (1),f (2),…, f ( M )]; Extract the first part of the initial acoustic feature sequence L 1 acoustic feature and use it as a high-weight feature candidate set S high Extracting the latter part of the initial acoustic feature sequence L 1 acoustic feature and use it as a low-weight feature candidate set S low .

[0008] Further, in S3, the high-weight feature candidate set is subjected to weight anomaly detection based on the designed dual-threshold outlier detection algorithm, and the anomaly detection results are compensated based on the designed dynamic weight compensation strategy to obtain a compensated feature dataset. The specific steps for optimizing the compensated feature dataset to obtain an optimized compensated feature dataset include: S31. Determine the candidate set of high-weight features. S high If each acoustic feature in the dataset conforms to a predefined global outlier removal strategy, the conforming acoustic feature is removed; otherwise, it is retained. The global outlier removal strategy is as follows: m w w (2) in, w The weighted average. w Standard deviation; w A global outlier threshold set based on the three-standard-deviation principle of feature distribution; S32. Calculate the high-weighted feature candidate set after processing in S31 using a sliding window. S high In the middle, the magnitude of the weight gradient change of adjacent acoustic features: Assuming the sliding window size is 5, the weights contained within the sliding window are... w t-2, w t-1, w t, w t+1, w t+2 Then the first t The magnitude of the weight gradient change of each acoustic feature g t The calculation formula is: g t= w t+1 -w t (3) Determine whether the magnitude of the weight gradient change of the currently calculated acoustic feature exceeds the set local outlier threshold; if so, mark the corresponding acoustic feature as a local anomalous feature. S33. The local abnormal features marked in S32 are compensated using a two-way compensation method to obtain the compensated feature dataset, which specifically includes: Determine whether the local abnormal feature is located in a set high-weight region or a low-weight region. If it is located in a high-weight region, then perform positive compensation on the weight of the local abnormal feature, that is: reduce its weight to the average weight of other acoustic features in the sliding window, and move its position in the feature sequence one position to the right; reduce the weight of the local abnormal feature to the average weight of other acoustic features in the sliding window except itself, and move the position of the local abnormal feature one position to the right in the direction of the low-weight region. If it is located in a low-weight region, the weight of the local abnormal feature is compensated in reverse, that is: its weight is increased to the average weight of other acoustic features in the sliding window, and its position in the feature sequence is moved forward one position. Let the compensated feature dataset be denoted as S comp , S comp The dimension is T × D , T For time steps, D For feature dimensions; S34. The energy threshold-based silence filtering method optimizes the compensated feature dataset to obtain an optimized compensated feature dataset, including: The short-time energy of each frame of speech corresponding to the acoustic features in the compensated feature dataset is calculated using the following formula: (4) in, x t ( n ) is the first t Frame audio signal; Determine if short-time energy meets the requirements ,max( E ) represents the set energy threshold. If it is true, the acoustic features corresponding to the speech signals that meet the conditions will be removed, and the optimized compensation feature dataset will be obtained.

[0009] Further, in S4, a temporal dependency verification model is established, and the training feature dataset is input into the temporal dependency verification model to obtain the classification results of each optimized compensation feature. Based on the classification results, the contribution of each acoustic feature in the training feature dataset to the classification results is analyzed, thereby updating the weight of each acoustic feature to obtain the final feature weights. The specific steps include: S41. A time-series dependency verification model is established based on a bidirectional LSTM network architecture to model the long-term dependency relationship of feature sequences. The time-series dependency verification model includes: an input layer, a bidirectional LSTM layer, a fully connected layer, and a Softmax activation function. The input layer is used to receive the training feature dataset; The bidirectional LSTM layer is used to capture and output the forward hidden state sequence and the backward hidden state sequence from the forward and backward directions; The fully connected layer is used to fuse the forward hidden state sequence and the backward hidden state sequence; The Softmax activation function is used to normalize the output of the fully connected layer and output the probability value corresponding to each category, which is the final classification decision result. S42. Calculate the feature importance score of each acoustic feature in the training feature dataset through gradient backpropagation. s m The calculation formula is: (5) in, The cross-entropy loss function; This represents the m-th acoustic feature value at the t-th time step; The final feature weights are obtained by updating the weights of each acoustic feature based on its importance score. The calculation formula is as follows: (6) in, l This is the fusion coefficient.

[0010] Further, in S5, the specific steps of classifying the acoustic features in the training feature dataset according to the final feature weights and truncating the classified acoustic features using an adaptive truncation strategy to obtain the final feature subset include: S51, Set threshold i 1, i 2, i The search range and search step size of 3 are determined by using a grid search method to traverse all possible threshold combinations within a set range. The acoustic features in the training feature dataset are temporarily classified by combining the threshold combinations of each traversal and the final feature weights, thereby selecting features that meet the conditions to form several temporary candidate feature subsets. That is, based on each threshold combination and the final feature weight w m The acoustic features in the training feature dataset are divided into four categories, including: Will w m ′≥ i The acoustic characteristics of 1 are classified as the first type of acoustic characteristics; Will i 2≤ w m ′< i The acoustic characteristics of 1 are used as the second type of acoustic characteristics; Will i 3≤ w m ′< i The acoustic characteristics of 2 are classified as the third type of acoustic characteristics; Will w m ′< i The acoustic features of type 3 are classified as the fourth type of acoustic features; A temporary subset of candidate features is constructed based on the first type of acoustic features and the second type of acoustic features; S52. Train the temporal dependency verification model based on several temporary candidate feature subsets. After training, calculate the cross-entropy loss of the temporal dependency verification model based on the verification set for all threshold combinations. S53. Compare the cross-entropy loss results under all threshold combinations, and take the threshold combination that minimizes the cross-entropy loss result as the optimal threshold combination. S54. Based on the optimal threshold combination, perform the final classification of the acoustic features in the training feature dataset; S55. The adaptive truncation method is used to truncate the first and second class acoustic features obtained from the final classification. The specific steps include: Set an initial cutoff point, and use the initial cutoff point to initially extract the first part of the first type of acoustic features and the second type of acoustic features. L One feature; Determine whether the weighted accuracy of the trained temporal dependency validation model on the validation set is lower than a set threshold. or If so, then gradually increase the amount of material initially captured. L The number of features up to L +Δ, until the weighted accuracy reaches the target, thus obtaining the final feature subset, represented as: S final ={ f (1), f (2),…, f ( L +Δ)}.

[0011] Beneficial Effects: This invention constructs high-weight and low-weight feature candidate sets based on initial feature weights, and classifies acoustic features in the training feature dataset according to the final feature weights. An adaptive truncation strategy is used to truncate the classified acoustic features, constructing a two-layer feature selection mechanism from coarse to fine screening. This effectively extracts the core acoustic manifestations most relevant to the pathophysiological process of dysarthria from massive and mixed general acoustic features, significantly improving the discriminative power and interpretability of the features. The designed dual-threshold outlier detection algorithm directly identifies and processes outliers in high-weight features. A dynamic weight compensation strategy corrects the weight bias caused by outliers while retaining the potential pathological information contained in the feature, thereby improving the generalization ability of the time-dependent validation model. This invention effectively overcomes the challenge of stably extracting discriminative pathological features from complex, high-dimensional, and highly individualized speech signals of stroke-related dysarthria, providing high-quality feature input for constructing more accurate and robust auxiliary diagnostic and assessment models. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 and Figure 2 This is a flowchart of a speech feature selection method for dysarthria in stroke based on dynamic weight compensation, as described in this invention. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0015] This embodiment provides a method for selecting speech features in stroke-related dysarthria based on dynamic weight compensation, such as... Figure 1 and Figure 2 As shown, the specific steps include: S1. Obtain acoustic feature data of patients in the stroke patient dataset; S2. Calculate the initial feature weights of the acoustic feature data, and then obtain a high-weight feature candidate set and a low-weight feature candidate set based on the acoustic feature data and the initial feature weights; In a specific embodiment, S2, calculating the initial feature weights of the acoustic feature data, and then obtaining a high-weight feature candidate set and a low-weight feature candidate set based on the acoustic feature data and the initial feature weights, includes the following specific steps: Suppose that the dataset of stroke patients contains a total of N There are samples, where the acoustic feature set is . F ={ f 1, f 2,… f m …, f M The tag set is C ={ c 1, c 2,…, c K}, then the first m Acoustic features f m Initial feature weights w m The calculation formula is: (1) in, and Features f m and categories c k The mean; M The total number of acoustic features; K Number of categories; This represents the m-th feature value of the i-th sample; This indicates that the i-th sample belongs to category [i]. c k The tag value; Specifically, formula (1) can effectively reflect the ability of acoustic features to differentiate patients with different degrees of severity by averaging the correlation coefficients of multiple categories. Furthermore, jitter may be highly correlated with the label in severe patients, but less correlated in mild patients, and its overall importance can be summarized by weighted averaging.

[0016] Acoustic features are assigned according to initial feature weights. w m The weights are sorted in descending order to obtain the initial acoustic feature sequence. S init =[ f (1), f (2),…,f ( M )]; In this embodiment, to reduce computational complexity, the first part of the initial acoustic feature sequence is extracted. L 1 acoustic feature and use it as a high-weight feature candidate set S high Extracting the latter part of the initial acoustic feature sequence L 1 acoustic feature and use it as a low-weight feature candidate set S low Preferably, L = M / 3.

[0017] Specifically, in this embodiment, the intermediate data other than the high-weight feature candidate set and the low-weight feature candidate set are not involved in the subsequent processing.

[0018] S3. Based on the designed dual-threshold outlier detection algorithm, perform weight anomaly detection on the high-weight feature candidate set, and compensate the anomaly detection results based on the designed dynamic weight compensation strategy to obtain the compensated feature dataset. Optimize the compensated feature dataset to obtain the optimized compensated feature dataset. Construct the training feature dataset based on the optimized compensated feature dataset and the low-weight feature candidate set. In a specific embodiment, in S3, the high-weight feature candidate set is subjected to weight anomaly detection based on the designed dual-threshold outlier detection algorithm, and the anomaly detection result is compensated based on the designed dynamic weight compensation strategy to obtain a compensated feature dataset. The specific steps for optimizing the compensated feature dataset to obtain an optimized compensated feature dataset include: Specifically, due to the high-weight feature candidate set S high The sample may contain anomalous features due to measurement errors or sample bias, such as extreme resonance peaks or fundamental frequency jitter. Therefore, this embodiment designs a dual-threshold outlier detection algorithm for high-weight feature candidate sets. S high Processing: S31. Determine the candidate set of high-weight features. S high If each acoustic feature in the dataset conforms to a predefined global outlier removal strategy, the conforming acoustic feature is removed; otherwise, it is retained. The global outlier removal strategy is as follows: m w w (2) in, wThe weighted average. w Standard deviation; w A global outlier threshold set based on the three-standard-deviation principle of feature distribution; S32. Calculate the high-weighted feature candidate set after processing in S31 using a sliding window. S high In the equation, the magnitude of the weighted gradient change of adjacent acoustic features is defined, where the sliding window size is set. W =5; Assuming the sliding window size is 5, the weights contained within the sliding window are... w t-2, w t-1, w t, w t+1, w t+2 Then the first t The magnitude of the weight gradient change of each acoustic feature g t The calculation formula is: g t= w t+1 -w t (3) Determine whether the magnitude of the weight gradient change of the currently calculated acoustic feature exceeds the set local outlier threshold. If so, mark the corresponding acoustic feature as a local anomalous feature.

[0019] In this embodiment, the local outlier threshold is set to twice the average gradient within the sliding window; Specifically, if the weight of a certain parameter suddenly increases within a local window, it may be due to data acquisition noise. In this case, the time-dependent verification model needs to be used to perform a final discrimination evaluation on all features.

[0020] S33. The local abnormal features marked in S32 are compensated using a two-way compensation method to obtain the compensated feature dataset, which specifically includes: Determine whether the local abnormal feature is located in the set high weight area (first 20%) or low weight area (last 20%). If it is located in the high weight area, then the weight of the local abnormal feature is positively compensated, that is: the weight of the local abnormal feature is reduced to the average weight of the other acoustic features in the sliding window except itself, and the position of the local abnormal feature in the feature sequence is shifted one position to the right in the direction of the low weight area. Specifically, the fundamental frequency feature in acoustic features has a small sample size, which may lead to its overestimation. Compensation can reduce its ranking.

[0021] If it is located in a low-weight region, the weight of the local abnormal feature is compensated in reverse, that is: the weight of the local abnormal feature is increased to the average weight of the other acoustic features in the sliding window except itself, and the position of the local abnormal feature in the feature sequence is moved forward one position in order to be closer to the high-weight region. Specifically, the vowel space area (VSA) in acoustic features may be underestimated due to skewed data distribution, and its importance can be enhanced through compensation.

[0022] Let the compensated feature dataset be denoted as S comp , S comp The dimension is T × D , T For time steps, D As a feature dimension, this embodiment uses a two-way compensation method to process local abnormal features, which can ensure that high-discrimination features are concentrated in the early part of the sequence.

[0023] S34. Specifically, since stroke speech often contains non-steady-state silence segments, which may interfere with the effectiveness of features, this embodiment optimizes the compensated feature dataset using an energy threshold-based silence filtering method to obtain an optimized compensated feature dataset, including: The short-time energy of each frame of speech corresponding to the acoustic features in the compensated feature dataset is calculated using the following formula: (4) in, x t ( n ) is the first t Frame audio signal; Determine if short-time energy meets the requirements ,max( E () indicates the set energy threshold. If the value is 0.1, then the acoustic features corresponding to the speech signals that meet the conditions will be removed, and the optimized compensation feature dataset will be obtained.

[0024] Specifically, the silence filtering method can reduce the interference of invalid features on the training of subsequent temporal dependency verification models, thereby improving the purity of the feature set.

[0025] S4. Obtain the pre-trained temporal dependency verification model, input the training feature dataset into the pre-trained temporal dependency verification model to obtain the classification results of each optimized compensation feature, analyze the contribution of each acoustic feature in the training feature dataset to the classification results based on the classification results, and update the weight of each acoustic feature to obtain the final feature weight. Specifically, in this embodiment, the training feature dataset is divided into a training set and a validation set.

[0026] In a specific embodiment, in S4, a temporal dependency verification model is established, and the training feature dataset is input into the temporal dependency verification model to obtain the classification results of each optimized compensation feature. Based on the classification results, the contribution of each acoustic feature in the training feature dataset to the classification results is analyzed, thereby updating the weight of each acoustic feature to obtain the final feature weights. The specific steps include: S41. Specifically, since the acoustic features of stroke patients have significant temporal dependence, it is necessary to verify the temporal discriminative power of the features through a deep learning model. In this embodiment, a temporal dependence verification model is established based on a bidirectional LSTM network architecture to model the long-term dependence of feature sequences. The temporal dependence verification model includes: an input layer, a bidirectional LSTM layer, a fully connected layer, and a Softmax activation function. The input layer is used to receive the training feature dataset; The bidirectional LSTM layer is used to capture and output the forward hidden state sequence and the backward hidden state sequence from the forward and backward directions; Specifically, in this embodiment, each LSTM layer contains 128 hidden units; The fully connected layer is used to fuse the forward hidden state sequence and the backward hidden state sequence; The Softmax activation function is used to normalize the output of the fully connected layer and output the probability value corresponding to each category, which is the final classification decision result.

[0027] Specifically, many abnormalities in articulation disorders are inherently dynamic and evolve over time. This embodiment utilizes a temporal dependency verification model to capture the joint change patterns and correlations of features over time, thereby enabling the assessment of which static or dynamic features truly play a key discriminative role in the continuous production of pathological speech. This makes the selected feature subset more capable of characterizing the dynamic pathological features of articulation disorders.

[0028] S42. Calculate the feature importance score of each acoustic feature in the training feature dataset through gradient backpropagation. s m The calculation formula is: (5) in, The cross-entropy loss function; This represents the m-th acoustic feature value at the t-th time step; Specifically, feature importance score s m The contribution of each acoustic feature in the training feature dataset to the temporal classification task was quantified; The final feature weights are obtained by updating the weights of each acoustic feature based on its importance score. The calculation formula is as follows: (6) in, l This is the fusion coefficient, with a default value of 0.5, used to balance statistical correlation with the importance of model-driven features.

[0029] S5. Classify the acoustic features in the training feature dataset according to the final feature weights, and use an adaptive truncation strategy to truncate the classified acoustic features to obtain the final feature subset.

[0030] In a specific embodiment, S5, the steps of classifying the acoustic features in the training feature dataset according to the final feature weights and truncating the classified acoustic features using an adaptive truncation strategy to obtain the final feature subset include: S51, Set threshold i 1, i 2, i The search range and search step size of 3 are determined by using a grid search method to traverse all possible threshold combinations within a set range. The acoustic features in the training feature dataset are temporarily classified by combining the threshold combinations of each traversal and the final feature weights, thereby selecting features that meet the conditions to form several temporary candidate feature subsets. Specifically, this embodiment sets a threshold. i 1, i 2, i The search ranges for 3 are [0.6, 1.0], [0.4, 0.8], and [0.2, 0.6], respectively, and all possible threshold combinations are traversed with a set step size; Specifically, based on each threshold combination and the final feature weight w m The acoustic features in the training feature dataset are divided into four categories, including: Will w m ′≥ i The acoustic features of 1 are designated as the first type of acoustic features, and this embodiment retains the first type of acoustic features; Specifically, the first type of acoustic feature is a high-discrimination feature, including fundamental frequency variability (Jitter) and formant concentration ratio (FCR). Will i 2≤ w m ′< i The acoustic characteristics of 1 are used as the second type of acoustic characteristics; Specifically, the second category of acoustic features is of medium discriminative power. During the actual model development, validation, or threshold optimization stages, prior clinical knowledge needs to be introduced to interpret, validate, or fine-tune the feature results. For example, the area of ​​vowel space (VSA) varies due to individual differences and requires further validation, i.e., calculating the correlation between this feature and the articulation disorder assessment score. When the correlation coefficient passes the statistical significance test, the feature is retained and included in the final candidate feature subset; otherwise, it will be classified as a low discriminative feature.

[0031] Will i 3≤ w m ′< i The acoustic characteristics of 2 are classified as the third type of acoustic characteristics; Specifically, the third type of acoustic feature is a low-discrimination feature, which is only used as a supplement in unbalanced data in practice; Will w m ′< i The acoustic features of type 3 are classified as the fourth type of acoustic features; Specifically, the fourth type of acoustic feature is a redundant or noise feature, which can be directly eliminated in practice.

[0032] This embodiment constructs a temporary subset of candidate features based on the first type of acoustic features and the second type of acoustic features; S52. Train the temporal dependency verification model based on several temporary candidate feature subsets. After training, calculate the cross-entropy loss of the temporal dependency verification model based on the verification set for all threshold combinations. S53. Compare the cross-entropy loss results under all threshold combinations, and take the threshold combination that minimizes the cross-entropy loss result as the optimal threshold combination. Specifically, this embodiment selects the threshold combination that minimizes the validation set loss as the optimal parameter. If multiple combinations have similar losses, the threshold combination with fewer features is preferred. This embodiment determines the threshold value based on the optimization results. i 1 = 0.8 i 2 = 0.6 i 3 = 0.4.

[0033] S54. Based on the optimal threshold combination, perform the final classification of the acoustic features in the training feature dataset; S55. To balance model efficiency and accuracy, this embodiment employs an adaptive truncation method to truncate the first and second class acoustic features obtained from the final classification. Specific steps include: Set an initial cutoff point, and use the initial cutoff point to initially extract the first part of the first type of acoustic features and the second type of acoustic features. L One characteristic, L = M / 3; Determine whether the weighted accuracy (WA) of the trained temporal dependency validation model on the validation set is lower than a set threshold. or , specifically, or =0.75, if so, then gradually increase the value of the initial cut-off. L The number of features up to L +Δ, Δ=10, until WA reaches the target, thus obtaining the final feature subset, represented as: S final ={ f (1), f (2),…, f ( L +Δ)}; Specifically, in this embodiment, the number of features is increased sequentially according to the order of the first type of acoustic features and the second type of acoustic features.

[0034] Specifically, WA is a commonly used classification performance evaluation metric when dealing with imbalanced data. It assigns different weights to each class based on the number of samples in each class, and its calculation formula is as follows: (7) in, N The total number of samples, K The total number of categories, Let the weight be the weight of the k-th class. Let be the number of true instances in class k.

[0035] The dynamic weight compensation algorithm proposed in this invention is a progressive, bidirectional feedback method. Its inherent connections are as follows: (1) Initial value calculation of feature weights provides a benchmark and focus range The initial feature weight calculation is the guiding principle of the entire algorithm. It provides an initial importance score for all features based on statistical priors through a multi-class correlation coefficient weighting method. (Initial acoustic feature sequence) S init It is the starting point for all operations, which involves extracting candidate sets of high and low weight features. S high and S lowThis focuses the optimization scope for the next stage of dynamic anomaly feature compensation. At this point, instead of inefficiently traversing all features, computational resources are concentrated on the subset of features most likely to have problems, improving algorithm efficiency. (2) Dynamic weight compensation strategy provides enhanced feature sequences The dynamic weight compensation strategy provides weight correction functionality. It receives a candidate set from the initial step and uses a dual-threshold mechanism to identify and correct unreliable weights caused by data noise, measurement errors, or sample bias. The output is the compensated sequence. S comp This is a more reliable feature ranking. This sequence is fed as high-quality input into the temporal dependency validation model, ensuring the reliability of the validation results; (3) Temporal dependency verification introduces dynamic changes This is the key step in the algorithm's integration of statistical static analysis and dynamic temporal modeling. The bidirectional LSTM network learns the temporal discriminative power of features from the training feature dataset and calculates the importance score through gradient backpropagation. s m It provides the contribution of features in the time dimension that statistical methods cannot capture. This is achieved by using statistical weights w m With time series score s m The features are then fused to generate the final feature weights. w' m This operation enhances the dynamic anomaly feature compensation from the previous stage, generating... w' m This serves as the basis for subsequent feature subset optimization steps; (4) Feature subset optimization and truncation strategy Based on the final feature weights w' m Weights are used to qualitatively partition features, and an adaptive truncation strategy dynamically determines the size of the optimal feature subset based on the real-time performance of the validation model on the validation set, considering temporal dependencies. The final output is a highly discriminative, dimensionally optimal feature subset. S final This completes the mapping from a high-dimensional redundant feature space to a stroke-specific feature space for dysarthria.

[0036] In summary, the present invention has the following beneficial effects: (1) Significantly improve recognition accuracy Through dynamic weight compensation and temporal verification, the most pathologically discriminative feature subset can be accurately selected from high-dimensional redundant features. On the MSDM database, the weighted accuracy of the four-class classification task for stroke dysarthria was improved from a baseline of 0.51 to 0.67, a relative improvement of over 30%. (2) Robustness of enhanced features The dual-threshold detection and bidirectional compensation mechanism effectively suppressed the abnormal feature weights caused by data noise or individual differences, and improved the stability of the feature set; the key features selected (such as Jitter, VSA, FCR) were significantly correlated with the Clinical Assessment Scale (FDA); (3) Optimize computational efficiency Compared to traditional wrapper-style feature selection methods (such as RFE), this invention reduces training time by about 35% while maintaining accuracy through dynamic compensation and adaptive truncation. The entire feature optimization process does not require much manual intervention, achieving end-to-end optimization from the original features to the optimal subset.

[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for selecting speech features in stroke-induced articulation disorders based on dynamic weight compensation, characterized in that, The specific steps include: S1. Obtain acoustic feature data of patients in the stroke patient dataset; S2. Calculate the initial feature weights of the acoustic feature data, and then obtain a high-weight feature candidate set and a low-weight feature candidate set based on the acoustic feature data and the initial feature weights; S3. Based on the designed dual-threshold outlier detection algorithm, perform weight anomaly detection on the high-weight feature candidate set, and compensate the anomaly detection results based on the designed dynamic weight compensation strategy to obtain the compensated feature dataset. Optimize the compensated feature dataset to obtain the optimized compensated feature dataset. Construct the training feature dataset based on the optimized compensated feature dataset and the low-weight feature candidate set. S4. Obtain the pre-trained temporal dependency verification model, input the training feature dataset into the pre-trained temporal dependency verification model to obtain the classification results of each optimized compensation feature, analyze the contribution of each acoustic feature in the training feature dataset to the classification results based on the classification results, and update the weight of each acoustic feature to obtain the final feature weight. S5. Classify the acoustic features in the training feature dataset according to the final feature weights, and use an adaptive truncation strategy to truncate the classified acoustic features to obtain the final feature subset.

2. The method for selecting speech features of dysarthria in stroke patients based on dynamic weight compensation according to claim 1, characterized in that, The specific steps for calculating the initial feature weights of the acoustic feature data, and then obtaining the high-weight feature candidate set and the low-weight feature candidate set based on the acoustic feature data and the initial feature weights, include: Suppose that the dataset of stroke patients contains a total of N There are samples, where the acoustic feature set is . F ={ f 1, f 2,… f m …, f M The tag set is C ={ c 1, c 2,…, c K }, then the first m Acoustic features f m Initial feature weights w m The calculation formula is: (1) in, and Features f m and categories c k The mean; M The total number of acoustic features; K Number of categories; This represents the m-th feature value of the i-th sample; This indicates that the i-th sample belongs to category [i]. c k The tag value; Acoustic features are assigned according to initial feature weights. w m The weights are sorted in descending order to obtain the initial acoustic feature sequence. S init =[ f (1), f (2),…, f ( M )]; Extract the first part of the initial acoustic feature sequence L 1 acoustic feature and use it as a high-weight feature candidate set S high Extracting the latter part of the initial acoustic feature sequence L 1 acoustic feature and use it as a low-weight feature candidate set S low .

3. The method for selecting speech features of dysarthria in stroke patients based on dynamic weight compensation according to claim 2, characterized in that, In S3, the high-weight feature candidate set is subjected to weight anomaly detection based on the designed dual-threshold outlier detection algorithm, and the anomaly detection results are compensated based on the designed dynamic weight compensation strategy to obtain the compensated feature dataset. The specific steps for optimizing the compensated feature dataset to obtain the optimized compensated feature dataset include: S31. Determine the candidate set of high-weight features. S high If each acoustic feature in the dataset conforms to a predefined global outlier removal strategy, the conforming acoustic feature is removed; otherwise, it is retained. The global outlier removal strategy is as follows: m w w (2) in, w The weighted average. w Standard deviation; w A global outlier threshold set based on the three-standard-deviation principle of feature distribution; S32. Calculate the high-weighted feature candidate set after processing in S31 using a sliding window. S high In the middle, the magnitude of the weight gradient change of adjacent acoustic features: Assuming the sliding window size is 5, the weights contained within the sliding window are... w t-2, w t-1, w t, w t+1, w t+2 Then the first t The magnitude of the weight gradient change of each acoustic feature g t The calculation formula is: g t= w t+1 -w t (3) Determine whether the magnitude of the weight gradient change of the currently calculated acoustic feature exceeds the set local outlier threshold; if so, mark the corresponding acoustic feature as a local anomalous feature. S33. The local abnormal features marked in S32 are compensated using a two-way compensation method to obtain the compensated feature dataset, which specifically includes: Determine whether the local abnormal feature is located in a set high-weight region or a low-weight region. If it is located in a high-weight region, then perform positive compensation on the weight of the local abnormal feature, that is: reduce its weight to the average weight of other acoustic features in the sliding window, and move its position in the feature sequence one position to the right; reduce the weight of the local abnormal feature to the average weight of other acoustic features in the sliding window except itself, and move the position of the local abnormal feature one position to the right in the direction of the low-weight region. If it is located in a low-weight region, the weight of the local abnormal feature is compensated in reverse, that is: its weight is increased to the average weight of other acoustic features in the sliding window, and its position in the feature sequence is moved forward one position. Let the compensated feature dataset be denoted as S comp , S comp The dimension is T × D , T For time steps, D For feature dimensions; S34. The energy threshold-based silence filtering method optimizes the compensated feature dataset to obtain an optimized compensated feature dataset, including: The short-time energy of each frame of speech corresponding to the acoustic features in the compensated feature dataset is calculated using the following formula: (4) in, x t ( n ) is the first t Frame audio signal; Determine if short-time energy meets the requirements ,max( E ) represents the set energy threshold. If it is true, the acoustic features corresponding to the speech signals that meet the conditions will be removed, and the optimized compensation feature dataset will be obtained.

4. The method for selecting speech features of dysarthria in stroke based on dynamic weight compensation according to claim 3, characterized in that, In S4, a temporal dependency verification model is established. The training feature dataset is input into the temporal dependency verification model to obtain the classification results of each optimized compensation feature. Based on the classification results, the contribution of each acoustic feature in the training feature dataset to the classification results is analyzed, thereby updating the weight of each acoustic feature to obtain the final feature weights. The specific steps include: S41. A time-series dependency verification model is established based on a bidirectional LSTM network architecture to model the long-term dependency relationship of feature sequences. The time-series dependency verification model includes: an input layer, a bidirectional LSTM layer, a fully connected layer, and a Softmax activation function. The input layer is used to receive the training feature dataset; The bidirectional LSTM layer is used to capture and output the forward hidden state sequence and the backward hidden state sequence from the forward and backward directions; The fully connected layer is used to fuse the forward hidden state sequence and the backward hidden state sequence; The Softmax activation function is used to normalize the output of the fully connected layer and output the probability value corresponding to each category, which is the final classification decision result. S42. Calculate the feature importance score of each acoustic feature in the training feature dataset through gradient backpropagation. s m The calculation formula is: (5) in, The cross-entropy loss function; This represents the m-th acoustic feature value at the t-th time step; The final feature weights are obtained by updating the weights of each acoustic feature based on its importance score. The calculation formula is as follows: (6) in, λ This is the fusion coefficient.

5. The method for selecting speech features of dysarthria in stroke based on dynamic weight compensation according to claim 4, characterized in that, In S5, the specific steps of classifying the acoustic features in the training feature dataset according to the final feature weights and truncating the classified acoustic features using an adaptive truncation strategy to obtain the final feature subset include: S51, Set threshold θ 1, θ 2, θ The search range and search step size of 3 are determined by using a grid search method to traverse all possible threshold combinations within a set range. The acoustic features in the training feature dataset are temporarily classified by combining the threshold combinations of each traversal and the final feature weights, thereby selecting features that meet the conditions to form several temporary candidate feature subsets. That is, based on each threshold combination and the final feature weight w m The acoustic features in the training feature dataset are divided into four categories, including: Will w m ′≥ θ The acoustic characteristics of 1 are classified as the first type of acoustic characteristics; Will θ 2≤ w m ′< θ The acoustic characteristics of 1 are used as the second type of acoustic characteristics; Will θ 3≤ w m ′< θ The acoustic characteristics of 2 are classified as the third type of acoustic characteristics; Will w m ′< θ The acoustic features of type 3 are classified as the fourth type of acoustic features; A temporary subset of candidate features is constructed based on the first type of acoustic features and the second type of acoustic features; S52. Train the temporal dependency verification model based on several temporary candidate feature subsets. After training, calculate the cross-entropy loss of the temporal dependency verification model based on the verification set for all threshold combinations. S53. Compare the cross-entropy loss results under all threshold combinations, and take the threshold combination that minimizes the cross-entropy loss result as the optimal threshold combination. S54. Based on the optimal threshold combination, perform the final classification of the acoustic features in the training feature dataset; S55. The adaptive truncation method is used to truncate the first and second class acoustic features obtained from the final classification. The specific steps include: Set an initial cutoff point, and use the initial cutoff point to initially extract the first part of the first type of acoustic features and the second type of acoustic features. L One feature; Determine whether the weighted accuracy of the trained temporal dependency validation model on the validation set is lower than a set threshold. η If so, then gradually increase the amount of material initially captured. L The number of features up to L +Δ, until the weighted accuracy reaches the target, thus obtaining the final feature subset, represented as: S final ={ f (1), f (2),…, f ( L +D)}。