Acoustic feature processing method and apparatus, computer device, readable storage medium, and program product

By processing the acoustic features of speech data from healthy individuals and Huntington's disease patients, selecting and weighting target acoustic features, and constructing speech index features, the problem of noise influence in tracking the progression of Huntington's disease was solved, and a more robust disease assessment was achieved.

CN122135744APending Publication Date: 2026-06-02YIHANG (GUANGZHOU) CLINIC CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YIHANG (GUANGZHOU) CLINIC CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

When using existing technologies to track the progression of Huntington's disease, isolated acoustic features are susceptible to environmental noise, reducing robustness.

Method used

Acoustic features were extracted from speech data of healthy individuals and patients with Huntington's disease. Stability and significant independent contribution were assessed to select target acoustic features. Based on weights, these features were weighted and scaled to construct a speech index feature to characterize the severity of Huntington's disease.

Benefits of technology

Integrating the multidimensional language impairments of Huntington's disease into clinically interpretable speech index features reduces the impact of environmental noise and improves the robustness of tracking Huntington's disease progression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135744A_ABST
    Figure CN122135744A_ABST
Patent Text Reader

Abstract

This application relates to the field of speech signal processing technology, providing an acoustic feature processing method, apparatus, computer device, readable storage medium, and program product. The method includes: extracting acoustic features from speech data of healthy individuals and Huntington's disease patients to obtain acoustic feature sets; evaluating the stability of the acoustic features in the acoustic feature sets to obtain candidate acoustic feature sets; evaluating the significant independent contribution of the candidate acoustic features in the candidate acoustic feature sets to obtain multiple target acoustic features and their corresponding weights; performing weighted and scaled processing on the multiple target acoustic features based on their weights to obtain speech index features; and using the speech index features to characterize the severity of Huntington's disease. This method can integrate the multidimensional language impairments of Huntington's disease into clinically interpretable speech index features that are significantly correlated with the severity of Huntington's disease, improving the robustness of tracking Huntington's disease progression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech signal processing technology, and in particular to an acoustic feature processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology

[0002] Research on speech function in Huntington's disease has primarily focused on isolated acoustic features, such as vowel duration prolongation, reduced basic frequency variability, and increased articulation errors. However, isolated acoustic features only cover a narrow aspect of Huntington's disease-related speech dysfunction, are easily affected by environmental noise, and reduce the robustness of tracking Huntington's disease progression. Summary of the Invention

[0003] Therefore, it is necessary to provide an acoustic feature processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product to address the aforementioned technical problems.

[0004] Firstly, this application provides an acoustic feature processing method, including:

[0005] Acoustic features were extracted from the speech data of healthy individuals and Huntington's disease patients to obtain acoustic feature sets.

[0006] The stability of the acoustic features in the acoustic feature set is evaluated to obtain a candidate acoustic feature set;

[0007] The candidate acoustic features in the candidate acoustic feature set are evaluated for significant independent contribution to obtain multiple target acoustic features and their corresponding weights.

[0008] The multiple target acoustic features are weighted and scaled based on the weights corresponding to the target acoustic features to obtain speech index features; the speech index features are used to characterize the severity of Huntington's disease.

[0009] Secondly, this application also provides an acoustic feature processing device, comprising:

[0010] The acoustic feature set acquisition module is used to extract acoustic features from the speech data of healthy people and Huntington's disease patients respectively, and obtain acoustic feature sets.

[0011] The candidate acoustic feature set acquisition module is used to evaluate the stability of the acoustic features in the acoustic feature set to obtain the candidate acoustic feature set.

[0012] The target acoustic feature and weight acquisition module is used to evaluate the significant independent contribution of the candidate acoustic features in the candidate acoustic feature set to obtain multiple target acoustic features and the weights corresponding to the target acoustic features.

[0013] The speech index feature acquisition module is used to perform weighted and scaled processing on the multiple target acoustic features based on the weights corresponding to the target acoustic features to obtain speech index features; the speech index features are used to characterize the severity of Huntington's disease.

[0014] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the above-described method.

[0015] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, the computer program being executed by a processor using the methods described above.

[0016] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that is executed by a processor using the methods described above.

[0017] The aforementioned acoustic feature processing methods, devices, computer equipment, computer-readable storage media, and computer program products, based on the stability and significant independent contribution of acoustic features, select target acoustic features from the acoustic feature set and determine the weights corresponding to the target acoustic features. Based on the weights corresponding to the target acoustic features, multiple target acoustic features are weighted and scaled to obtain speech index features characterizing the severity of Huntington's disease. Integrating the multidimensional language impairments of Huntington's disease into clinically interpretable speech index features significantly correlated with the severity of Huntington's disease effectively averages the noise and variability associated with individual acoustic features, reduces the impact of environmental noise, and improves the robustness of tracking Huntington's disease progression. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a diagram illustrating the application environment of an acoustic feature processing method in one embodiment.

[0020] Figure 2 This is a flowchart illustrating an acoustic feature processing method in one embodiment;

[0021] Figure 3 This is a schematic diagram of the correlation study design and experimental workflow in one embodiment;

[0022] Figure 4a Receiver operating characteristic curves for the performance of speech index features in categorical healthy controls and Huntington gene carriers in one embodiment;

[0023] Figure 4b This is a forest plot of confidence intervals for the coefficients of the second multivariate logistic regression model (used to construct speech index features) in one embodiment;

[0024] Figure 4c A radar image showing the acoustic features of key targets in four participating groups in one embodiment;

[0025] Figure 4d A heatmap showing partial correlations between acoustic features of four key targets and clinical scales in one embodiment;

[0026] Figure 4e This is a bar chart of the speech index feature scores for each participant group in one embodiment;

[0027] Figure 5a This is one of the schematic diagrams of a clinical scale and neuroimaging features for predicting the severity of Huntington's disease using speech index features in one embodiment;

[0028] Figure 5b This is a second schematic diagram of a clinical scale and neuroimaging features for predicting the severity of Huntington's disease based on speech index features in one embodiment.

[0029] Figure 5c This is the third illustration of a clinical scale and neuroimaging features for predicting the severity of Huntington's disease based on speech index features in one embodiment.

[0030] Figure 5d This is a fourth illustration of a clinical scale and neuroimaging features for predicting the severity of Huntington's disease based on speech index features in one embodiment.

[0031] Figure 5e This is the fifth illustration of a clinical scale and neuroimaging features for predicting the severity of Huntington's disease using speech index features in one embodiment.

[0032] Figure 6 This is a structural block diagram of an acoustic feature processing device in one embodiment;

[0033] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0035] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0036] The acoustic feature processing method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Terminal 102 can extract acoustic features from the speech data of healthy individuals and Huntington's disease patients respectively, obtaining acoustic feature sets; it can perform stability assessment and significant independent contribution assessment on the acoustic features in the acoustic feature sets, obtaining multiple target acoustic features and their corresponding weights; based on the weights corresponding to the target acoustic features, it can perform weighted and scaled processing on the multiple target acoustic features to obtain speech index features characterizing the severity of Huntington's disease. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0037] In one exemplary embodiment, such as Figure 2 As shown, an acoustic feature processing method is provided, which can be applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps S201 to S204. Wherein:

[0038] Step S201: Extract acoustic features from the speech data of healthy individuals and Huntington's disease patients to obtain acoustic feature sets.

[0039] Huntington's disease (HD) is an autosomal dominant neurodegenerative disorder caused by an abnormal amplification of the cytosine-adenine-guanine trinucleotide repeat (CAG) in the Huntington's disease gene (HTT). Approximately 4 out of every 100,000 people are affected by Huntington's, exhibiting the classic triad of motor, cognitive, and psychiatric symptoms. Language impairment is a central feature of Huntington's disease, manifesting as dysarthria, leading to decreased comprehension by others. As a complex motor and cognitive function dependent on the coordinated activity of specific cortical areas, basal ganglia, and cerebellar networks, speech is affected early in Huntington's disease. Impairments in rhythm, articulation, timing, and vocal stability typically precede overt motor symptoms, reflecting the underlying pathology of the disease.

[0040] The study participants, including healthy individuals and those with Huntington's disease (collectively referred to as participants), were asked to read a Chinese text of approximately 400 characters at a normal speed, with no time limit and prior notification of the reading requirements. Voice data from the reading process was recorded using a voice recorder or smartphone, resulting in audio files for each participant in both groups, which were then used as the voice data for the healthy individuals and Huntington's disease patients.

[0041] It can preprocess speech data from healthy individuals and Huntington's disease patients, and extract acoustic features from the preprocessed speech data to obtain prosodic features, temporal features, speech intelligibility features, and timbre features.

[0042] The specific steps for preprocessing voice data from healthy individuals and Huntington's disease patients are as follows:

[0043] First, each audio file is processed through a bandpass filter to maintain frequencies between 80Hz and 8000Hz. Next, background noise in each audio file is suppressed using a noise reduction algorithm, retaining only audio segments with a signal-to-noise ratio (SNR) exceeding 15dB. Finally, to control loudness variations caused by different recording distances, each audio file is normalized, with the root mean square amplitude (RMS) uniformly adjusted to 0.1. To analyze the temporal characteristics of the speech data, a speech activity detection (VAD) tool is used to perform VAD analysis on each audio file, segmenting each file into high-resolution speech and non-speech timelines. To distinguish between significant pauses and normal breathing intervals, all silence intervals are subject to a duration threshold of 200 milliseconds; intervals shorter than 200 milliseconds are considered apnea and disregarded. Conversely, the duration and frequency of silence intervals exceeding 200 milliseconds are quantified as characteristics representing significant pauses or hesitations.

[0044] A multi-stage preprocessing pipeline is employed for each audio file in the speech data. This standardizes the audio files, reduces the impact of environmental noise and reverberation, and isolates speech segments for feature extraction, thereby ensuring the robustness and comparability of the audio files and improving their quality. Specifically, each audio file in the speech data can be processed using audio and music signal analysis tools, resampling it at a uniform sampling rate of 16000Hz to standardize the temporal resolution of all speech signals. Furthermore, the audio can be converted to mono format through channel averaging.

[0045] The prosodic features, temporal features, language intelligibility features, and timbre features of speech data can be standardized to obtain an acoustic feature set. Specifically, a robust scaler can be used to standardize the acoustic features by subtracting the median from each acoustic feature and scaling it according to the interquartile range (IQR). This reduces the potential impact of outliers in the speech data and enhances the robustness of extreme values ​​in subsequent processing steps.

[0046] Step S202: The stability of the acoustic features in the acoustic feature set is evaluated to obtain the candidate acoustic feature set.

[0047] The stability of acoustic features in the acoustic feature set can be evaluated by guided sampling combined with a logistic regression model using the Least Absolute Shrinkage and Selection Operator (LASSO). Acoustic features that meet the stability requirements in the acoustic feature set can be selected as initial candidate acoustic features to form a candidate acoustic feature set.

[0048] Step S203: Evaluate the significant independent contribution of the candidate acoustic features in the candidate acoustic feature set to obtain multiple target acoustic features and their corresponding weights.

[0049] The independent contribution of each candidate acoustic feature can be evaluated by using the first multivariate logistic regression model (multivariate logistic regression model) to distinguish between Huntington's disease patients and healthy controls, thereby verifying the independent role of each candidate acoustic feature. Candidate acoustic features that do not show significant independent contribution can be removed from the candidate acoustic feature set to obtain multiple target acoustic features.

[0050] The multivariate logistic regression model can be adjusted based on the number of feature types of the target acoustic features; the adjusted multivariate logistic regression model can be cross-validated a set number of times based on multiple target acoustic features; and the coefficients in the adjusted multivariate logistic regression model after the last validation can be used as the weights corresponding to the target acoustic features.

[0051] Step S204: Based on the weights corresponding to the target acoustic features, multiple target acoustic features are weighted and scaled to obtain speech index features; speech index features are used to characterize the severity of Huntington's disease.

[0052] Based on the weights corresponding to the target acoustic features, multiple target acoustic features can be enhanced or weakened to obtain speech index features. These speech index features are used to characterize the severity of Huntington's disease.

[0053] In the aforementioned acoustic feature processing method, target acoustic features are selected from the acoustic feature set based on their stability and significant independent contribution, and the weights corresponding to the target acoustic features are determined. Multiple target acoustic features are then weighted and scaled based on their weights to obtain speech index features characterizing the severity of Huntington's disease. Integrating the multidimensional language impairments of Huntington's disease into clinically interpretable speech index features significantly correlated with the severity of Huntington's disease effectively averages the noise and variability associated with individual acoustic features, reduces the impact of environmental noise, and improves the robustness of tracking Huntington's disease progression.

[0054] In one embodiment, acoustic features are extracted from the speech data of healthy individuals and Huntington's disease patients to obtain an acoustic feature set. The specific steps are as follows: based on the fundamental frequency profile of the speech data of healthy individuals and Huntington's disease patients, prosodic features of the speech data are extracted; temporal features, speech intelligibility features, and timbre features of the speech data are extracted; based on the prosodic features, temporal features, speech intelligibility features, and timbre features of the speech data, an acoustic feature set is obtained.

[0055] To construct a comprehensive and clinically relevant speech index, a multidimensional feature selection strategy can be employed, with the primary screening criteria based on biomarkers proven effective in analyzing Mandarin speech in Parkinson's disease patients. To enhance the generality of the acoustic features, acoustic features proven effective in analyzing articulation disorders in non-tonal language environments (such as studies on Parkinson's disease and Huntington's disease) were included.

[0056] It is important to note that the selected acoustic features were chosen based on their resilience to variations in recording conditions. Since the speech data was collected in an outpatient setting, traditional speech quality metrics such as jitter, flicker, and harmonic-to-noise ratio (HNR) were intentionally excluded because they are highly sensitive to the recording environment and microphone quality, in order to avoid introducing potential confounding factors. Instead, prosodic features, temporal features, speech intelligibility features, and timbre features—these acoustic feature dimensions have proven to provide reliable insights even in relatively controlled acoustic environments.

[0057] Based on the fundamental frequency (F0) profile of speech data from healthy individuals and Huntington's disease patients, the characteristics of the pitch and loudness domains can be quantified to obtain the prosodic features of the speech data.

[0058] One approach is to use a speech analysis toolkit to extract the average fundamental frequency (representing average pitch), the standard deviation of the fundamental frequency (representing pitch variability), and the average value of the absolute fundamental frequency slope from the fundamental frequency profile of the speech data, which are the main acoustic factors related to perceived pitch. These values ​​are then used as pitch features. The average value of the absolute fundamental frequency slope can capture introverted dynamics and measure the degree of pitch modulation in the reading task.

[0059] The mean, standard deviation, slope, and kurtosis of the intensity profile of speech data can be extracted from the fundamental frequency profile of the speech data as loudness features. These loudness features quantify the intensity of the speech signal and can describe the loudness dynamics.

[0060] The 13 Mel-frequency cepstral coefficients of the speech data can be extracted step by step according to the short-term analysis window. The average standard deviation of coefficients 1 to 12 is calculated as the time variability feature of the Mel-frequency cepstral coefficients of the speech data. This feature can capture the average fluctuation of the spectrum shape throughout the speech process.

[0061] The temporal variability of Mel-Frequency Cepstral Coefficients (MFCCs) in speech data can be used as a timbre feature, which can be used to evaluate timbre instability.

[0062] Pause features in voice data can be extracted using the automatic methods in the WebRTC Voice Activity Detection Library (webrtcvad). The minimum pause duration is set to 200 milliseconds.

[0063] Speech activity detection can extract many features related to pauses and speech segments from speech data. For example, speech activity detection can extract speech time and pause time from speech data, which can be used to calculate the ratio of speech to pauses and the average or variation length.

[0064] The net syllable rate (NSR) can be calculated by dividing the number of syllables in the text transcribed by Automatic Speech Recognition (ASR) by the net speech time (total duration excluding pauses).

[0065] Features associated with pauses and speech segments in speech data can be used as temporal features of the speech data. Specifically, the temporal features of the speech data are defined as unvoiced and non-consonant signals, including breath intervals following established standards.

[0066] Speech data can be transcribed using an automatic speech recognition model within a large speech model. The transcription results are then compared with purified reference text using a word error rate calculation tool to obtain the substitution rate, deletion rate, and insertion rate of the automatic speech recognition results. This allows for the determination of transcription accuracy and, consequently, the linguistic intelligibility of the speech data.

[0067] Internal confidence scores, such as weighted average and minimum segment log probability, can be extracted from automatic speech recognition models. These scores can serve as indicators of transcription reliability, reflecting the inherent difficulty in recognizing speech signals.

[0068] The replacement rate, deletion rate, and insertion rate of automatic speech recognition results, as well as the internal confidence score of the automatic speech recognition model, can be used as language intelligibility features of speech data.

[0069] The extracted acoustic features are shown in Table 1.

[0070] Table 1

[0071]

[0072] Among them, MAFS stands for Mean Absolute Fundamental Frequency Slope, F0std stands for Fundamental Frequency Standard Deviation, F0Mean stands for Fundamental Frequency Mean, INTSd stands for Intensity Standard Deviation, INTMean stands for Intensity Mean, INTKurt stands for Intensity Kurtosis, INTSkew stands for Intensity Kurtosis, TSDur stands for Intensity Skewness, TSDur stands for Total Speech Duration, NSR stands for Net Syllable Rate, SPR stands for Speech-Pause Ratio, MPL stands for Speech-Pause Ratio, TPDur stands for Total Pause Duration, SdurCV stands for Speech Duration Coefficient of Variation, DPI stands for Duration of Pause Intervals, and SUB stands for Substitution. Rate, WALP stands for Weighted Average Log Power, MSLP stands for Minimum Segment Log Power, DEL stands for Delay, INS stands for Insertion, and MTV stands for Temporal Variability of Mel-Frequency Cepstral Coefficients.

[0073] In this embodiment, prosodic features of speech data are extracted based on the fundamental frequency profile of speech data from healthy individuals and Huntington's disease patients. Based on the prosodic features, temporal features, language intelligibility features, and timbre features of the speech data, an acoustic feature set is obtained that can capture both the language-specific features of speech data and general markers of motor articulation disorders.

[0074] In one embodiment, the acoustic features in the acoustic feature set are evaluated for stability to obtain a candidate acoustic feature set. The specific steps are as follows: the acoustic feature set is subjected to replacement resampling to obtain multiple acoustic features to form an intermediate acoustic feature set; according to the LASSO Regression model, the acoustic features in the intermediate acoustic feature set are screened for importance to obtain the selection frequency of each acoustic feature in the intermediate acoustic feature set; acoustic features with selection frequencies greater than the stability threshold are selected from the intermediate acoustic feature set as initial candidate acoustic features; and a candidate acoustic feature set is obtained based on the initial candidate acoustic features and pre-set demographic covariates.

[0075] The acoustic feature set can be replaced by resampling to obtain 1000 acoustic features; and an intermediate acoustic feature set can be obtained from the 1000 acoustic features.

[0076] A LASSO Regression model can be trained, also known as a minimum absolute shrinkage and selection operator (L1 regularized) logistic regression model. This model can classify participants into Huntington's disease gene carriers and healthy controls. The regularization parameter (C value) of this model can be optimized using five-fold hierarchical cross-validation. The inherent L1 penalty term of the LASSO Regression model precisely zeroes the coefficients of minor features, thus enabling automatic feature selection.

[0077] The importance of acoustic features in the intermediate acoustic feature set can be screened using the LASSO Regression (Minimum Absolute Shrinkage and Selection Operator Logistic Regression) model. The selection frequency of each acoustic feature is determined by the proportion of times it is identified as an important feature (with a non-zero coefficient) in multiple repeated (e.g., 1000 iterations) importance screenings. A higher selection frequency indicates a more stable acoustic feature.

[0078] A stability threshold can be set according to the actual situation.

[0079] Acoustic features with frequencies greater than the stability threshold can be selected from the intermediate acoustic feature set as initial candidate acoustic features.

[0080] It can extract the clinical and demographic characteristics of participants to obtain pre-defined demographic covariates.

[0081] Initial candidate acoustic features and pre-defined demographic covariates can be used as candidate acoustic features to obtain a candidate acoustic feature set.

[0082] In this embodiment, guided sampling and the LASSO Regression model are used to identify stable and concise candidate acoustic features from the acoustic feature set, thereby improving the robustness of tracking the progression of Huntington's disease.

[0083] In one embodiment, the candidate acoustic features in the candidate acoustic feature set are evaluated for significant independent contribution to obtain the target acoustic features and their corresponding weights. The specific steps are as follows: the candidate acoustic features in the candidate acoustic feature set are input into a first multivariate logistic regression model to obtain the significant independent contribution level of each candidate acoustic feature; candidate acoustic features with a significant independent contribution level greater than the independent contribution threshold are selected from the candidate acoustic feature set as target acoustic features; the first multivariate logistic regression model is adjusted according to the number of feature types of the target acoustic features to obtain a second multivariate logistic regression model; the second multivariate logistic regression model is cross-validated a set number of times based on multiple target acoustic features; and the weights corresponding to the target acoustic features are obtained based on the coefficients in the second multivariate logistic regression model after the last validation.

[0084] Candidate acoustic features from the candidate acoustic feature set can be input into the first multivariate logistic regression model to obtain the degree of insignificant independent contribution of each candidate acoustic feature. Based on the degree of insignificant independent contribution of each candidate acoustic feature, the degree of significant independent contribution of each candidate acoustic feature can be obtained.

[0085] An independent contribution threshold can be set according to the actual situation.

[0086] Candidate acoustic features with a significant independent contribution greater than the independent contribution threshold can be selected from the candidate acoustic feature set and used as target acoustic features.

[0087] The first multivariate logistic regression model can be adjusted according to the number of feature types of the target acoustic features to obtain a multivariate logistic regression model that fits the number of feature types of the target acoustic features, which can then be used as the second multivariate logistic regression model.

[0088] A target acoustic feature training set can be obtained based on multiple target acoustic features. The target acoustic feature training set can be divided into a set number of target acoustic feature training subsets; one target acoustic feature training subset can be used as the validation set and the other target acoustic feature training subsets can be used as the training set to perform a set number of cross-validations on the second multivariate logistic regression model.

[0089] For example, the target acoustic feature training set can be divided into five target acoustic feature training subsets; one of the target acoustic feature training subsets can be used as the validation set and the other target acoustic feature training subsets can be used as the training set to perform five cross-validations on the second multivariate logistic regression model.

[0090] After five cross-validation comparisons, the mean area under the receiver operating characteristic (AUC) of the second multivariate logistic regression model (6 features) reached 0.846, which was better than the AUC of 0.838 of the first multivariate logistic regression model (8 features). This confirms that the model refinement process successfully reduced the complexity and noise of the model, thereby improving the generalization of the second multivariate logistic regression model.

[0091] The second multivariate logistic regression model contains four core acoustic features and two demographic features. Multicollinearity can then be evaluated on these six features. The mean variance inflation factor (VIF) of the second multivariate logistic regression model is much lower than the standard threshold of 2, confirming the stability of the second multivariate logistic regression model.

[0092] The coefficients in the second multivariate logistic regression model from the last validation can be used as the weights corresponding to the target acoustic features.

[0093] In this embodiment, the candidate acoustic features in the candidate acoustic feature set can be evaluated for significant independent contribution based on the first multivariate logistic regression model to verify the independent role of each candidate acoustic feature. This yields multiple target acoustic features that have significant independent contribution to distinguishing Huntington's disease patients from healthy controls, as well as the weights corresponding to the target acoustic features. This allows the multidimensional language impairments of Huntington's disease to be integrated into clinically interpretable speech index features that are significantly associated with the severity of Huntington's disease.

[0094] In one embodiment, the target acoustic features include at least one of total speech duration features, substitution rate features, gender features, mean absolute fundamental frequency slope features, time variability features of Mel frequency cepstral coefficients, and age features.

[0095] In the gender characteristics, the gender code for females is 0, and the gender code for males is 1.

[0096] A series of statistical analyses can be performed to assess the clinical relevance and effectiveness of speech index features. Specifically, the performance of speech index features can be validated by Huntington staging, clinical scores (TMS, cUHDRS, SDMT), and caudate and putamen volumes on magnetic resonance imaging. SDMT stands for Symbol Digit Modalities Test.

[0097] Box plots can be used to visualize the distribution of speech index features across the healthy control group, the pre-onset Huntington's disease group, the Huntington's disease stage I group, and the Huntington's disease stage II group. Pairwise comparisons between groups can be performed using the nonparametric Mann-Whitney U test. To control for the false discovery rate (FDR) from multiple comparisons, the degree of insignificant independent contribution (p-value) was adjusted using the Benjamini-Hochberg method.

[0098] To explore the relationship between speech index features and standard clinical scales and magnetic resonance imaging (MRI) parameters for Huntington's disease (e.g., putamen volume and caudate nucleus volume), Spearman's rank correlation coefficient (ρ) was calculated to robustly assess monotonic relationships without assuming linearity. The degree of insignificant independent contribution from all these correlations was adjusted for multiple comparisons using a false discovery rate method. The resulting correlation matrix was visualized as a heatmap to show the interrelationships between features. The standard clinical scales for Huntington's disease can be the Unified Huntington's Disease Rating Scale - Clinical Version (cUHDRS) and the Total Motor Score (TMS).

[0099] To account for the nonlinear relationship between speech index features and the clinical severity score of Huntington's disease, a generalized additive model (GAM) was employed. An independent GAM was constructed for each clinical outcome, with speech index features modeled using a smoothing term. The effective degrees of freedom of the smoothing term were consistently greater than 1, confirming the existence of a significant nonlinear association. The explanatory power of the model was reported using pseudo-R-squared (R²), and the significance of speech index features was assessed using an F-test.

[0100] Validation results showed that the speech index feature significantly increased during the Huntington phase (p<0.001). This speech index feature was strongly correlated with clinical indicators (cUHDRS: ρ=-0.80, p<0.001; TMS: ρ=0.70, p<0.001; SDMT: ρ=-0.74, p<0.001) and neuroimaging biomarkers (caudate nucleus: ρ=-0.54, p<0.001; putamen: ρ=-0.60, p<0.001). A generalized additive model confirmed the high predictive value of the speech index feature for these results (pseudo-R² up to 0.643). Here, p represents the degree of insignificant independent contribution, and ρ represents the Spearman rank correlation coefficient.

[0101] In one embodiment, the target acoustic features include total speech duration features, replacement rate features, gender features, mean absolute fundamental frequency slope features, time variability features of Mel frequency cepstral coefficients, and age features. Based on the weights corresponding to the target acoustic features, multiple target acoustic features are weighted and scaled to obtain speech index features. The specific steps are as follows: Based on the weights corresponding to the total speech duration features, replacement rate features, and gender features, feature enhancement processing is performed on the total speech duration features, replacement rate features, and gender features respectively to obtain target acoustic feature enhancement results; Based on the weights corresponding to the mean absolute fundamental frequency slope features, time variability features of Mel frequency cepstral coefficients, and age features respectively, feature weakening processing is performed on the mean absolute fundamental frequency slope features, time variability features of Mel frequency cepstral coefficients, and age features respectively to obtain target acoustic feature weakening results; The target acoustic feature enhancement results and target acoustic feature weakening results are weighted to obtain speech index features.

[0102] Based on the weights of 1.880 for the total speech duration feature, 1.102 for the replacement rate feature, and 1.033 for the gender feature, feature enhancement processing can be performed on the total speech duration feature, replacement rate feature, and gender feature respectively, resulting in the target acoustic feature enhancement result as: 1.880 × total speech duration feature + 1.102 × replacement rate feature + 1.033 × gender feature.

[0103] Based on the weights of -0.970 for the average absolute fundamental frequency slope feature, 0.710 for the time variability feature of the Mel frequency cepstral coefficient, and 0.011 for the age feature, feature weakening can be performed on the average absolute fundamental frequency slope feature, the time variability feature of the Mel frequency cepstral coefficient, and the age feature, respectively. The target acoustic feature weakening result is: -0.970 × average absolute fundamental frequency slope feature + 0.710 × time variability feature of the Mel frequency cepstral coefficient + 0.011 × age feature.

[0104] The results of target acoustic feature enhancement and weakening can be weighted to obtain a linear combination of weighted and scaled features, which serves as the speech index feature. This involves integrating multidimensional acoustic features (total speech duration, substitution rate, mean absolute fundamental frequency slope, and time variability of Mel-frequency cepstral coefficients) with key demographic data (gender and age) into a continuous score, which is then used as the speech index feature. Typically, the weights for target acoustic feature enhancement and weakening are each 1. Therefore, the speech index feature = 1.880 × total speech duration + 1.102 × substitution rate + 1.033 × gender - 0.970 × mean absolute fundamental frequency slope + 0.710 × time variability of Mel-frequency cepstral coefficients + 0.011 × age.

[0105] In this study, the coefficients from the second multivariate logistic regression model in the last validation were rounded to three decimal places and used as the weights corresponding to the target acoustic features for reporting purposes. To achieve complete repeatability, the numerical precision coefficients used in the speech index feature calculation are shown in Table 2.

[0106] Table 2

[0107]

[0108] In this embodiment, the total speech duration feature, replacement rate feature, and gender feature are respectively subjected to feature enhancement processing to obtain the target acoustic feature enhancement result; the mean absolute fundamental frequency slope feature, Mel frequency cepstral coefficient time variability feature, and age feature are respectively subjected to feature weakening processing to obtain the target acoustic feature weakening result, so as to obtain speech index feature. The multidimensional language impairment of Huntington's disease is integrated into clinically interpretable speech index feature that is significantly associated with the severity of Huntington's disease. This effectively averages the noise and variability associated with individual acoustic features, reduces the influence of environmental noise, and can improve the robustness of tracking the progression of Huntington's disease.

[0109] To better understand the above method, an application embodiment of the acoustic feature processing method of this application is described in detail below.

[0110] Huntington's disease is entering an era of disease-modifying therapy, requiring sensitive, scalable, and objective biomarkers for patient stratification and follow-up. Currently, the most widely used assessment tool for Huntington's disease is the Unified Huntington's Disease Rating Scale (UHRS), but clinical scales are time-consuming and dependent on the rater, while neuroimaging is costly and difficult to obtain.

[0111] In response, this embodiment provides an acoustic feature processing method, which develops and verifies a speech index feature based on automatic acoustic analysis to characterize the severity of Huntington's disease and determine the stage of Huntington's disease.

[0112] The participants in the clinical trial were identified as healthy individuals and patients with Huntington's disease (collectively referred to as participants). The Huntington's disease group consisted of 141 Huntington's gene carriers (37 patients with pre-onset Huntington's disease and 104 patients with dominant Huntington's disease), while the healthy group consisted of 69 healthy controls (HC).

[0113] Participants read a standardized passage aloud, and the recordings were processed to extract 20 acoustic features. A speech index feature was constructed for feature selection using guided sampling, minimum absolute shrinkage, and a selection operator regression model and a multivariate logistic regression model.

[0114] Huntington's disease (HD) is an autosomal dominant neurodegenerative disorder caused by an abnormal amplification of the cytosine-adenine-guanine trinucleotide repeat (CAG) in the huntingtin gene (HTT). Approximately four people per 100,000 suffer from Huntington's disease and exhibit the classic triad of motor, cognitive, and psychiatric symptoms. However, widely used measurements such as the Unified Clinical Rating Scale for Huntington's Disease are inadequate, being both time-consuming and subjective. Furthermore, neuroimaging techniques such as magnetic resonance imaging (MRI) and the detection of the mutant huntingtin protein (mHTT) are costly and sometimes unavailable, hindering participation in care and clinical trials.

[0115] Language impairment is a core feature of Huntington's disease, manifesting as dysarthria that leads to decreased comprehension by others. As a complex motor and cognitive function dependent on the coordinated activity of specific cortical areas, the basal ganglia, and cerebellar networks, speech is affected early in Huntington's disease. Impairments in rhythm, articulation, timing, and vocal stability typically precede overt motor symptoms, reflecting the underlying pathology of the disease. While previous studies have identified specific acoustic abnormalities in Huntington's disease, the clinical utility of speech analysis has been limited by its overemphasis on individual acoustic features, failing to reflect the multifaceted nature of speech degeneration associated with Huntington's disease.

[0116] Current research on speech function in Huntington's disease primarily focuses on isolated acoustic features, such as vowel duration prolongation, reduced fundamental frequency variability, and increased articulation errors. However, understanding the potential pitfalls of these fragmented approaches in translational applications remains crucial. First, isolated indicators only cover a narrow aspect of Huntington's disease-related speech dysfunction, making findings susceptible to environmental noise and reducing robustness in tracking overall clinical disease progression. For example, a patient might exhibit a high substitution rate (indicating articulation deficit) while maintaining normal fundamental frequency variability (measuring prosodic function), or exhibit prolonged speech duration (temporal impairment) but maintain speech intelligibility. Interpreting such incongruous single-feature patterns to assess the overall disease burden has been subjective and imprecise. Furthermore, the richness of underlying acoustic parameters, including over 50 commonly reported speech features, presents significant dimensional challenges, potentially leading to overfitting in small patient populations and inconsistencies between studies. Therefore, despite strong evidence linking speech dysfunction to Huntington's disease progression and neuropathology, speech analysis remains a descriptive tool rather than a standardized, objective endpoint for routine patient monitoring or clinical trials.

[0117] To address the aforementioned limitations, this embodiment develops a machine learning-derived speech index feature, integrating multiple key acoustic features into a single continuous score. Using a data-driven approach, guided resampling, minimum absolute contraction, and selection operator regression and multivariate logistic regression models are employed to stably select features. This integrates information from different speech domains, including total speech duration (timing feature), substitution rate (language intelligibility), mean absolute fundamental frequency slope (prosody), and time variation of Mel-frequency cepstral coefficients (timbre), forming a unified speech index feature.

[0118] This embodiment investigates the correlation between speech index features and traditional clinical and neuroimaging biomarkers, and proposes that speech index features can serve as a quantitative, non-invasive, and clinically interpretable measure of the severity of Huntington's disease.

[0119] Research design and experimental workflow, such as Figure 3As shown. The study included 141 Huntington's disease amplification carriers (including 37 patients with pre-symptom Huntington's disease and 104 patients with dominant Huntington's disease) and 69 age- and sex-matched healthy controls. All participants were required to read a standardized article to obtain speech data. The speech data were preprocessed through splicing, truncation, denoising, and dereverberation. Using an open-source toolkit, 20 acoustic and linguistic features were extracted from the speech data. These features were divided into four domains: prosodic features, encompassing three pitch features and four loudness features; timbre features, represented by the time variation of Mel frequency cepstral coefficients; temporal features, including seven features related to speech rate and pause structure; and linguistic intelligibility features, including two transcription confidence scores and three word error rates. To obtain robust and unbiased estimates of the model's classification performance, a five-fold stratified cross-validation procedure was employed across the entire dataset. Stratification ensured that the proportions of Huntington's disease patients and healthy individuals were preserved in each fold, thus reliably assessing the model's generalization ability to new data. The speech index feature was initially used to distinguish Huntington's disease carriers from healthy controls. A robust speech index feature was constructed by selecting target acoustic features through guided resampling, minimum absolute shrinkage, and a choice operator regression model and a multivariate logistic regression model. The practicality of the speech index feature in predicting clinical scores (such as cUHDRS, TMS, SDMT) and its correlation with caudate and putamen volumes was validated.

[0120] The study involves three key objectives: (1) developing speech index features; (2) determining whether speech index features reflect the severity of Huntington's disease; and (3) determining whether speech index features reflect neurodegeneration indicated by magnetic resonance imaging.

[0121] I. Obtaining Voice Data:

[0122] The participants included 69 age- and sex-matched healthy controls and 141 Huntington's disease carriers (including 37 patients with pre-symptom Huntington's disease, 67 patients with stage I Huntington's disease, and 37 patients with stage II Huntington's disease). Participants were drawn from the Guangzhou Clinical Trial Preparation Huntington's Disease Cohort. A convenience sample of 210 individuals was included in the cross-sectional study and divided into three groups: age- and sex-matched healthy controls, the pre-symptom Huntington's disease (preHD) group, and the manifest Huntington's disease (mHD) group. The manifest Huntington's disease group was further divided into stage I and stage II Huntington's disease groups based on overall functional score. The inclusion criteria for the Huntington's disease patient population were: 1) ≥36 trinucleotide repeats in the sequence confirmed by Huntington's gene testing; 2) normal vision / hearing; 3) ability to understand and read text; 4) Mandarin proficiency; and 5) willingness to provide written informed consent. Patients with other neurological disorders unrelated to Huntington's disease were excluded from the study.

[0123] Participants underwent clinical profiling, including trinucleotide repeat length measurement and analysis using the standard Huntington's Disease Scale, including the Total Motor Score (TMS), Total Functional Capacity (TFC), Independence Score (IS), Symbol Digit Modalities Test (SDMT), Category Fluency Test (CFT), the three components of the Stroop Test (Stroop Word Test (SWT), Stroop Color Test (SCT), and Stroop Interference Test (SIT)), the Mini-Mental State Examination (MMSE), and Trail Making Test - Part (TMT) A and B. The overall UHDRS score (cUHDRS), trinucleotide repeat sequence (CAG), and CAG Age Product (CAP) were also assessed. Each participant then read a standardized Mandarin passage aloud, which was recorded as audio data for subsequent analysis. In addition, 96 patients with Huntington's disease underwent 3-Tesla Magnetic Resonance Imaging (3TMRI), and the volume of the putamen and caudate nucleus was quantified using brain imaging analysis software.

[0124] The clinical and demographic characteristics of the participants are shown in Table 3. There were no significant differences in sex distribution among the groups, but patients in both the stage I and stage II Huntington's disease groups were older than the healthy control group and the pre-symptom group (p<0.001). Furthermore, patients in the stage II Huntington's disease group had a longer trinucleotide repeat sequence count than patients in the pre-symptom and stage I groups. As expected, advanced-stage patients had lower CAP, TMS, TFC, and IS scores, and performed worse on cognitive assessments (SDMT, CFT, SWT, SCT, SIT, MMSE, and TMT) than the pre-symptom group, reflecting more severe motor and cognitive impairment. Data are presented as mean ± standard deviation (SD).

[0125] Table 3

[0126]

[0127] 2. Extracting and analyzing acoustic features:

[0128] Speech data underwent preprocessing (resampling, noise reduction, and normalization) before acoustic feature extraction, yielding a set of 20 multidimensional prosodic, timbre, temporal, and linguistic features. The selection strategy was guided by two key research areas to ensure comprehensive coverage of speech pathology. First, effective features from studies analyzing reading tasks in Mandarin-speaking Parkinson's disease patients were included. Second, established indicative features from current acoustic analysis of Huntington's disease were included, particularly those validated in remote data collection environments. To ensure robustness, features known to be sensitive to recording quality were excluded based on previous findings. Feature extraction was performed using open-source toolkits.

[0129] Speech features differed among different Huntington's disease groups and the healthy control group. To explore the speech features at different stages of Huntington's disease, a set of speech features was first extracted, as shown in Table 4. These features were divided into four categories: prosodic features, temporal features, speech intelligibility features, and timbre features. In the prosodic domain, five of the seven features showed significant differences among the different groups, including MAFS (reflecting pitch dynamics), F0std (capturing pitch changes), INTSd (reflecting intensity fluctuations), INTMean (representing overall loudness), and INTSkew (representing asymmetry in intensity distribution). Regarding temporal features, five features showed significant differences, including TSDur (total speech duration), NSR (syllable production rate), MPL (mean pause duration), TPDur (total pause duration), and DPI (pause habit). Most significantly, all five features in speech intelligibility showed significant differences among the three groups. These include SUB (reflecting articulation and articulation accuracy), WALP (overall ASR confidence level), MSLP (lowest ASR confidence level), DEL (transcriptional omission), and INS (transcriptional insertion). Finally, the MTV (articulation clarity / stability) feature within timbre also showed differences between groups.

[0130] Table 4

[0131]

[0132] III. Extracting speech index features:

[0133] To determine a stable and concise set of speech index features to distinguish between Huntington's gene carriers and healthy controls, a multi-stage model development and validation process was implemented. First, guided sampling and minimal absolute shrinkage with a selection operator logistic regression (1000 iterations) were used to evaluate the stability of all acoustic features. This procedure identified acoustic features with a selection rate exceeding 95% as initial candidate acoustic features.

[0134] Next, the initial candidate acoustic features are combined with pre-specified demographic covariates (age and gender) to obtain candidate acoustic features. Based on the candidate acoustic features, a first multivariate logistic regression model (which can be called the preliminary multivariate logistic regression model) is fitted. The first multivariate logistic regression model is then used to remove all statistically insignificant acoustic predictors (p>0.05) from the candidate acoustic features to obtain the target acoustic features. The first multivariate logistic regression model is adjusted according to the number of feature types of the target acoustic features to obtain the second multivariate logistic regression model.

[0135] Finally, a target acoustic feature training set can be obtained based on multiple target acoustic features. Using this training set, the performance and robustness of the second multivariate logistic regression model are evaluated by generating non-folded predictions within a five-fold cross-validation framework. The coefficients from the last validated second multivariate logistic regression model are used as the weights corresponding to the target acoustic features.

[0136] Based on the weights corresponding to the target acoustic features, multiple target acoustic features are weighted and scaled to obtain speech index features used to characterize the severity of Huntington's disease.

[0137] IV. Clinical validation of speech index features:

[0138] The performance of the speech index feature was validated using Huntington's staging, clinical scores (TMS, cUHDRS, SDMT), and caudate and putamen volumes on MRI. Validation results showed a significant increase in the speech index feature during Huntington's stage (p<0.001). This speech index feature was strongly correlated with clinical indicators (cUHDRS: ρ=-0.80, p<0.001; TMS: ρ=0.70, p<0.001; SDMT: ρ=-0.74, p<0.001) and neuroimaging biomarkers (caudate nucleus: ρ=-0.54, p<0.001; putamen: ρ=-0.60, p<0.001). A generalized additive model confirmed the high predictive value of the speech index feature for these results (pseudo-R² up to 0.643). Here, p represents the degree of insignificant independent contribution, and ρ represents the Spearman rank correlation coefficient. The aforementioned fully automated and interpretable speech index features can serve as an effective digital biomarker for the severity of Huntington's disease, and are expected to enable remote monitoring, enrich clinical trials, and objectively assess treatment effectiveness.

[0139] Specifically, the distribution of speech index features in participant groups (healthy control group, pre-symptom group of Huntington's disease, and overt Huntington's disease group) and their correlation with clinical and magnetic resonance imaging indicators were assessed, and the nonlinear relationship was validated using Spearman ranking correlation and generalized additive models. Statistical significance was set at p<0.05.

[0140] Next, the study investigated whether speech index features could distinguish between Huntington's gene carriers and healthy controls. It was observed that speech index features effectively distinguished between Huntington's gene carriers and healthy controls, achieving an AUC of 0.846 (95% confidence interval 0.787-0.896, estimated at 1000 bootstrap resampling, p < 0.001). Figure 4a (As shown).

[0141] Given the complexity of Huntington's disease symptoms, relying on a single speech feature is challenging. Therefore, this embodiment develops a comprehensive speech index feature (which may be called a speech index feature). A set of parsimony and stable target acoustic features is selected and constructed using guided sampling, minimum absolute contraction, and a logistic regression model with selection operators, as well as a multivariate logistic regression model, while controlling for age and gender as pre-specified covariates. Four key target acoustic features were identified: TSDur, SUB, MAFS, and MTV, which, along with demographic variables, constitute the speech index feature (e.g., Figure 4b (As shown). It is worth noting that the four key target acoustic features—TSDur, SUB, MAFS, and MTV—represent different speech dimensions, and the TSDur and SUB scores increase with the progression of Huntington's disease (e.g., Figure 4c (As shown). Specifically, the duration of TSDur was significantly longer in the stage I and stage II Huntington's disease groups than in the dominant Huntington's disease group, although there was no significant difference between the dominant Huntington's disease group and the healthy control group. In contrast, MAFS decreased with disease progression. MTV showed a significant pattern: it was higher in the dominant Huntington's disease group compared to the healthy control group, but decreased in the stage II Huntington's disease group, suggesting a possible early compensatory mechanism that weakens in the later stages.

[0142] Further investigation could be conducted to determine whether the four target acoustic features—TSDur, SUB, MAFS, and MTV—are associated with symptoms of Huntington's disease. It is possible to observe the correlation between TSDur and all cognitive scales (such as...). Figure 4dThe speech index features showed a high correlation with TSDur (r=-0.70), particularly with SDMT. Since SDMT primarily reflects information processing speed, this aligns well with TSDur as an indicator of speech pause duration. TSDur also showed a high correlation with TMS (r=0.57), consistent with the expectation that more severe motor symptoms lead to longer pauses. Furthermore, SUB was highly correlated with both motor and cognitive scales, which is understandable given that speech intelligibility and articulation accuracy are influenced by both motor control and cognitive processing. Meanwhile, MTV correlated more strongly with TMS than with cognitive scales. MTV serves as an important indicator of articulation stability, which is closely related to motor function. Of the four target acoustic features included in the speech index features, three strongly correlated with TMS, and one more closely correlated with cognitive performance, indicating that the speech index features capture the interleaved motor and cognitive impairments in Huntington's disease. Finally, the differences in speech index features across different disease stages were assessed. Figure 4e As shown, speech index features increased with disease severity, with significant differences (both <0.001) between the pre-symptom group and the dominant Huntington's disease group (Stage I), and between the dominant Huntington's disease group (Stage I) and the dominant Huntington's disease group (Stage II). Taken together, these findings highlight the potential of speech index features as a comprehensive biomarker for Huntington's disease.

[0143] in, Figure 4a The Receiver Operating Characteristic (ROC) curve shows the performance of the speech index feature in classifying healthy controls and Huntington's gene carriers, achieving an area under the curve (AUC) of 0.846. The shaded area represents the 95% confidence interval, estimated based on 1000 bootstrap resampling. The dashed line represents the random classifier. False Positive Rate, specificity, True Positive Rate, and Sensitivity (D) represent the false positive rate, specificity, true positive rate, and sensitivity, respectively. Figure 4bThe forest plot shows the confidence intervals of the coefficients (β) of the second multivariate logistic regression model. Logistic regression coefficients represent logistic regression coefficients, and features represent features. Feature selection employs a two-stage process: First, the most stable initial candidate acoustic features are identified from the acoustic feature set by guided sampling combined with minimum absolute shrinkage and selection operators in the logistic regression model. Second, these stable initial candidate acoustic features, along with fixed demographic covariates (age and sex), are entered into the first logistic regression model as candidate acoustic features. This model is then refined by removing all features with insignificant p-values ​​(p>0.05) to obtain the target acoustic features and the second logistic regression model. The coefficients in this second logistic regression model define the weights of each variable in calculating the speech index features. The selected target acoustic features and their weights include total speech duration (1.880), substitution rate (1.102), mean absolute fundamental frequency slope (-0.970), temporal variability of Mel frequency cepstral coefficients (0.710), sex (1.033), and age (0.011). Figure 4c This is a radar chart showing the acoustic characteristics of four participating groups: healthy control (HC), pre-HD (pre-Hunting disease), stage I (Hunting disease), and stage II (Hunting disease). After adjusting for age and sex using analysis of covariance (ANCOVA), the acoustic characteristics showed a significant difference in MTV between the pre-HD and healthy control groups as disease severity decreased or increased. These changes reflect the worsening of language impairment as the disease progresses. The green line represents the healthy control group; the blue line represents the pre-HD group; the purple line represents stage I; and the red line represents stage II. Figure 4d The heatmaps show partial correlations between four acoustic features and a set of clinical scales. These correlations have been adjusted for age and sex, and p-values ​​for multiple comparisons have been corrected using a false discovery rate method. Figure 4e This is a bar graph showing the Speech Index characteristic scores for each participant group: healthy control (HC), pre-HD (Pre-HD), Stage 1 (Stage 1), and Stage 2 (Stage 2). The Speech Index characteristic increases significantly with disease severity, with significant differences between the Pre-HD, Stage 1, and Stage 2 groups. Higher Speech Index characteristics indicate more severe speech impairment. The error bars represent the standard deviation of the mean.

[0144] The speech index feature was validated to predict disease severity in clinical and neuroimaging parameters. Specifically, the sensitivity and specificity of the speech index feature in tracking disease severity were further explored. Good correlations were observed between the speech index feature and the volumes of the TMS, cUHDRS, SDMT, caudate nucleus, and putamen (ρ = 0.70, -0.80, -0.74, -0.54, and -0.60, p < 0.0001). This indicates that the speech index feature closely reflects motor and cognitive impairments, as well as striatal atrophy—a core neuropathological feature of Huntington's disease. Regression analysis further confirmed a high degree of agreement between the speech index feature and the actual volumes of the TMS, cUHDRS, SDMT, caudate nucleus, and putamen (pseudo R² = 0.464, 0.643, 0.566, 0.297, and 0.361, p < 0.0001). Figures 5a-5e (As shown). In summary, these findings support the use of speech index features as a sensitive, non-invasive, and scalable tool for monitoring Huntington's disease progression and assessing the effectiveness of Huntington's disease treatment responses.

[0145] This embodiment utilizes machine learning methods to develop speech index features as a potential quantitative biomarker for Huntington's disease and validates the correlation between speech index features and clinical indicators such as cUHDRS and TMS. Consistent with the progression of Huntington's disease, the speech index features were found to sequentially increase from the pre-symptom stage to stage I and then to stage II. More importantly, the speech index features are strongly negatively correlated with striatal atrophy, especially the loss of volume in the caudate nucleus and putamen, which are the earliest degenerate nuclei in the pathogenesis of Huntington's disease. These results indicate that speech analysis not only captures functional impairments but also reflects underlying brain changes.

[0146] Compared to previous studies on speech function in Huntington's disease (primarily focusing on isolated acoustic features), this embodiment addresses a key gap by integrating multidimensional language impairments into a single, clinically interpretable indicator—the speech index feature—using machine learning methods. The association between the speech index feature and overall Huntington's disease severity is significantly enhanced. Guided resampling combined with minimum absolute contraction and selection operator regression techniques was employed to develop the speech index feature. This feature not only achieves transparent feature weighting (prioritizing substitution errors and prolonged speech duration as primary discriminative features) but also avoids the inherent interpretability limitations of "black box" deep learning models. Furthermore, it effectively averages noise and variability associated with individual indicators, directly mitigating the weaknesses of single-feature methods. Notably, a conceptually similar multidimensional framework has been validated in Mandarin-speaking Parkinson's disease patients, enabling cross-disease insights. Consistent with previous studies, this embodiment reveals significant MAFS manifestations in Huntington's disease, indicating a common prosodic disorder in both high motor (HD) and low motor (PD) articulation disorders. However, among the speech index features, the significant weight of MTV reflects spectral instability, while SUB is used for speech intelligibility, which is uniquely linked to the choreiform movements and articulation inaccuracies characteristic of Huntington's disease. This distinction highlights the value of the domain-adaptive approach: instead of adopting a one-size-fits-all feature set, it prioritizes indicators that align with the unique pathophysiology of Huntington's disease, ensuring that speech index features are both universal and disease-specific.

[0147] A key finding is the strong correlation between speech index features and clinical disease severity, particularly with the comprehensive UHDRS (cUHDRS; ρ = -0.80, p < 0.001). As a validated endpoint in Huntington's disease clinical trials, cUHDRS integrates motor, cognitive, and functional abilities. Speech index features alone explain approximately 64.3% of the variance in cUHDRS scores (GAM model, pseudo-R² = 0.643), highlighting the potential of speech index features as a surrogate endpoint. The robust performance of speech index features can be interpreted as a reflection of the neuropathology of Huntington's disease. Speech production is a complex motor and cognitive task involving cortical, striatal, and cerebellar networks, all of which are affected by Huntington's disease. Speech index features are likely to reflect functional impairment caused by progressive neurodegenerative diseases. Their strong discriminative power suggests that language regression is a highly sensitive functional endpoint that resonates with biological disease progression, making it an ideal tool for clinical research and treatment development.

[0148] Striatal atrophy is the earliest neuroimaging marker of neurodegenerative changes associated with Huntington's disease. Neuroanatomically, the cortex and caudate nucleus regulate the initiation and execution of speech motor functions via the corticocaudate nucleus connection. Huntington's disease selectively degenerates intermediate-sized spinous neurons in these regions, disrupting the cortico-striatal-thalamic-cortical circuit; this disruption manifests as speech impairments (such as rhythmic variations, inaccurate articulation, and prosodic abnormalities). This embodiment demonstrates a strong correlation between speech index features and the volume of key basal ganglia structures. Previous studies of Huntington's disease speech either focused solely on describing speech abnormalities or employed simple correlation analyses with limited acoustic features, resulting in weak associations between the brain and speech (ρ = 0.30–0.45). In contrast, the speech index features in this embodiment show a significantly enhanced correlation with structural neuroimaging indices, most notably with shell volume (ρ = 0.60). This indicates that overall speech assessment accurately reflects the inherent distributed neural network dysfunction in Huntington's disease. The 36.1% variation further indicates that the speech index feature goes beyond simply detecting striatal atrophy, reflecting the impact of downstream striatal degeneration on cortical function and network connectivity. Therefore, the speech index feature is not only a related indicator of atrophy, but also a functional indicator of how striatal damage disrupts the brain's integrative networks necessary for fluent speech.

[0149] Digital biomarkers, including speech analysis, video-based assessments, and wearable sensor technologies, have emerged as promising tools for quantifying the severity of Huntington's disease. Video biomarkers quantify overt motor symptoms, such as chorea, facial motor dysfunction, and oculomotor dysfunction, which are highly correlated with UHDRS motor scores and capture nonverbal impairments. Wearable devices, utilizing accelerometers, can continuously monitor motor activity, sleep patterns, and daily functioning. In contrast, the speech index feature of this embodiment offers a unique, comprehensive, and practical alternative: it integrates motor and language impairments through a short, standardized task (≤5 minutes) that can be performed on demand via a smartphone. Furthermore, the acoustic feature processing method provided in this embodiment avoids the burden of continuous use and high data storage requirements of wearable devices, as well as the visual privacy risks associated with video tools, while integrating motor, cognitive, and language neurodomains into a single interpretable indicator. In addition to practical advantages, the speech index feature demonstrates robust performance relative to established Huntington's disease biomarkers. For example, plasma NfL, as a gold standard peripheral biomarker, is associated with clinical severity scores ρ=0.50–0.70, but requires invasive blood collection and centralized laboratory processing, while the voice index feature achieves comparable or better correlations in a non-invasive, remote manner (e.g., ρ=0.70 for TMS).

[0150] In summary, this embodiment developed and validated a machine learning-derived speech index feature as a sensitive, objective, and interpretable digital biomarker for the severity of Huntington's disease. Its high correlation with clinical-scale and neuroimaging biomarkers, coupled with its remote operability, makes it a powerful tool for improving the efficiency of clinical trials and enabling frequent, objective monitoring of disease progression and treatment response.

[0151] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0152] Based on the same inventive concept, this application also provides an acoustic feature processing apparatus for implementing the acoustic feature processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more acoustic feature processing apparatus embodiments provided below can be found in the limitations of the acoustic feature processing method described above, and will not be repeated here.

[0153] In one exemplary embodiment, such as Figure 6 As shown, an acoustic feature processing device is provided, wherein:

[0154] The acoustic feature set acquisition module 601 is used to extract acoustic features from the speech data of healthy people and Huntington's disease patients respectively to obtain acoustic feature sets;

[0155] The candidate acoustic feature set acquisition module 602 is used to evaluate the stability of the acoustic features in the acoustic feature set to obtain the candidate acoustic feature set.

[0156] The target acoustic feature and weight acquisition module 603 is used to evaluate the significant independent contribution of the candidate acoustic features in the candidate acoustic feature set to obtain multiple target acoustic features and the weights corresponding to the target acoustic features.

[0157] The speech index feature acquisition module 604 is used to perform weighted and scaled processing on the multiple target acoustic features based on the weights corresponding to the target acoustic features to obtain speech index features; the speech index features are used to characterize the severity of Huntington's disease.

[0158] In one embodiment, the acoustic feature set acquisition module 601 is used to: extract prosodic features of the speech data based on the fundamental frequency profile of speech data from healthy individuals and Huntington's disease patients; extract temporal features, speech intelligibility features, and timbre features of the speech data; and obtain an acoustic feature set based on the prosodic features, temporal features, speech intelligibility features, and timbre features of the speech data.

[0159] In one embodiment, the candidate acoustic feature set acquisition module 602 is configured to: perform replacement resampling on the acoustic feature set to obtain multiple acoustic features to form an intermediate acoustic feature set; perform importance screening on the acoustic features in the intermediate acoustic feature set according to the LASSO Regression model to obtain the selection frequency of each acoustic feature in the intermediate acoustic feature set; select acoustic features with selection frequencies greater than a stability threshold from the intermediate acoustic feature set as initial candidate acoustic features; and obtain a candidate acoustic feature set based on the initial candidate acoustic features and a pre-set demographic covariate.

[0160] In one embodiment, the target acoustic feature and weight acquisition module 603 is configured to: input candidate acoustic features from the candidate acoustic feature set into a first multivariate logistic regression model to obtain the significant independent contribution level of each candidate acoustic feature; select candidate acoustic features from the candidate acoustic feature set whose significant independent contribution level is greater than the independent contribution threshold as the target acoustic features; adjust the first multivariate logistic regression model according to the number of feature types of the target acoustic features to obtain a second multivariate logistic regression model; perform a set number of cross-validations on the second multivariate logistic regression model based on multiple target acoustic features; and obtain the weights corresponding to the target acoustic features based on the coefficients in the second multivariate logistic regression model after the last validation.

[0161] In one embodiment, the target acoustic features include at least one of total speech duration features, replacement rate features, gender features, mean absolute fundamental frequency slope features, time variability features of Mel frequency cepstral coefficients, and age features.

[0162] In one embodiment, the speech index feature acquisition module 604 is configured to: perform feature enhancement processing on the total speech duration feature, the replacement rate feature, and the gender feature based on the weights corresponding to the total speech duration feature, the replacement rate feature, and the gender feature, respectively, to obtain a target acoustic feature enhancement result; perform feature weakening processing on the average absolute fundamental frequency slope feature, the time variability feature of the Mel frequency cepstral coefficient, and the age feature based on the weights corresponding to the average absolute fundamental frequency slope feature, the time variability feature of the Mel frequency cepstral coefficient, and the age feature, respectively, to obtain a target acoustic feature weakening result; and perform weighted processing on the target acoustic feature enhancement result and the target acoustic feature weakening result to obtain the speech index feature.

[0163] Each module in the aforementioned acoustic feature processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0164] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores data for embodiments of the acoustic feature processing method. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an acoustic feature processing method.

[0165] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0166] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0167] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0168] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0169] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0170] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0171] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0172] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An acoustic feature processing method, characterized in that, The method includes: Acoustic features were extracted from the speech data of healthy individuals and Huntington's disease patients to obtain acoustic feature sets. The stability of the acoustic features in the acoustic feature set is evaluated to obtain a candidate acoustic feature set; The candidate acoustic features in the candidate acoustic feature set are evaluated for significant independent contribution to obtain multiple target acoustic features and their corresponding weights. The multiple target acoustic features are weighted and scaled based on the weights corresponding to the target acoustic features to obtain speech index features; the speech index features are used to characterize the severity of Huntington's disease.

2. The method according to claim 1, characterized in that, The acoustic feature set is obtained by extracting acoustic features from the speech data of healthy individuals and Huntington's disease patients, including: Based on the fundamental frequency profile of speech data from healthy individuals and Huntington's disease patients, prosodic features of the speech data are extracted. Extract the temporal features, language intelligibility features, and timbre features from the speech data; An acoustic feature set is obtained based on the prosodic features, temporal features, language intelligibility features, and timbre features of the speech data.

3. The method according to claim 1, characterized in that, The stability evaluation of the acoustic features in the acoustic feature set to obtain a candidate acoustic feature set includes: The acoustic feature set is subjected to replacement resampling to obtain multiple acoustic features, thereby forming an intermediate acoustic feature set; Based on the LASSO Regression model, the acoustic features in the intermediate acoustic feature set are screened by importance to obtain the selection frequency of each acoustic feature in the intermediate acoustic feature set. Acoustic features with frequencies greater than a stability threshold are selected from the intermediate acoustic feature set as initial candidate acoustic features. Based on the initial candidate acoustic features and the pre-defined demographic covariates, a candidate acoustic feature set is obtained.

4. The method according to claim 1, characterized in that, The candidate acoustic features in the candidate acoustic feature set are evaluated for significant independent contribution to obtain the target acoustic features and their corresponding weights, including: The candidate acoustic features in the candidate acoustic feature set are input into the first multivariate logistic regression model to obtain the significant independent contribution of each candidate acoustic feature. Candidate acoustic features with a significant independent contribution greater than the independent contribution threshold are selected from the candidate acoustic feature set and used as the target acoustic features; The first multivariate logistic regression model is adjusted according to the number of feature types of the target acoustic features to obtain the second multivariate logistic regression model; The second multivariate logistic regression model is cross-validated a set number of times based on multiple target acoustic features; The weights corresponding to the target acoustic features are obtained based on the coefficients in the second multivariate logistic regression model in the last validation.

5. The method according to claim 1, characterized in that, The target acoustic features include at least one of the following: total speech duration feature, replacement rate feature, gender feature, mean absolute fundamental frequency slope feature, time variability feature of Mel frequency cepstral coefficients, and age feature.

6. The method according to claim 5, characterized in that, The target acoustic features include total speech duration features, substitution rate features, gender features, mean absolute fundamental frequency slope features, time variability features of Mel frequency cepstral coefficients, and age features. The speech index features are obtained by weighting and scaling the multiple target acoustic features based on their corresponding weights, including: Based on the weights corresponding to the total speech duration feature, the replacement rate feature, and the gender feature, feature enhancement processing is performed on the total speech duration feature, the replacement rate feature, and the gender feature respectively to obtain the target acoustic feature enhancement result; Based on the weights corresponding to the average absolute fundamental frequency slope feature, the time variability feature of the Mel frequency cepstral coefficient, and the age feature, feature weakening processing is performed on the average absolute fundamental frequency slope feature, the time variability feature of the Mel frequency cepstral coefficient, and the age feature respectively to obtain the target acoustic feature weakening result. The target acoustic feature enhancement result and the target acoustic feature weakening result are weighted to obtain speech index features.

7. An acoustic feature processing device, characterized in that, The device includes: The acoustic feature set acquisition module is used to extract acoustic features from the speech data of healthy people and Huntington's disease patients respectively, and obtain acoustic feature sets. The candidate acoustic feature set acquisition module is used to evaluate the stability of the acoustic features in the acoustic feature set to obtain the candidate acoustic feature set. The target acoustic feature and weight acquisition module is used to evaluate the significant independent contribution of the candidate acoustic features in the candidate acoustic feature set to obtain multiple target acoustic features and the weights corresponding to the target acoustic features. The speech index feature acquisition module is used to perform weighted and scaled processing on the multiple target acoustic features based on the weights corresponding to the target acoustic features to obtain speech index features; the speech index features are used to characterize the severity of Huntington's disease.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.