A depression degree evaluation method based on acoustic feature sparse manifold topology

By constructing a depression assessment method based on sparse manifold topology and utilizing hybrid kernel functions and sharpening mechanisms, the prediction bias caused by individual heterogeneity in existing technologies is solved, achieving a high-precision and high-reliability depression assessment.

CN122493890APending Publication Date: 2026-07-31NANHU BRAIN COMPUTER CROSS RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANHU BRAIN COMPUTER CROSS RES INST
Filing Date
2026-04-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for assessing the severity of depression suffer from problems such as high subjectivity, long time consumption, difficulty in achieving high-frequency daily monitoring, and limited prediction accuracy. In particular, there are prediction biases due to differences in acoustic representation caused by individual heterogeneity and atypical sample prediction bias caused by ignoring the local topological structure of the data.

Method used

An evaluation method based on acoustic feature sparse manifold topology is adopted. A sparse association topology is constructed using an improved Nadaraja-Watson estimation method. The similarity potential is calculated by a hybrid kernel function and a boundary threshold and sharpening mechanism are introduced to construct a sparse manifold topology, thereby improving the accuracy and robustness of depression level prediction.

Benefits of technology

It improves the accuracy and reliability of depression assessment, reduces the mean absolute error and variance of prediction results, enhances the ability to focus on highly similar samples, and achieves a high degree of consistency with clinical scale scores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493890A_ABST
    Figure CN122493890A_ABST
Patent Text Reader

Abstract

This invention discloses a method for assessing depression levels based on sparse manifold topology of acoustic features. The method includes: collecting speech data and corresponding clinical scale scores of a target subject at a preset time frequency; extracting acoustic feature vectors to form a historical reference sample library; collecting speech data of the target subject at the time of the test; extracting acoustic features to obtain the acoustic feature vector to be tested; combining Euclidean distance and cosine similarity, using a hybrid kernel function to calculate the hybrid similarity potential between the acoustic feature vector to be tested x and the acoustic feature vectors of each historical reference sample; based on the hybrid similarity potential, introducing a boundary threshold and sharpening mechanism to construct a sparse manifold topology and calculate normalized weights; and calculating a predicted depression score based on an improved Nadaraja-Watson estimation model to achieve depression level assessment. This invention can more accurately reflect the patient's current depression level and improve the accuracy and robustness of depression level prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer data processing and medical auxiliary diagnosis technology, specifically to a method for assessing depression based on acoustic feature sparse manifold topology. Background Technology

[0002] Depression is a mental illness characterized by persistent low mood, lack of energy, and loss of interest. It is the leading cause of disability and a major contributor to the disease burden worldwide. Currently, the assessment of the severity of depression in patients mainly relies on physician interviews and standardized scales such as the Hamilton Depression Rating Scale. However, this traditional diagnostic and treatment model suffers from problems such as high subjectivity, time-consuming nature, and difficulty in achieving high-frequency daily monitoring.

[0003] In recent years, with the development of artificial intelligence technology, automatic assessment of depression levels based on speech signals has become a research hotspot. Most existing solutions employ parametric models such as machine learning or deep learning models to construct a global mapping function from acoustic features to depression scores. However, the acoustic representation of depression exhibits strong individual heterogeneity, with significant differences among patients. Existing parametric models suffer from limited prediction accuracy due to their inability to eliminate baseline differences between individuals, and they tend to overlook local topological structures within the data, leading to significant prediction biases for atypical samples. Traditional non-parametric models, typically based on linear weighting of the entire sample, tend to average out predictions and struggle to capture subtle variations in features. Therefore, constructing a reliable and accurate assessment method for predicting patient depression levels remains a challenging problem to solve. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by providing a method for assessing depression levels based on acoustic feature-based sparse manifold topology. This invention utilizes an improved Nadaraja-Watson estimation method to construct a sparse association topology, and employs weighted inference based on the pairwise potential energy of the test sample and historical samples, thereby improving the accuracy and robustness of depression level prediction.

[0005] The objective of this invention is achieved through the following technical solution: a method for assessing depression levels based on acoustic feature sparse manifold topology, comprising the following steps: (1) Collect the speech data of the target object and the corresponding clinical scale scores according to the preset time frequency, extract the acoustic feature vector, and form a historical reference sample library; (2) Collect the speech data of the target object at the time of the test, extract the acoustic features, and obtain the acoustic feature vector to be tested; (3) Combining Euclidean distance and cosine similarity, the mixture kernel function is used to calculate the mixture similarity potential between the acoustic feature vector x to be tested and the acoustic feature vectors of each historical reference sample; (4) Based on the hybrid similarity potential, a boundary threshold and sharpening mechanism are introduced to construct a sparse manifold topology and calculate the normalized weights; (5) Calculate the predicted value of depression score based on the improved Nadalaya-Watson estimation model to realize the assessment of depression level.

[0006] Furthermore, step (1) specifically includes: Voice data and corresponding clinical scale scores of target subjects were collected in the form of follow-up at a preset time frequency during the follow-up period; acoustic feature vectors of voice data were extracted and standardized with zero mean and unit variance; a historical reference sample library was constructed based on the standardized acoustic feature vectors and their corresponding clinical scale scores.

[0007] Furthermore, the clinical scales include the Hamilton Depression Rating Scale, the Montgomery-Esenberg Depression Rating Scale, the Patient Health Questionnaire Depression Scale, the Beck Depression Rating Scale, and the Zung Self-Rating Depression Scale. The acoustic feature vectors were extracted using the open-source speech feature extraction tool OpenSMILE. The historical reference sample library is dynamically updated and expands as the number of follow-ups increases.

[0008] Furthermore, the formula for calculating the hybrid similarity potential energy is as follows: In the formula, To achieve the potential energy of mixed similarity, Let x be the standardized acoustic feature vector of the i-th historical reference sample, and let x be the acoustic feature vector to be measured. As a balance factor, For radial basis kernel functions based on Euclidean distance, It is a direction-based cosine kernel function.

[0009] Furthermore, the radial basis kernel function based on Euclidean distance is: In the formula, is the bandwidth parameter of the radial basis kernel function.

[0010] Furthermore, the direction-based cosine kernel function is: .

[0011] Furthermore, step (4) specifically includes the following sub-steps: (4.1) The mixed similarity potential is truncated by a boundary threshold to construct a sparse similarity matrix: In the formula, Let be the sparsed similarity matrix of the i-th historical reference sample. is the mixed similarity potential between the acoustic feature vector to be tested and the acoustic feature vector of the i-th historical reference sample, and m is the preset boundary threshold. (4.2) Sharpen the s-th power similarity matrix: In the formula, Let s be the similarity score of the i-th historical reference sample after power-law sharpening, where s is the sharpening factor and s≥1; (4.3) Calculate the normalized weights: In the formula, The normalized weights for the i-th historical reference sample. Let be the similarity score of the j-th historical reference sample after power-law sharpening, and N be the total number of historical reference samples. To prevent constants with a denominator of zero.

[0012] Furthermore, the formula for calculating the predicted value of the depression severity score is as follows: In the formula, This is a predictor of the severity of depression. is the clinical scale score corresponding to the i-th historical reference sample.

[0013] The beneficial effects of this invention are as follows: This invention uses a hybrid kernel function that combines Euclidean distance and cosine similarity to simultaneously capture the spatial distance and directional consistency of acoustic features, avoiding the limitations of a single metric form, and adapting to more complex speech performance; by introducing a boundary threshold m, a sparse manifold topology is constructed, which effectively reduces the interference of low-relevance samples on the prediction results and improves the robustness of the model; by using a power function sharpening mechanism, adaptive focusing on high-similarity samples is achieved, and compared with traditional linear weighting, this invention can more accurately reflect the patient's current level of depression. Attached Figure Description

[0014] Figure 1 This is a flowchart of the depression assessment method based on acoustic feature sparse manifold topology of the present invention. Detailed Implementation

[0015] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0016] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0017] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0018] The present invention will now be described in detail with reference to the accompanying drawings. Unless otherwise specified, the features of the following embodiments and implementations can be combined with each other.

[0019] See Figure 1 The depression severity assessment method based on acoustic feature sparse manifold topology of the present invention specifically includes the following steps: (1) Collect the speech data of the target object (i.e., specific patients) and the corresponding clinical scale scores according to the preset time frequency, extract the acoustic feature vector, and form a historical reference sample library.

[0020] Specifically, firstly, voice data and corresponding clinical scale scores of specific patients are collected at preset time frequencies during the follow-up period in the form of follow-up visits. The clinical scales include, but are not limited to: Hamilton Depression Rating Scale (HAMD), Montgomery-Esenberg Depression Rating Scale (MADRS), Patient Health Questionnaire Depression Scale (PHQ-9), Beck Depression Rating Scale (BDI), Zung Self-Rating Depression Scale (SDS), etc. The clinical scales assess the patient's current depressive state from multiple dimensions (characteristic dimensions such as depressive mood, self-blame, motivation, anxiety, somatization symptoms, and dullness). Then, acoustic feature extraction was performed on the collected speech data to extract acoustic feature vectors, which were then subjected to zero-mean unit variance standardization (Z-score standardization). The acoustic feature vectors were extracted using the open-source speech feature extraction tool OpenSMILE, extracting features including fundamental frequency (F0), jitter, Mel-frequency cepstral coefficients (MFCCs), and shimmer. The mean and variance used in the standardization process were the mean and variance of the historical reference sample library for that specific patient individual, thus eliminating differences in vocal baselines between different patients. Finally, a historical reference sample library D was constructed based on the standardized acoustic feature vectors and their corresponding clinical scale scores. In the formula, Let be the standardized acoustic feature vector of the i-th historical reference sample. Let be the clinical scale score corresponding to the i-th historical reference sample, and N be the total number of historical reference samples. This historical reference sample database D is dynamically updated and expands with the increase in follow-up visits, continuously refining the patient's personalized baseline.

[0021] (2) Collect the speech data of the target object (i.e., a specific patient) at the time of the test, perform the same acoustic feature extraction and standardization process, and obtain the acoustic feature vector x to be tested.

[0022] (3) Combining Euclidean distance and cosine similarity, a hybrid kernel function is used to calculate the hybrid similarity potential between the acoustic feature vector x to be tested and the acoustic feature vectors of each historical reference sample. : In the formula, The hybrid similarity potential energy has a value range of [0,1] and is used to characterize the similarity potential energy between the acoustic feature vector to be tested and the acoustic feature vector of the historical reference sample. It is a balancing factor, and its value range is [0,1]. For radial basis kernel functions based on Euclidean distance, It is a direction-based cosine kernel function.

[0023] Furthermore, the radial basis kernel function based on Euclidean distance is: In the formula, is the bandwidth parameter of the radial basis kernel function, used to control the rate at which similarity decays as the Euclidean distance between feature vectors increases.

[0024] Furthermore, the direction-based cosine kernel function maps its values ​​to the range [0,1] through a linear transformation, and its expression is: (4) Based on hybrid similarity potential We introduce boundary thresholds and sharpening mechanisms to construct sparse manifold topology and calculate normalized weights.

[0025] (4.1) Using the boundary threshold m to measure the mixing similarity potential Truncation is performed to construct a sparse similarity matrix. : In the formula, Let be the sparsed similarity matrix of the i-th historical reference sample. is the mixed similarity potential between the acoustic feature vector to be tested and the acoustic feature vector of the i-th historical reference sample, and m is a preset boundary threshold with a value range of (0,1], in order to eliminate the interference of low-relevance samples on the prediction results.

[0026] (4.2) For the sparsed similarity matrix Sharpen to the power of s: In the formula, Let s be the similarity score of the i-th historical reference sample after power-law sharpening, and let s be the sharpening factor, where s≥1, to increase the weight of high-similarity samples during normalization, making the weight of high-similarity samples more concentrated. The sharpening factor s utilizes the nonlinear amplification property of the power function to achieve an attention mechanism; as s increases, the weight distribution exhibits a peak shape, and the Nadalaya-Watson estimation model focuses more on the most similar historical state of the test sample, thereby improving prediction accuracy; when s approaches 1, the Nadalaya-Watson estimation model degenerates into a standard hard-threshold-based truncated kernel regression.

[0027] (4.3) Calculate the normalized weights: In the formula, The normalized weights for the i-th historical reference sample. Let be the similarity score of the j-th historical reference sample after power-law sharpening. To prevent constants with a denominator of zero.

[0028] (5) Calculate the predicted value of depression score based on the improved Nadalaya-Watson estimation model to realize the assessment of depression level.

[0029] Furthermore, the formula for calculating the predictive value of the depression severity score is as follows: In the formula, This provides a predicted depression level score. The depression level assessment is now complete. This invention uses a hybrid kernel function to simultaneously capture the spatial distance and directional consistency of acoustic features. By using a boundary threshold m, a sparse manifold topology is constructed, and a sharpening mechanism is employed to enhance the weight parameters of highly similar samples, thereby improving the reliability and accuracy of the depression level prediction score.

[0030] It should be noted that the weights of the original Nadalaya-Watson estimation model were given by the kernel function. In this embodiment, the kernel function is replaced by a hybrid kernel function constructed by the radial basis function and the cosine kernel function. Furthermore, truncation and sharpening mechanisms are added. This is an improvement on the Nadalaya-Watson estimation model, which eliminates noise sample interference and focuses on highly similar samples.

[0031] For example, this embodiment was validated on a clinical dataset, which included follow-up voice data from 120 patients with depression. Experimental results showed that compared with the traditional kernel regression method, the mean absolute error (MAE) of the present invention was reduced by 18.6%; compared with the method without sparsity processing, the variance of the prediction results of the method described in the present invention was reduced by 22.3%; and by introducing a sharpening factor s=2, the Nadaraja-Watson estimation model significantly improved its attention concentration on highly similar historical samples.

[0032] The predicted scores of this invention at different levels of depression show a high degree of consistency with actual clinical scale scores, verifying the effectiveness and reliability of this invention.

[0033] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for assessing depression levels based on acoustically characteristic sparse manifold topology, characterized in that, Includes the following steps: (1) Collect the speech data of the target object and the corresponding clinical scale scores according to the preset time frequency, extract the acoustic feature vector, and form a historical reference sample library; (2) Collect the speech data of the target object at the time of the test, extract the acoustic features, and obtain the acoustic feature vector to be tested; (3) Combining Euclidean distance and cosine similarity, the mixture kernel function is used to calculate the mixture similarity potential between the acoustic feature vector x to be tested and the acoustic feature vectors of each historical reference sample; (4) Based on the hybrid similarity potential, a boundary threshold and sharpening mechanism are introduced to construct a sparse manifold topology and calculate the normalized weights; (5) Calculate the predicted value of depression score based on the improved Nadalaya-Watson estimation model to realize the assessment of depression level.

2. The method for assessing depression based on sparse manifold topology of acoustic features according to claim 1, characterized in that, Step (1) specifically includes: Voice data and corresponding clinical scale scores of target subjects were collected in the form of follow-up at a preset time frequency during the follow-up period; acoustic feature vectors of voice data were extracted and standardized with zero mean and unit variance; a historical reference sample library was constructed based on the standardized acoustic feature vectors and their corresponding clinical scale scores.

3. The method for assessing depression based on acoustically characteristic sparse manifold topology according to claim 1 or 2, characterized in that, The clinical scales include the Hamilton Depression Rating Scale, the Montgomery-Essenberg Depression Rating Scale, the Patient Health Questionnaire Depression Scale, the Beck Depression Rating Scale, and the Zung Self-Rating Depression Scale. The acoustic feature vectors were extracted using the open-source speech feature extraction tool OpenSMILE. The historical reference sample library is dynamically updated and expands as the number of follow-ups increases.

4. The method for assessing depression based on sparse manifold topology of acoustic features according to claim 1, characterized in that, The formula for calculating the hybrid similarity potential energy is: In the formula, To achieve the potential energy of mixed similarity, Let be the standardized acoustic feature vector of the i-th historical reference sample, and let x be the acoustic feature vector to be measured. As a balance factor, For radial basis kernel functions based on Euclidean distance, It is a direction-based cosine kernel function.

5. The method for assessing depression based on acoustic feature sparse manifold topology according to claim 4, characterized in that, The radial basis kernel function based on Euclidean distance is: In the formula, is the bandwidth parameter of the radial basis kernel function.

6. The method for assessing depression based on sparse manifold topology of acoustic features according to claim 4, characterized in that, The direction-based cosine kernel function is: 。 7. The method for assessing depression based on acoustic feature sparse manifold topology according to claim 1, characterized in that, Step (4) specifically includes the following sub-steps: (4.1) The mixed similarity potential is truncated by a boundary threshold to construct a sparse similarity matrix: In the formula, Let be the sparsed similarity matrix of the i-th historical reference sample. is the mixed similarity potential between the acoustic feature vector to be tested and the acoustic feature vector of the i-th historical reference sample, and m is the preset boundary threshold. (4.2) Sharpen the s-th power similarity matrix: In the formula, Let s be the similarity score of the i-th historical reference sample after power-law sharpening, where s is the sharpening factor and s≥1; (4.3) Calculate the normalized weights: In the formula, The normalized weights for the i-th historical reference sample. Let be the similarity score of the j-th historical reference sample after power-law sharpening, and N be the total number of historical reference samples. To prevent constants with a denominator of zero.

8. The method for assessing depression based on acoustic feature sparse manifold topology according to claim 1, characterized in that, The formula for calculating the predicted value of the depression severity score is as follows: In the formula, This is a predictor of the severity of depression. is the clinical scale score corresponding to the i-th historical reference sample.