A method for detecting depression based on voice characteristics
By constructing depression diagnosis problems, recording audio signals and using CNN model training, the problem of low accuracy of depression detection in the existing technology is solved, and efficient and accurate depression detection and data set assembly is achieved, reducing testing time.
Patent Information
- Application Number
- CN202211474725.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-11-23
AI Technical Summary
The detection method of depression based on audio signals in the prior art has low accuracy, and traditional support vector machines and decision tree models lead to data redundancy, making it difficult to improve the specificity and sensitivity of detection.
29 depression diagnosis problems were constructed, the subject's audio signal was recorded through the microphone, and after processing using the audio noise reduction algorithm, feature extraction was performed using Open SMILE, and the CNN model was trained to select the model with the best specificity and sensitivity for depression detection.
It improves the accuracy of depression detection, reduces testing time, helps patients to test independently, relieves hospital visits, and forms data sets to help research.
Smart Images

Figure CN116364116B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to a method for detecting depression based on speech features. Background Art
[0002] Depression is a common mental illness, which is mainly manifested as a series of symptoms such as low mood, slow thinking, reduced will activity, sleep disorders, and decreased appetite. In this state, the communication between patients and normal people will show differences. Therefore, we can compare the audio signals of patients and normal people to help diagnose depression. By analyzing relevant features such as the waveform and frequency of the audio, the audio signal can be effectively extracted and analyzed, and the audio of patients and normal people can be distinguished to achieve the classification effect. Whether it is for assisting doctors in making decisions for diagnosis or helping experimenters discover the laws and features existing in the audio signal, the analysis of the audio signal is an essential stage.
[0003] The paper "Liu Z, Wang D, Zhang L, et al. A Novel Decision Tree for Depression Recognition in Speech[J]. 2020." uses audio analysis to detect depression and provides an overall idea for detecting depression based on the features of depressive audio. First, an interview questionnaire containing 29 questions is designed, and the audio signals of the subjects are recorded by means of face-to-face interviews. Each subject has 29 audio signals. After preprocessing and feature extraction operations on the audio signals, dimensionality reduction operations are performed on the extracted features. Since the number of features is too large, the purpose of the dimensionality reduction operation is to screen out the top 20 dimensions of features that have a greater impact on the classification effect for analysis. And a traditional support vector machine is used to train and test the data to obtain 29 training models. The training models are sorted according to specificity and sensitivity, and a decision tree model is used to detect depression. Therefore, the paper mainly uses the first 20-dimensional main features of the audio and a traditional classifier for training, and then uses the specificity and sensitivity of the model to further determine whether there is depression.
[0004] In the paper, the main classification method is to use the traditional support vector machine method, and the relief algorithm is used in feature extraction. However, the results obtained by these two algorithms are not very good, and the correct rate can only reach 70%. Moreover, in the design of the decision tree, the method only uses the first two better models of sensitivity and specificity at most, which may lead to a large amount of data redundancy. Therefore, how to improve this detection method has become a major problem.
[0005] Based on the paper, if the accuracy of the algorithm can be improved, such as specificity and sensitivity, and the model is appropriately modified to make the detection results more accurate, this will provide great help for the hospital to independently detect depression in the future, thereby reducing manpower and material resources. At the same time, the more data collected will also provide effective help for future research. Summary of the Invention
[0006] To overcome the deficiencies of the prior art, the present invention provides a method for detecting depression based on voice features, which records the audio information during the process of a subject answering depression diagnosis questions and stores it in a cloud library, and then extracts the key features in the audio, and detects depression through a trained model. The present invention can help patients independently detect depression, relieve the pressure on hospital visits, speed up the complex interrogation process, and at the same time be able to build a dataset of depression to provide help for depression research.
[0007] The technical solution adopted by the present invention to solve its technical problems includes the following steps:
[0008] Step 1: Construct 29 depression diagnosis questions based on the depression scale;
[0009] The 29 depression diagnosis questions include three types of questions: interview, reading, and picture description;
[0010] The interview questions include 18 questions, taken from the depression scale;
[0011] The reading questions include 1 short passage and 6 groups of words; the 6 groups of words respectively express positive emotions, neutral emotions, and negative emotions; the words expressing positive emotions and negative emotions are respectively selected from the affirmative words and negative words in the emotional ontology corpus created by Professor Lin Hongfei, and the words expressing neutral emotions are selected from the Chinese emotional word extreme value table to represent neutral words;
[0012] The picture description question is to select three pictures from the Chinese Facial Affective Picture System CFAPS, which respectively represent positive, neutral, and negative faces, and the last picture comes from the Thematic Apperception Test TAT;
[0013] Step 2: Conduct interviews with all subjects, so that each subject answers the 29 depression questions constructed in Step 1, and record the audio signal of the subject answering the questions through a microphone;
[0014] Step 3: Package and upload the recorded audio to the cloud for storage in sequence, and give each subject a unique identifier;
[0015] Step 4: Use an audio noise reduction algorithm to process all audio files to make the audio files free from noise interference;
[0016] Step 5: Extract features from all the audio using Open SMILE, and at the same time label each audio data with the corresponding label. Patients with depression are labeled as 1, while healthy individuals are labeled as 0;
[0017] Step 6: Group the features extracted from all the audio according to the questions. Each group only contains the audio of all the subjects answering one question, and each group constitutes a data set. A total of 29 data sets are formed; then divide each data set into a training set, a validation set and a test set;
[0018] Step 7: Input the 29 training sets into the CNN model for training respectively, and use the validation set and the test set for validation and testing to obtain 29 depression detection models;
[0019] Step 8: Calculate the specificity and sensitivity of each depression detection model, and select the model with the largest specificity value and the top two models with the largest sensitivity value. A total of three depression detection models are used as the final depression detection models;
[0020] The calculation formulas are as follows:
[0021]
[0022]
[0023] TP: True positive, which is the number of samples in the validation set where the subject's diagnosis result is depression and the model's prediction result is depression;
[0024] FP: False positive, which is the number of samples in the validation set where the subject's diagnosis result is healthy and the model's prediction result is depression;
[0025] FN: False negative, which is the number of samples in the validation set where the subject's diagnosis result is depression and the model's prediction result is healthy;
[0026] TN: True negative, which is the number of samples in the validation set where the subject's diagnosis result is healthy and the model's prediction result is healthy;
[0027] Step 9: Only retain the questions corresponding to the three final depression detection models as the final depression diagnosis questions to test the patients.
[0028] Preferably, the depression scale includes the Self-Rating Depression Scale (SDS), the Patient Health Questionnaire-9 (PHQ-9), the Hamilton Depression Rating Scale (HAMD), and the Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition (DSM-IV).
[0029] Preferably, the audio storage format is a WAV file.
[0030] Preferably, the ratio of the dataset division is: training set: validation set: test set = 28:10:13.
[0031] Preferably, the structure of the CNN model is as follows: First, it goes through two convolutional layers, including convolutional operations with 16 and 32 filters of size 3×1 respectively. The third layer is a max pooling layer of 2×1; the fourth and fifth layers afterwards include convolutional operations with 64 filters of size 3×1, and then a max pooling layer of 2×1; the seventh layer is a fully connected layer for the final model output.
[0032] The beneficial effects of the present invention are as follows:
[0033] The present invention mainly helps patients to detect depression independently, relieve the pressure on hospital visits, speed up the complex consultation process, and at the same time can build a dataset of depression to provide assistance for research. Multimodal data can also be added in the research to improve accuracy and efficiency; the method of the present invention not only ensures the correct rate, but also can greatly reduce the test time and complete the detection of depression efficiently and quickly. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of the method of the present invention.
[0035] Figure 2 It is a structural diagram of the CNN model of the present invention.
[0036] Figure 3 It is the software interface of the present invention Figure 1 .
[0037] Figure 4 It is the software interface of the present invention Figure 2 .
[0038] Figure 5 It is the software interface of the present invention Figure 3 . DETAILED DESCRIPTION OF THE INVENTION
[0039] The present invention will be further described below with reference to the drawings and embodiments.
[0040] In order to assist doctors in diagnosing depression and hoping to help patients diagnose depression independently through this system, the present invention proposes a depression detection software based on voice features. The so-called voice features refer to the audio information mainly collected during the interview, reading, and description processes of the subjects when they answer questions, read, and describe. The purpose of the present invention is to utilize the collected voice information, extract the voice features therein, and judge depression through a trained model, so as to help doctors accurately judge whether a patient has depression and provide certain research assistance to doctors. For testers, by analyzing the voice features of the subjects and their specific situations, the internal connection between the two can be obtained, which provides help for relevant conclusions.
[0041] As Figure 1 shown, a depression detection method based on voice features proposed by the present invention includes a data acquisition module, a data storage module, a data analysis module, and an output module.
[0042] The audio signals used in the present invention are recorded through a normal microphone, and accurate recording is carried out in each link. For example, at the beginning of answering an interview question, the microphone is turned on to record the audio, and after the subject clicks "end", the microphone is turned off to end the recording.
[0043] At the same time, in the data acquisition module, the interview questions adopt the "face-to-face" conversation method. Using the existing face animation generation technology, the face pictures and the audio of the questions are synthesized to achieve the effect of face-to-face communication. The face animation generation technology mainly includes modules such as face recognition, model training, and output. The face recognition module uses the s3fd algorithm, which is inspired by the anchor-based face detection algorithm. It classifies and regresses the detection targets through some predefined anchors, and improves the face recall rate by modifying the anchor matching strategy. Model training mainly trains the convolutional neural network and generative adversarial network of the set audio and pictures. The output module mainly uses the above-trained model, takes the face pictures and audio as inputs, and outputs a more realistic face animation video for subsequent research. This virtual intelligent doctor can relieve the pressure of doctors' consultations and at the same time reduce the psychological defenses of the subjects, collecting more real emotional data.
[0044] The data collected in this aspect mainly include the audio and video data of the subjects, which are mainly collected by a conventional computer camera and a professional microphone. The intelligent doctor communicates with the subjects about a total of 29 questions, which are designed based on a series of depression scales, such as SDS (Self-Rating Depression Scale), PHQ-9 (Patient Health Questionnaire-9), HAMD (Hamilton Depression Scale), etc. Only the relevant physiological signals of the subjects are recorded during the collection process, and the complete conversation process is not recorded. At the same time, the data collection environment needs to be kept relatively quiet, and the collection process is completed in a noise-free environment. The data currently used by the system is audio data, and the video data will be used in the subsequent multi-modal model.
[0045] In the experimental design, two factors were examined: speaking style and emotional valence. The speaking styles were three speaking modes involved in the study: interview, word reading, and picture description. Each of them had three emotional states: positive, neutral, and negative. To counteract the sequence effect, the speech orders with different emotional valences were randomly assigned. The experimental language was Chinese, and the whole experiment lasted about 25 minutes. More details about these three speaking styles are described as follows: 1) The interview had 18 questions from DSM-IV, HRSD, and other scales. 2) The reading consisted of a short passage and six groups of words, with each group having 10 common Chinese words from the Emotion Ontology Corpus and the Chinese Emotion Word Extreme Value Table. 3) The picture description included three facial expression pictures from the Chinese Facial Affect Picture System (CFAPS).
[0046] There were three parts in a fixed order: interview, reading, and picture description. The text materials were displayed on the computer screen, and the participants were required to complete the experiment according to the instructions.
[0047] 1) Interview: This task included 18 questions with positive, neutral, and negative meanings. These topics were from DSM-IV and some depression scales commonly used in this field. For example: If you had a vacation, please describe your travel plan. What was the best gift you received? How do you feel? Please describe one of your friends, including age, job, personality, and hobbies. How do you evaluate yourself? What do you want to do when you can't fall asleep? What makes you despair?
[0048] 2) Reading: This part consisted of a short story named "The North Wind and the Sun", which was taken from the booklet "Principles of the International Phonetic Association" and was often used in acoustic analysis in international multilingual clinical research. Six groups of words expressed positive emotions, neutral emotions, and negative emotions respectively. Affirmative and negative words were selected from the Emotion Ontology Corpus created by Professor Lin Hongfei, and neutral words were selected from the Chinese Emotion Word Extreme Value Table. All these words were common words in Chinese to avoid the influence of education level, and the number of strokes of the six groups of words was similar. The subjects were required to read a story, and the words appeared in their usual way.
[0049] 3) Picture description: The materials for this task include a total of four pictures. Three pictures were selected from the Chinese Facial Affective Picture System (CFAPS), representing positive, neutral, and negative faces respectively. The last picture is from the Thematic Apperception Test (TAT). The TAT was created by Murray in 1935. The Thematic Apperception Test is a projective personality test. Among many psychological test theories, the psychological projection theory is one of the most widely applied, and the psychological projection techniques derived from it are often used in personality measurement. In this task, the subjects were required to freely describe these four pictures.
[0050] In the experiment, the 29 records of each subject were named from 1 to 29 in a determined order. Specifically as follows: The positive, neutral, and negative interview records were named 1 - 6, 7 - 12, and 13 - 18 respectively. The record of the short story was named 19. The reading of the six - word groups was named 20 - 21, 22 - 23, and 24 - 25 in the order of positive, neutral, and negative emotions. 26 - 28 are the picture descriptions in the same order as the reading part. The record number of the Thematic Apperception Test (TAT) is 29.
[0051] In the data storage module, the client needs to send a data transmission request to the cloud server. After the request is passed, the entire recorded audio is packaged and uploaded to the cloud for storage in order, and a unique identifier is given to each subject. The audio storage format is a WAV file and can be directly played using audio playback software.
[0052] In the data analysis module, first, the cloud server needs to process all audio files using an audio noise reduction algorithm to ensure that the audio files used are free from interference by other noises. Secondly, the audio data needs to be feature - extracted and selected. Finally, a trained model is used to judge the mental state of the subjects.
[0053] In terms of model training, first, all the audio in the dataset is used to extract features using the Open SMILE software in order. Open SMILE is a tool for speech feature extraction and has a relatively wide range of applications in fields such as speech recognition (feature extraction, keyword recognition, etc.), affective computing (emotion recognition, sensitive virtual agents, etc.), and music information retrieval (chord annotation, beat tracking, onset detection, etc.). This tool can currently calculate up to 6000 audio features, such as frame energy, frame intensity, Mel - frequency cepstral coefficients, and a series of other speech features, and can identify the subtle differences in emotions to achieve the purpose of emotion analysis. At the same time, the tool also includes the open CV library, which can be used for video processing and feature extraction.
[0054] Using the Open SMILE software, speech features of over 1000 dimensions can be obtained. Therefore, it is necessary to screen the main features for subsequent analysis and utilization. The purpose of feature selection is to eliminate irrelevant feature data and reduce redundant features, thereby reducing the number of features, improving the model accuracy, and reducing the running time. At the same time, there are various feature selection methods, so it is necessary to find a suitable method for feature selection and processing. Before feature selection, corresponding labels need to be assigned to each audio data. For example, patients with depression are labeled as 1, while normal people are labeled as 0. After feature selection is completed, the processed data is used as the input of the model for model training.
[0055] Dataset division: Training set: Validation set: Test set = 28:10:13.
[0056] The training set is mainly used to train the model, the validation set is mainly used to evaluate the accuracy, sensitivity, and specificity of the model, and the test set is mainly used to evaluate the accuracy of the entire final model.
[0057] The audio data after feature selection is used as the input of the model. As Figure 2 shown, it first passes through two convolutional layers, including convolutional operations with 16 and 32 filters of size 3×1 respectively. The third layer is a max-pooling layer of 2×1. The subsequent fourth and fifth layers include convolutional operations with 64 filters of size 3×1, followed by a max-pooling layer of 2×1. The seventh layer is a fully connected layer for the final model classification output.
[0058] Among the 29 trained models, select the model with the best performance in specificity and the top two models with the best performance in sensitivity, a total of three depression detection models as the final depression detection models for the final data processing and analysis of the system. The model specificity and sensitivity need to be calculated according to the confusion matrix as follows:
[0059]
[0060] The calculation formulas are as follows:
[0061]
[0062]
[0063] Determination of the four values in the specificity and sensitivity calculation formulas: (All four data are calculated on the validation set)
[0064] TP: True positive, the number of samples in the validation set where the professional doctor's diagnosis result is depression and the model prediction result is depression;
[0065] FP: False positive, the number of samples in the validation set where the professional physician's diagnosis result is healthy, but the model's prediction result is depression;
[0066] FN: False negative, the number of samples in the validation set where the professional physician's diagnosis result is depression, but the model's prediction result is healthy;
[0067] TN: True negative, the number of samples in the validation set where the professional physician's diagnosis result is healthy and the model's prediction result is also healthy;
[0068] Therefore, the final system only needs the subject to communicate three questions with the intelligent doctor to predict the mental state. This method not only ensures the accuracy rate, but also can greatly reduce the testing time and complete the detection of depression efficiently and quickly.
[0069] In the system output module, after the cloud server completes data processing and analysis, it sends a data transmission request to the client. The subject waits for a while. After the request is approved, the client will display the received result on the system interface to complete the result output process. Specific embodiments:
[0071] The purpose of the present invention is to use a computer program to utilize the subject's audio and other related physiological signals to help patients independently detect depression, relieve the hospital's medical pressure, and speed up the complex medical consultation process. At the same time, it can build a dataset of depression to provide help for future research, and multi-modal data can also be added in subsequent research to improve the accuracy rate and efficiency.
[0072] The present invention is divided into two major parts. One is the model training data acquisition tool, and the other is the depression detection software based on voice features.
[0073] The data acquisition tool includes three modules: the data acquisition and storage module and the data analysis module. The data acquisition tool is only for system developers to use.
[0074] In the data acquisition and storage module, first, the subject needs to click the "Start" button on the interface, call the explain() function and ask the subject to read the software usage notes. After clicking "OK", call the information() function and ask the subject to enter relevant information (name, gender, age). Click the "Confirm" button to call the strat() function. After saving the information, the conversation can start. The system will automatically assign a number to the subject using the mkdir method and create a corresponding folder. For example, the number of the first depression patient is "A0001", and the number of the first healthy subject is "B0001". The subject's information will be saved in the data_information.xlsx file. After starting, call the player method of the VLC library to display the intelligent doctor synthesized by audio and image to ask questions. During the conversation, 29 questions will be asked by the intelligent doctor. After each question is asked, the subject clicks the "Start" button at the bottom of the interface to call the my_record() function to start answering. After answering, click the "End" button. After each question is answered, call the save_file() function to save the audio. The audio file will be named according to the question number and saved in the corresponding folder. For example, after the first depression patient answers the first question, the audio file "01.wav" will be saved in the folder "A0001". After each question is answered, call the rest() function, and there will be a five-second break. After the break, call the start() function again to enter the next question. After 29 questions are answered, call the complete() function, and the system will prompt "Data acquisition has been completed". After clicking the "OK" button, return to the main interface, and the data acquisition process is completed. After the acquisition is completed, the attending doctor needs to fill in the subject's status in the data_information.xlsx file. For example, fill in 0 for healthy and 1 for depression.
[0075] In the data analysis module, first, the openSMILE tool is used to call the FeatureExtraction() function to configure the emobase2010.conf file in the toolkit, extract audio features one by one, and save them as the feature file audio_feature.txt. Then, the relief() function is used to assign different weights to the features according to the correlation between each feature and the category. The top 20 features with larger extracted feature weights are used as the input of the classification model. After training the model, it is sorted according to the specificity and sensitivity of the model. Specificity represents the probability of correctly judging a non-patient, and sensitivity represents the probability of correctly judging a patient. Specificity and sensitivity are obtained from formulas (1) and (2). The present invention selects two models with the best sensitivity performance and one model with good specificity performance for the final depression prediction software. The corresponding software only requires the subject to answer three questions. While maintaining an 80% correct rate, the detection time is greatly shortened.
[0076] As Figures 3 - 5 shown, the depression detection software based on voice features includes four modules: a data acquisition module, a data storage module, a data analysis module, and an output module.
[0077] In the data acquisition module, click "Start" without entering relevant information. After starting, call the player method of the VLC library to display the intelligent doctor synthesized by audio and image to ask questions. After each question is asked, the subject clicks the "Start" button to call the my_record() function to start answering and start recording audio. After answering, the subject clicks the "End" button to complete the audio recording. After answering, the rest() function will be called, and there will be a five-second break. After the break, call the start() function again to enter the next question. When all three questions are answered, the data acquisition process is completed.
[0078] In the data storage module, after each question is recorded, the system calls the request() function to send a data transmission request to the cloud server. After the server receives the data, it stores the data in a specified folder. For example, the first audio of the first depression patient is saved as the "01.wav" file in the "A0001" folder.
[0079] In the data analysis module, using the open SMILE tool, call the FeatureExtraction() function to configure the emobase2010.conf file in the toolkit, extract audio features one by one, save them as the feature file audio_feature.txt, and call the readFile() function to read the main features of the txt file and save them in the feature_data array. Use the feature_data array as the input, call the read_model() function to read the three trained models, and use the feature array as the input to predict the results. The results are obtained by voting.
[0080] Finally, in the output module, the cloud server calls the request() function to send the result data to the client system. After the system receives the result data, it displays it on the screen to complete the output.
Claims
1. A method for detecting depression based on speech features, characterized in that, It includes the following steps: Step 1: Based on the depression scale, construct 29 depression diagnosis questions; The 29 depression diagnosis questions include three types of questions: interview, reading, and picture description; The interview questions include 18 questions, which are taken from the depression scale; The reading questions include 1 short passage and 6 groups of words; the 6 groups of words express positive emotions, neutral emotions, and negative emotions respectively; the words expressing positive emotions and negative emotions are selected as affirmative words and negative words respectively from the emotional ontology corpus created by Professor Lin Hongfei, and the words expressing neutral emotions are selected as neutral words from the Chinese emotional word extreme value table; The picture description question is to select three pictures from the Chinese Facial Affective Picture System CFAPS, which represent positive, neutral, and negative faces respectively, and the last picture comes from the Thematic Apperception Test TAT; Step 2: Conduct interviews with all subjects, so that each subject answers the 29 depression questions constructed in Step 1, and record the audio signal of the subject answering the questions through a microphone; Step 3: Package and upload the recorded audio to the cloud for storage in sequence, and give each subject a unique identifier; Step 4: Use an audio noise reduction algorithm to process all audio files to make the audio files free from noise interference; Step 5: Extract features from all audio using Open SMILE, and at the same time label each audio data with the corresponding label, label the depression patients as 1, and the healthy people as 0; Step 6: Group the features extracted from all audio according to the questions. Each group only contains the audio of all subjects answering one question, and each group constitutes a data set, and a total of 29 data sets are formed; then divide each data set into a training set, a validation set, and a test set; Step 7: Input the 29 training sets into the CNN model for training respectively, and use the validation set and the test set for validation and testing to obtain 29 depression detection models; Step 8: Calculate the specificity and sensitivity of each depression detection model, and select the model with the largest specificity value and the top two models with the largest sensitivity value, a total of three depression detection models as the final depression detection models; The calculation formula is as follows: TP: True positive, the number of samples in the validation set where the subject's diagnosis result is depression and the model's prediction result is depression; FP: False positive, the number of samples in the validation set where the subject's diagnosis result is healthy and the model's prediction result is depression; FN: False negative, the number of samples in the validation set where the subject's diagnosis result is depression and the model's prediction result is healthy; TN: True negative, the number of samples in the validation set where the subject's diagnosis result is healthy and the model's prediction result is healthy; Step 9: Only retain the questions corresponding to the three final depression detection models as the final depression diagnosis questions to test the patients.
2. The method for detecting depression based on voice features according to claim 1, wherein The depression scale includes the Self-Rating Depression Scale SDS, the Patient Health Questionnaire-9 PHQ-9, the Hamilton Depression Scale HAMD, and the Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition DSM-IV.
3. The method for detecting depression based on voice features according to claim 1, wherein The audio storage format is a WAV file.
4. A method for detecting depression based on voice features according to claim 1, characterized in that, The ratio of the dataset division is: training set: validation set: test set = 28:10:
13.
5. A method for detecting depression based on voice features according to claim 1, characterized in that, The structure of the CNN model is as follows: First, it goes through two convolutional layers, which respectively include convolutional operations with 16 and 32 filters of size 3×1. The third layer is a max pooling layer of 2×1; the fourth and fifth layers afterwards include convolutional operations with 64 filters of size 3×1, followed by a max pooling layer of 2×1; the seventh layer is a fully connected layer for the final model output.
Citation Information
Patent Citations
Depression automatic evaluation system and method based on phonetic features and machine learning
CN106725532A
KR20200092166A