Voice depression detection system based on deep reinforcement learning
By introducing Q-learning and an adaptive reward function into the voice depression detection system, combined with 1D convolution and BiLSTM, the data imbalance problem was solved, the detection accuracy and robustness were improved, and better depression detection results were achieved.
Patent Information
- Application Number
- CN202511261562.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-12-12
AI Technical Summary
Existing voice depression detection systems tend to learn features of the majority class when faced with data imbalance, leading to a decline in detection performance. Furthermore, traditional methods may introduce irrelevant information or increase the risk of overfitting.
We introduce a deep reinforcement learning method based on Q-learning, which focuses on minority class samples through an adaptive reward function. We combine 1D convolution and bidirectional long short-term memory network (BiLSTM) to extract deep depression-related features in speech, construct a depression detection agent, and train it using the interaction mechanism of reinforcement learning.
It improves the accuracy and robustness of voice depression detection, mitigates the impact of class imbalance, and enhances the model's generalization ability and detection performance.
Smart Images

Figure CN121122333A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a voice depression detection system based on deep reinforcement learning, belonging to the field of voice depression detection technology. Background Technology
[0002] Depression is a common mental illness affecting approximately 280 million people worldwide, and in severe cases, it can lead to suicide. Although effective treatments exist for mild, moderate, and severe depression, in low- and middle-income countries, over 75% of patients do not receive treatment due to a lack of professional mental health professionals and other factors. Therefore, monitoring individual mental health and detecting depressive symptoms as early as possible is crucial. Since the process of speech production can be influenced by cognitive and physiological factors caused by depression, acoustic features are an effective diagnostic tool for depression. Furthermore, communicating with patients non-contactly using only a microphone improves accessibility and protects patient privacy. Therefore, speech-based diagnostic technologies for depression are gaining popularity and can be combined with current diagnostic techniques to improve diagnostic accuracy.
[0003] In recent years, the development of deep learning has provided a reliable method for speech-based depression detection. However, training deep depression detection models requires a large amount of diverse data, while depression datasets are often small due to privacy concerns. Furthermore, the low frequency of depression in the general population makes data imbalance a challenge. Common methods for addressing data imbalance fall into two categories: data-level oversampling / undersampling and algorithmic methods that use cost-sensitive methods to reset class weights in the loss function. However, these methods may have limitations, such as introducing irrelevant information or increasing the risk of overfitting. Therefore, introducing a custom reward function using reinforcement learning becomes a viable approach, as its learning mechanism and specific reward function can easily focus more on the minority classes by assigning higher rewards or penalties.
[0004] This system introduces a deep reinforcement learning framework based on Q-learning into the speech depression detection task. By using an adaptive reward function, it focuses more on minority class samples during training, learning key depression-related features in speech, thereby mitigating the impact of class imbalance in speech depression detection. Reinforcement learning is a method of learning through interaction, aiming to train an agent to continuously learn knowledge by interacting with the environment and receiving rewards or penalties, thus becoming more adaptable to the environment. Existing research has shown that in classification tasks, deep reinforcement learning can eliminate noisy data and learn better features, thereby improving classification performance. Therefore, the speech depression detection task can be viewed as a speech depression classification problem under Markov decision-making, using a reinforcement learning framework to train an agent capable of accurately detecting depression. Summary of the Invention
[0005] This invention provides a speech depression detection system based on deep reinforcement learning. In practical applications, speech depression detection is often plagued by class imbalance, which causes the model to tend to learn features of the majority class during training, thus reducing detection performance. To address this problem, this invention introduces a deep reinforcement learning method based on Q-learning. By utilizing the interactive mechanism and adaptive reward function of reinforcement learning, the system can give greater importance to minority class samples, thereby more effectively learning depression-related features in speech and improving the accuracy and robustness of depression detection.
[0006] The usage process of the speech depression detection system based on deep reinforcement learning provided by this invention includes the following steps: First, low-level acoustic features are extracted from the speech data. Then, statistical functions are used to obtain 88-dimensional high-level acoustic features for each speech segment. Next, the extracted high-level acoustic features are concatenated with a fixed number of speech segments from a subject to form a two-dimensional acoustic feature sample. Simultaneously, the subject's depression label serves as the label for the two-dimensional acoustic feature sample. These two elements together constitute the basic components of the reinforcement learning environment. In this environment, the speech depression detection task is treated as a classification problem under Markov decision, where states, actions, adaptive rewards, and policies are explicitly defined. Next, a deep network consisting of 1D convolutions and a bidirectional long short-term memory (BiLSTM) network is constructed and initialized to approximate the Q-value. The depression detection agent interacts with the environment, obtaining the quintuple for each state and storing it in replay memory. During training, samples are randomly selected from the replay memory to calculate the expected and actual reward values, and the network is updated using temporal difference loss. After training, acoustic features are extracted from the data of the subjects to be tested and two-dimensional acoustic feature samples are constructed. These samples are then fed into the intelligent agent for testing, and the results are used for majority voting to obtain the final individual-level depression category. The technical solution adopted in this invention can be further refined. The process of extracting high-level acoustic features from speech, as mentioned above, is specifically as follows: First, low-level acoustic features were extracted from the raw speech data of the j-th sentence of the i-th subject, including 25 low-level acoustic features such as frequency-related parameters, energy / amplitude-related parameters, and spectral parameters. Next, all low-level acoustic features were smoothed over time using a 3-frame-long symmetric moving average filter. Then, the arithmetic mean and standard deviation were calculated for each low-level acoustic feature, resulting in 50 high-level acoustic features. In addition, loudness and pitch were calculated separately for different low-level acoustic features, yielding an additional 27 high-level acoustic features. Furthermore, the alpha ratio of unvoiced segments, the Hammarberg index, the arithmetic mean of the spectral slopes of 0-500Hz and 500-1500Hz, the equivalent sound level, and 6 other time-related features were considered, totaling 88 parameters, constituting the high-level acoustic features of the speech. The aforementioned construction of the environment in reinforcement learning specifically involves concatenating the extracted high-level acoustic feature vectors from a subject according to a fixed number of speech segments to form a sample. Where n i Let represent the total number of individual-level two-dimensional acoustic feature samples for the i-th subject, m represent the number of speech segments, and 88 be the dimension of the high-level acoustic feature vector. Then, the depression labels of all subjects are considered as labels on the two-dimensional acoustic feature sample data. The two-dimensional acoustic feature sample data and its labels constitute the environment in reinforcement learning, where the samples are... Depression is labeled as n represents the total number of subjects. 7. The aforementioned view of the voice depression detection task as a classification problem under Markov decision-making is as follows: Markov decision-making defines the next state as only related to the current state and independent of previous states. The following details how to describe the voice depression detection task as a Markov decision problem. The four elements defining a Markov decision problem are: the samples in the environment are defined as the state space S = {s1, s2, ..., s...}. t}, state s t It is a sample randomly selected from the sample space at time t, which is the two-dimensional acoustic feature sample s. t =X t ∈X, each time step t corresponds to one sample; the set of depression labels is defined as the action space A={0,1}, a t Represents the state s at time t. t The corresponding labels are depressed or non-depressed; the reward function R is an adaptive reward function constructed to solve the class imbalance problem. Among them, s t ,at ,l t Let D represent the state, action, and sample label at time t, respectively. min D represents the minority class sample set. max Let represent the majority class sample set, ρ be the reward adjustment rate, and λ be a penalty factor that controls the penalty value during the initial training phase. The policy function is the basis for the depression detection agent to select actions based on the Q-value, denoted by π(a...). t |s t Let s represent the state s of the agent at time t. t Choose action a t The strategy used in this system is the epsilon greedy strategy: Where q is a random number between (0,1), ε is a value that gradually decreases with the time step, and Q(s) t ,a) is the Q-value function corresponding to each action calculated by the Q-network at time t. The aforementioned network construction is a Q-network for approximating the Q-value, consisting of 1D convolutions and a bidirectional long short-term memory (BiLSTM) network. The 1D convolutional layers extract short-order temporal context relationships from high-level acoustic features, and after two-dimensional max pooling, obtain deep representations. These are then further transformed by convolution to finally obtain the deep embedding representation F. f ∈R n×88 Then, the deep embedding representation F is... f The data is cascaded with high-level acoustic features and then fed into a bidirectional long short-term memory (BiLSTM) network to compute long-order contextual relations, resulting in a deep representation P∈R containing depressive information. 1×H Where H represents twice the dimension of the hidden layer in a Bidirectional Long Short-Term Memory (BiLSTM) network. The Q-value is then calculated using the representation P through a fully connected layer. The calculation process for the deep representation P is as follows: P t =BiLSTM(Concatenate(Conv(s)) t ),s t )) After the network is built, its parameters need to be initialized as the current network θ in Double DQN, and a copy of the model is made as the target network θ′ for subsequent training and updates. Next, parameters in the reinforcement learning framework, such as replay memory and episodes, also need to be initialized. Additionally, the ratio of the minority class to the majority class samples in the sample space is defined as the initial reward adjustment rate. The algorithm guides the calculation of rewards for the interaction between the agent and the environment in the first round of depression detection. The process by which the depression detection agent calculates its Q-value, makes decisions, and interacts with the environment, as mentioned above, is as follows: This deep reinforcement learning framework uses a Dueling DQN structure to calculate the Q-value. Therefore, after the Q-network, composed of 1D convolutions and a bidirectional long short-term memory network (BiLSTM), extracts the deep representation P of depression, two fully connected layers are added to the end of the network, representing the state value V(s) and the advantage value A(s,a) corresponding to each action, respectively. The values obtained from these two fully connected layers are used to calculate the final Q-value to reduce unnecessary exploration by the agent. The formula for calculating the Q-value is as follows: Subsequently, the depression detection AI selects an action based on the previously defined epsilon greedy policy function and determines the current state s based on the termination function. t Will this cause the current round to end? After determining this, the AI will set the quintuple (s) of the current time t. t ,a t ,r t ,s t+1 ,terminal t Stored in replay memory for subsequent updates. The termination function is defined as follows: The aforementioned update process for training the speech depression detection agent is as follows: During training, the depression detection agent randomly samples the replay memory at fixed time step intervals to obtain a batch (s) b ,a b ,r b ,s b+1 ,terminal b The network is updated using a quintuple, employing a Double DQN structure, where the state s is calculated from the current network θ. b The expected reward is used as the basis for the selection of the next state s by the target network θ′ based on the action that maximizes the value of the next state calculated by the current network θ. b+1 The reward, and compared with the current reward value r b The sum is the actual reward. If the terminal in the quintuple... b =true, the actual reward is the reward value r b The difference between the actual reward and the expected reward is used to update the parameters of the current network, and after a fixed time step I, the parameters of the current network are directly assigned to the target network to reduce the correlation between networks. The formula for calculating the loss is as follows: The aforementioned process trains the depression detection agent according to the order in the sample space. However, when the termination condition of the termination function is triggered, the round ends, and the minority class D in memory is recalculated and replayed. rmin With the majority class D rmax The proportion of the sample serves as a new reward adjustment rate. This guides the reward settings for the next round, and the sample space is shuffled to start the training for the next round, until the number of rounds (episodes) ends.
[0007] The beneficial effects of this invention are as follows: Compared to using sampling or cost-sensitive methods to address the data imbalance problem in speech depression detection datasets, the proposed deep reinforcement learning-based speech depression detection system can assign higher reward or penalty values by formulating a specific reward function, thus focusing more on the minority category. This allows for full utilization of the information in the sample data without introducing irrelevant information. Furthermore, the combination of 1D convolution and bidirectional long short-term memory (BiLSTM) networks is used to further extract deeper depression-related features from the speech, improving the accuracy of speech depression detection and the model's generalization ability. Attached Figure Description
[0008] Figure 1 This is a flowchart of a speech depression detection system based on deep reinforcement learning.
[0009] Figure 2 This is a diagram of a Q-network model based on 1D convolution and bidirectional long short-term memory (BiLSTM). Detailed Implementation
[0010] To more clearly describe the content of this invention, further explanation is provided below with reference to examples and accompanying drawings. The embodiments described below are not intended to limit the scope of this invention. A speech depression detection system based on deep reinforcement learning according to this invention includes the following steps:
[0011] Step 1: First, preprocess the speech in the dataset to extract high-level acoustic features eGeMAPS. Then, construct individual-level two-dimensional acoustic feature samples according to a fixed number of speech segments, and use the subject's depression label as the label of the corresponding two-dimensional acoustic feature sample. Together, they form the reinforcement learning environment. The specific steps are as follows:
[0012] 1.1 For a speech segment S∈R 1×TThe calculation includes 25 low-level descriptors, such as frequency-related parameters, energy / amplitude-related parameters, and spectral parameters. Next, all low-level descriptors are smoothed over time using a 3-frame-long symmetric moving average filter. Then, the arithmetic mean and standard deviation are calculated for each low-level descriptor, resulting in 50 high-level statistical functions. Furthermore, loudness and pitch are calculated separately for different low-level descriptors, yielding an additional 27 high-level statistical functions. Above these, the alpha ratio of unvoiced segments, the Hammarberg exponent, the arithmetic mean of the spectral slopes from 0-500Hz and 500-1500Hz, the equivalent sound level, and six other time-related features are considered, totaling 88 parameters that constitute the high-level acoustic features of the speech, eGeMAPS samples S∈R. 1×88 .
[0013] 1.2 Assuming that the number of speech segments for each participant in the dataset is the same, for example, interview recordings with fixed questions, then the high-level acoustic features (eGeMAPS) of all speech segments for each participant are concatenated to form a two-dimensional acoustic feature sample X∈R. m×88 'm' represents the number of speech segments. If the number of speech segments for each participant in the dataset is different, such as in a single interview recording, the entire recording needs to be divided into many speech segments according to timestamps. Then, the high-level acoustic features (eGeMAPS) of each speech segment are extracted, and an appropriate number of speech segments is selected for concatenation. The resulting two-dimensional acoustic feature sample X represents the speech information of a participant, and therefore the label is the participant's depression label. Together, they constitute the environment in reinforcement learning.
[0014] Step 2: Analogize the voice depression detection task to a Markov decision problem and define the elements in the reinforcement learning framework. The specific steps are as follows:
[0015] 2.1 Definition of State Space S: In the depression detection problem, the state space is the set of all acoustic feature samples, where each sample represents a state s. During the initial training phase, the state space is randomly shuffled, and a state is randomly selected as the initial state for this training. At each time step t, there is a state s. t With sample x t,t Correspondingly, when the current state transitions to the next state, the change in state follows the change in the sample.
[0016] 2.2 Definition of Action Space A: The agent's action 'a' is the label for depression detection; therefore, for state 's'... t Action a to be performed t ∈{0,1}, which is a binary classification problem, where 1 indicates that the category of the sample is depression.
[0017] 2.3 Reward Function R Definition: An adaptively adjusted reward function was designed. For minority class samples, a larger reward or penalty is given when the depression detection agent predicts correctly or incorrectly; while for majority class samples, a reward adjustment rate is applied. Furthermore, an additional term is added to the reward function: in the initial stage of training, misclassified samples receive a larger penalty to encourage the agent to adjust more quickly. This penalty term gradually decreases as the time step increases, so it is represented by λ multiplied by the reciprocal of the time step. The reward function R is as follows:
[0018] 2.4 Definition of policy function P: Policy function π(a t |s t The Q-value is the basis for the agent to select actions, representing the agent's state s at time t. t Choose action a t The strategy used in this system is the epsilon greedy strategy: Where q is a random number between (0,1), ε is a value that gradually decreases with the time step, and Q(s) t ,a) is the Q-value function corresponding to each action calculated by the Q-network at time t.
[0019] Step 3: Construct and initialize the Q-network in reinforcement learning, which, together with the policy function, constitutes the agent for speech depression detection, interacting with the environment to learn. Simultaneously, initialize other parameters in the reinforcement learning framework. The specific steps are as follows:
[0020] 3.1 The deep Q-network with approximate Q-value consists of two parts: 1D convolution and a bidirectional long short-term memory (BiLSTM) network. The 1D convolution includes two convolutional layers and one two-dimensional max-pooling layer. The first convolutional layer has 64 channels, a kernel size of 3, and a stride of 1, followed by a ReLU activation function. The max-pooling layer has a 3x3 pooling window and a stride of 1. The second convolutional layer has 88 channels, a kernel size of 3, and a stride of 1, used to align with the feature dimensions of the original high-level acoustic features eGeMAPS. After this 1D convolution extracts the deep speech depression information embedding, it is concatenated with the original two-dimensional acoustic feature samples and then fed into the BiLSTM network. The BiLSTM network has two layers, a hidden layer size of 128 dimensions, and a dropout of 0.1. The state s at time t... t Deep depressive characteristics P were extracted using 1D convolution and a bidirectional long short-term memory network (BiLSTM). tThen, the Q-values corresponding to each classification action are calculated through a fully connected layer.
[0021] 3.2 After the Q network is built, the network parameters θ are initialized as the current network, and the network structure and parameters are copied into another network θ′ as the target network for subsequent updates of network parameters using the Double DQN structure. Furthermore, an initial reward adjustment rate needs to be defined based on the ratio of the minority class to the majority class in the training data sample space. The size of the replay memory is initialized to M = 100000, and the number of episodes in reinforcement learning is also initialized.
[0022] Step 4: Begin a round by shuffling the state space S and randomly selecting a state value as the initial state. Train the depression detection agent and update the network parameters synchronously. The specific steps are as follows:
[0023] 4.1 The training process of the depression detection agent is as follows: The state s t After passing through a Q-network with a Dueling structure, the state value V(s) and the advantage value A(s,a) for each action are calculated using two fully connected layers. The final Q-value is then calculated using these two values according to the following formula: Subsequently, the depression detection agent follows the epsilon greedy policy function π(a t |s t According to the current state s t The corresponding Q-value selection action a t It interacts with the environment to obtain an adaptive reward value r. t At the same time, the current state s is determined based on the termination function. t Will this cause the current round to terminate? Then, the quintuple (s) of the current time t will be... t ,a t ,r t ,s t+1 ,terminal t Store it in the replay memory. The definition of the termination function is:
[0024] 4.2 The network update process is as follows: At the start of the update, after training at each time step, a batch of 16 data points (s) different from the current time t are randomly selected from the replay memory. b ,a b ,r b ,s b+1 ,terminal b The quintuple data is used for network updates. In this quintuple, 'a'... b,r b ,terminal b Represents state s b The actions to be performed, the reward value obtained, and whether to terminate are determined. First, the current network θ determines the actions based on the state s. b and action a b Calculate the expected reward value Q(s) b ,a b ;θ), calculate the next state s b+1 The action a that maximizes Q value b+1 =argmax a′ Q(s b+1 ,a′;θ);and the action a b+1 The actual reward value for the next state is calculated as the action of the target network, multiplied by a discount factor γ = 0.2, and then compared with the reward value r. b Adding them together gives the state s b The actual reward value r b +γQ(s b+1 argmax a′ Q(s b+1 ,a′;θ);θ′). The difference between the actual reward and the expected reward is used as the basis for updating the network parameters. Among them, if terminal b =true, representing state s b If the current round is terminated, then state s b The corresponding actual reward value is the actual reward r obtained. b The loss function for updating the parameters of the current network θ is as follows: The parameters of the target network θ′ are copied from the current network θ after I = 1000 steps.
[0025] Step 5: Repeat step 4 until the terminal. t =true; After each round, calculate the ratio of minority class to majority class samples in the replay memory as the new reward adjustment rate. This guides the learning for the next round, and returns to step 4 until the number of rounds (episodes) ends.
[0026] Step 6: After the depression detection agent is trained, the samples in the test set are processed into two-dimensional acoustic feature samples in the same way and fed into the depression detection agent for depression detection. Specifically, the ε parameter in the epsilon greedy strategy is set to 0, meaning only the category that obtains the maximum Q value is selected.
[0027] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention. Experimental Design
[0028] Experimental Dataset and Split Introduction: The dataset used in the experiment is the Distress Analysis Interview Corpus, Wizard of Oz (DAIC-WoZ), the official dataset designated by the AVEC2016 Depression Detection Challenge. This dataset is a publicly available English depression dataset containing audio and text transcriptions of 189 participants. Each audio file ranges in length from 7 to 33 minutes, with a fixed sampling rate of 16000Hz. Participants were labeled using PHQ-8 scores based on their questionnaire responses; participants with scores greater than or equal to 10 were considered to have depression. The DAIC-WoZ dataset includes a training set of 107 participants (30 depressed, 77 non-depressed) and a development set of 35 participants (12 depressed, 25 non-depressed). We followed the same random sampling strategy as DepAudioNet, using the training set for training and the development set for testing. Accuracy, precision, recall, the harmonic mean of precision and recall (F1 score), and the area under the ROC curve (AUC) are used as evaluation metrics for the model. Experimental results
[0029] Experimental Results under Different Voice Depression Detection Networks: The voice depression detection method based on deep reinforcement learning proposed in this invention was compared with other voice depression detection methods, and the results are shown in Table 1. Compared with existing advanced voice depression detection methods such as SVM, RF, DepAudioNet, CNN-AE, DEPA, Mfcc-LSTM, and ConvbiLSTM, the proposed method shows varying degrees of improvement in precision, recall, and F1 score. Compared with DepAudioNet and ConvbiLSTM methods, which also use convolutional and long short-term memory networks, the proposed method improves the F1 score by 12%, demonstrating that the proposed method is more advanced than traditional temporal models based on convolutional and long short-term memory networks. Compared to the D3QN reinforcement learning framework which uses a fixed reward function (a reward of 1 for a correct prediction and -1 for an error), the method of this invention improves precision by 6.1%, F1 score by 4.5%, and AUC by 4.3%, demonstrating that the adaptively adjusted reward function proposed in this method is helpful for improving depression detection performance for minority class samples (non-depressive classes) and the overall system. Table 1. Indicator values of the method of the present invention and various comparison algorithms on the DAIC-WoZ dataset.
Claims
1. A deep reinforcement learning based speech depression detection system, characterized in that, The method comprises the following steps: S1 extracts low-level acoustic features from the jth segment of raw speech data of the ith subject, and generates the high-level acoustic feature vector of the segment by statistical method, denoted as d i,j ∈R 88 wherein 88 represents the dimension of the high-level acoustic feature vector; S2 constructs an individual-level two-dimensional acoustic feature sample for the i-th subject by splicing together a fixed number of speech segments m based on the advanced acoustic features of the i-th subject. Where n i Let represent the total number of individual-level two-dimensional acoustic feature samples for the i-th subject, and m represent the number of speech segments. Simultaneously, the depression labels of all subjects are used as labels for the two-dimensional acoustic feature samples, which, together with the two-dimensional acoustic feature samples, constitute the environment in reinforcement learning. The sample is... Depression is labeled as Where n represents the total number of subjects. S3 treats the depression detection task as a Markov decision problem and defines the following elements: the samples in the environment are defined as the state space S = {s1, s2,..., s t}, at each time step t, s t = X t ∈ X, the set of depression labels is defined as the action space A = {0, 1}; an adaptive reward function R is designed and an epsilon-greedy policy is used as the basis for action selection; S4 setting the size M of the replay memory in reinforcement learning, the update interval I, and the number of training episodes episo des; At the same time, define the number of most class samples D in the environment max and the number of minority class samples D min , used to initialize the reward adjustment rate in the adaptive reward function R S5 initializing a deep network constructed by a 1D convolution and a bidirectional long short-term memory network as a Q network in the depression detection agent; S6 shuffling the state space S and randomly selecting a state as the initial state in the current episode; S7 will state s t input into the Q network to calculate the Q value, and according to the Q value and epsilon greedy strategy to decide the action a t under the current state s t ; after that, the action a t and the corresponding sample label l t in the environment are compared, the reward value r t is obtained according to the adaptive reward function R, and the next state value s t+1 is obtained from the state space S; at the same time, according to the time step t and the state s t , action a t and label l t , define the termination variable terminal t , and store the five-tuple (s t , a t , r t , s t+1 , terminal t ) in the replay memory; S8 define state s t = s t+1 ; S9 depression detection agent randomly samples replay memory to obtain a batch of (s b ,a b ,r b ,s b+1 ,terminal b ) quintuples, calculates the expected reward of (s b ,a b ,r b ,s b+1 ,terminal b ) quintuples using the structure of Double DQN, and updates the Q network with the difference between the actual reward and the expected reward. S10 repeats S7 to S9 until terminal t = true; and calculates the number of majority class samples D in the replay memory rma and the number of minority class samples D rmin to update the reward adjustment rate in the adaptive reward function R The current round number is incremented by 1, and the process returns to S6. S11 repeating S6 to S10 until the number of episodes episo des is completed; S12 in the test process, the sample data of the subjects are input into the trained depression detection agent for depression classification, and the classification results of each group of sample data are finally majority voted to obtain the depression detection result of a single subject.
2. The speech depression detection system based on deep reinforcement learning according to claim 1, wherein, The high-level acoustic features in step S1 are constructed in the following process: first, 25 low-level acoustic features such as frequency-related parameters, energy / amplitude-related parameters, and spectral parameters are extracted from the speech; then, all the low-level acoustic features are smoothed in the time dimension using a 3-frame long symmetric moving average filter, and then the arithmetic mean and standard deviation are applied to each low-level acoustic feature to generate 50 corresponding high-level acoustic features; for loudness and pitch, 27 additional high-level acoustic features are calculated according to different low-level acoustic features; in addition, there are 88 parameters including the Alpha ratio of unvoiced segments, the Hammarberg index, the arithmetic mean of the 0-500HZ and 500-1500HZ spectral slopes, the equivalent sound level, and another 6 time-related features, which constitute the high-level acoustic features.
3. The speech depression detection system based on deep reinforcement learning according to claim 1, wherein, The four elements in the depression detection task in step S3 are defined as follows in the Markov decision problem: the samples in the environment are defined as the state space S = {s1, s2,..., s t}, where s t is a sample randomly selected from the sample space at time t, and each time step t corresponds to a sample; the set of depression labels is defined as the action space A = {0, 1}, where a t represents that the corresponding label of the state s t at time t is depression or non-depression; the reward function R is an adaptive reward function constructed to solve the class imbalance problem: where s t , a t , and y t denote the state, action and sample label at time t, respectively, D min denotes the minority class sample set, D max denotes the majority class sample set, p is the reward adjustment rate, and l is a penalty factor to control the size of the penalty value at the beginning of training; the policy function is the basis for the depression detection agent to select actions according to the Q value, and is denoted by p(a t | s t ), which represents the policy of the agent selecting action a t at time t under state s t , and the epsilon greedy policy is used in the system: where q is a random number between (0, 1), ε is a value gradually reduced with time steps, Q(s t a) is the Q value function corresponding to each action calculated by the Q network in the state at time t.
4. The speech depression detection system based on deep reinforcement learning according to claim 1, wherein, The Q network in the depression detection agent constructed in step S5 is a time sequence calculation network containing a 1D convolution and a bidirectional long short-term memory network (BiLSTM), and there is a jump connection path between the 1D convolution and the bidirectional long short-term memory network (BiLSTM) to preserve the original input and fuse it with the convolution features. After 1D convolution calculation, the deep representation extracted is concatenated with the high-level acoustic features and then input into the bidirectional long short-term memory network (BiLSTM) for Q value calculation; after construction, the current network θ needs to be initialized, and the network parameters are copied to another network with the same structure, called target network θ'.
5. The speech depression detection system based on deep reinforcement learning according to claim 1, wherein, The Q network with Dueling structure in step S7 refers to adding two fully connected layers at the end of the network when calculating the Q value using the Q network, which are used to calculate the state value v(s) and the advantage value A(s,a) corresponding to each action, respectively, and use these two values to calculate the final Q value to reduce unnecessary exploration of the agent. The calculation formula of the Q value is as follows: Subsequently, the depression detection AI will select the current state s according to the previously defined epsilon greedy strategy. t The following action a t , will action a t The corresponding sample label in the environment t The comparison is performed, and the reward value r is obtained according to the adaptive reward function R. t And obtain the next state value s from the state space S. t+1 Meanwhile, based on the time step t and the state s t Action a t and tag l t Define the terminator variable terminal t and (s t ,a t ,r t ,s t+1 ,terminal t The quintuple is stored in the replay memory, where the termination function is defined as follows:
6. The speech depression detection system based on deep reinforcement learning according to claim 1, wherein, The Double DQN structure is used in step S9 for network updating by random sampling from the replay memory, that is, the current network θ calculates the expected reward of the state s b , the target network θ' calculates the reward of the next state s b+1 , and the actual reward is obtained by adding the reward value r b ; if terminal b = true in the five-tuple (s b , a b , r b+1 , s b , terminal b ), the actual reward is the reward value r b , the difference between the actual reward and the expected reward is used for parameter updating of the current network, and the parameters of the current network are directly assigned to the target network after a fixed time step I, so as to reduce the correlation between the networks; the calculation formula of Loss is as follows: