Voice-based depression detection model construction method and depression detection system

By constructing the DisNet framework, combining the learnable frequency domain filter group and the hierarchical speech feature extraction module, and using a self-supervised learning strategy, the problem of insufficient interpretability and generalization performance of existing depression detection methods is solved, and efficient depression detection and feature visualization is achieved.

CN120356490AActive Publication Date: 2025-07-22NANJING MEDICAL UNIV

Patent Information

Application Number
CN202510838622.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The existing deep learning-based depression detection methods lack interpretability, cannot effectively identify depression-related speech patterns, and are too dependent on labeled data, resulting in insufficient generalization performance of the model.

Method used

A depression screening network framework (DisNet) is built, and the learning frequency domain filter group module and hierarchical speech feature extraction module are combined with a self-supervised learning strategy to learn depression characteristics in speech end-to-end, so as to achieve the interpretability of features and the generalization performance of the model.

Benefits of technology

It improves the accuracy and interpretability of depression detection, can clearly reveal the acoustic characteristics related to depression in the pronunciation, provides an intuitive basis for clinical diagnosis, reduces dependence on labeled data, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356490A_ABST
    Figure CN120356490A_ABST
Patent Text Reader

Abstract

The invention discloses a depression detection model construction method based on voice and a depression detection system, which are used for performing voice depression detection by constructing a depression detection model and using voice data, can learn depression characteristics existing in voice end to end, and realizes quantitative and visual representation of depression biomarkers at the same time. Besides, in order to solve the problem of weak model interpretability caused by less clinical data, the invention also improves the interpretability of the features and the generalization performance of the model by using a novel voice self-supervised learning strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the intersection of psychiatry, artificial intelligence and machine learning, and in particular to a method for constructing a speech-based depression detection model and a depression detection system. Background Art

[0002] Depression is a common mental health problem and a major contributor to the global burden of disease. Depression usually begins in adolescence, may persist or relapse in adulthood, and often becomes a lifelong chronic mental disorder. Currently, the detection and screening of depression mainly rely on subjective methods such as questionnaires and clinical interviews, such as the Patient Health Questionnaire-9 (PHQ-9), the Hamilton Depression Rating Scale 17-item (HAMD-17), and the Depression Anxiety Stress Scale-21 (DASS-21). However, these methods are easily affected by the interviewer's experience, the quality of the questionnaire, and the willingness of the subjects. Therefore, it is necessary to develop an objective method for detecting depression.

[0003] As an easily accessible biological data, speech provides an effective perspective for depression detection, because various emotional states can be expressed through different vocalization patterns. Previous research reports have shown that the speech of patients with depression exhibits low pitch, slow pitch changes, slurred speech, or long pauses. In summary, designing a speech-based depression detection (SDD) method to identify acoustic biomarkers related to depression has important practical significance for improving detection performance.

[0004] Recently, there has been significant development of deep learning (DL) methods for SDD tasks, which can be divided into two main categories: developing models based on manually extracted spectral features, and building models that can directly process raw signals.

[0005] 1) The first category of methods utilizes spectral features such as power spectrograms, Mel filter banks (MFbanks), and MFCCs as input, and processes these features through convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) combined with attention mechanisms to mine and identify key local speech features.

[0006] 2) The second type of method avoids relying on hand-crafted spectral features and instead learns features directly from the original signal. For example, by using multiple convolutional layers to reduce the dimensionality of the speech signal, an input of similar size to the spectral features is generated. However, due to the high dimensionality of the input signal, such models usually require larger convolution kernels and more complex network structures, resulting in an excessive number of model parameters.

[0007] In addition, although the SDD method based on the above deep learning method has made certain progress, a major limitation of these models is the lack of interpretability, that is, the inability to identify meaningful depressive speech patterns or acoustic metrics. In fact, in actual clinical practice, being able to discover objective biomarkers can provide more clinical support for depression detection than simply slightly improving the prediction accuracy. Summary of the Invention

[0008] The present invention aims to solve at least one of the technical problems existing in the related art to a certain extent.

[0009] An object of the present invention is to provide a method for constructing a depression detection model based on speech, and construct an interpretable depression screening network framework (Learning interpretable depression representations in speech, DisNet) for detecting depression using speech, which can end-to-end learn the depressive features existing in speech. At the same time, by using a novel speech self-supervised learning strategy (SLRD), the interpretability of the features and the generalization performance of the model are improved, and the problem of weak interpretability of the model caused by less clinical data is solved.

[0010] Another object of the present invention is to provide a depression detection system applying the above construction method to perform speech-based depression detection and realize the quantification and visual representation of depression biomarkers.

[0011] In order to achieve the above object, on the one hand, the present invention provides a method for constructing a depression detection model based on speech, including the following steps:

[0012] S100. Collect speech data and add labels indicating whether it belongs to depressive speech; preprocess the collected speech data to construct a speech data set;

[0013] S200. Construct a depression screening network framework, where the depression screening network framework at least includes a learnable frequency-domain filter bank module for extracting depression-related spectral features from speech signals, and a hierarchical speech feature extraction module for extracting depression-related representations and locating key pronunciation regions;

[0014] The speech data is input into the depression screening network framework, and the learnable frequency-domain filter bank module adaptively adjusts the frequency response of the filter to select depression-related spectral features from the speech data;

[0015] The spectral features extracted by the learnable frequency-domain filter bank module are input into the hierarchical speech feature extraction module;

[0016] The hierarchical speech feature extraction module performs block processing on the input spectral features, and performs compression and sparsification operations through forward propagation, gradually discarding features irrelevant to the depression detection task and retaining key depression-related features;

[0017] S300. Train the constructed depression screening network framework in two stages;

[0018] Among them, the first stage is the pre-training stage. An unsupervised learning strategy is adopted. Speech samples are randomly sampled from the speech dataset, and non-overlapping segments are extracted as the input data of the depression screening network framework. The sampled segments are respectively input into the learnable frequency-domain filter bank module and the hierarchical speech feature extraction module to extract emotional representations and speaker representations, and perform interaction and reconstruction operations on the extracted representations. Update the model parameters of the depression screening network framework according to the designed loss function, and save the weights of the depression screening network framework;

[0019] The second stage is the fine-tuning stage. Initialize the depression screening network framework with the network weights in the pre-training stage. Adopt a self-supervised learning strategy. On the labeled speech dataset, use the known labels to fine-tune the model parameters of the depression screening network framework to obtain a trained depression detection model.

[0020] A further preferred technical solution of the present invention is that the preprocessing of the collected speech data in step S100 to construct a speech dataset is specifically as follows:

[0021] Normalize the collected speech data, adjust the sampling rate to a unified standard of 16 kHz, and perform non-overlapping sliding window processing on the speech signal. Each segment has a length of 10 s;

[0022] Define the complete speech signal of the i-th subject as , where , represents the total length of the speech, and I is the total number of subjects; After sliding window processing, it is divided into a set composed of several segments. The label of each segment is the same as the label of the original speech. Among them, J represents the total number of segments; the length of each segment is , and the label indicates whether the segment belongs to depressive speech. The predicted label of the speech is obtained by majority voting on all segment labels.

[0023] Preferably, in step S200, the learnable frequency-domain filter bank module selects spectrum features related to depression from speech data by adaptively adjusting the frequency response of the filters; specifically:

[0024] S210. Perform Fourier transform on the speech segment to convert it into the corresponding Fourier spectrum ;

[0025] Introduce a learnable frequency-domain filter bank to perform frequency-domain filtering, where the center frequency and bandwidth of a single filter are both learnable parameters for adaptive adjustment in response to different speech features, and obtain the feature , expressed as:

[0026] ;

[0027] where is the label of the filter, is the number of filters, ⊙ is the element-wise multiplication operation; the symbol represents the mean square absolute value operation, which is used to emphasize the amplitude information in the filter; is defined as:

[0028] ;

[0029] where 𝑏 represents the subscript of the sub-band bin; and respectively represent the learnable center frequency and standard deviation;

[0030] All filters are evenly distributed in all sub-band sets of length 𝐵, and the length of each filter is ; in addition, the difference and between the center frequencies of adjacent filters is constrained to be non-negative; the learnable frequency-domain filter bank is established by stacking a series of ; for a given STFT spectrum, the output feature of the learnable frequency-domain filter bank is expressed as:

[0031] ;

[0032] where each serves as an aggregation function of the power spectrum, and by adjusting its importance weight, relevant frequency information is retained or attenuated; represents the event frame;

[0033] S220. For each element of the output feature, along the filter axis, by using the exponential moving average of past values Normalize it, which is expressed as:

[0034] ;

[0035] The value of is updated by the following formula:

[0036] ;

[0037] After normalization, the dynamic range of the features of is reduced through a compression operation, which is expressed as: ;

[0038] ;

[0039] where 𝑡∈[0,𝑇] represents the time frame range, is a small constant; , , and are the gain, exponent, offset, and smoothing coefficient respectively, and are parameterized together and participate in the optimization and adjustment as the learnable parameters of the learnable frequency-domain filter bank module.

[0040] Preferably, the hierarchical speech feature extraction module described in step S200 performs block processing on the input spectral features, and performs compression and sparsification operations through forward propagation, gradually discarding the features irrelevant to the depression detection task and retaining the key depression-related features; specifically:

[0041] S230. Assume that the speech representation is distributed in the union of low-dimensional subspaces, and these subspaces are projected through the orthogonal basis set where , represents the dimension of the token, represents the dimension of the subspace after projection;

[0042] Define that the hierarchical speech feature extraction module maximizes the following objective function:

[0043] ;

[0044] where, represents the mapping function of each layer of the hierarchical speech feature extraction module, represents the parameter space of the mapping function, represents the expected value, and represent the coding rates between and within the representations respectively; represents the number of layers of the network, is obtained by taking The tokenized representation obtained by inputting into the patch embedding module based on the fully connected layer. Specifically, is divided into multiple patch blocks, and the dimension of each patch block is , is the dimension of each patch block after passing through the embedding module. Subsequently, a fully connected patch embedding layer serializes these blocks, where , represents the classification token, represents the 0th layer, that is, at the first input, the 1st to the Pth input tokens;

[0045] The hierarchical speech feature extraction module constructs a two-stage optimization network structure for the final representation, where each stage contains an incremental operation process that performs a forward mapping based on the distribution of its input ; ;

[0046] For a given subspace projection matrix , the hierarchical speech feature extraction module first compresses the token into the specified subspace using the multi-head attention mechanism in the gradient descent iteration process of minimizing the objective function to obtain the compressed representation , expressed as:

[0047] ;

[0048] where, represents the learning rate, and the multi-head attention mechanism MSA contains self-attention mechanisms, that is, ;

[0049] ;

[0050] ;

[0051] where, represents the gradient, is the total number of patch blocks, is a small constant;

[0052] S240. The hierarchical speech feature extraction module maximizes through a loose proximal gradient descent step to sparsify the token to obtain a sparse and structured output representation , and the sparsification strategy is:

[0053] 1) Introduce a learnable sparse dictionary Achieve a trade - off between representational diversity and sparsity, that is ;

[0054] 2) Introduce a norm constraint to the token representation to enhance its sparsity;

[0055] Based on these two strategies, re - define as:

[0056] ;

[0057] where is the regularization term for sparsification;

[0058] For the first term, when the dictionary is approximately orthogonal, that is , then:

[0059] ;

[0060] Accordingly, solve it according to the following formula:

[0061] ;

[0062] Relax the above formula into a convex optimization problem with a learnable re - weighted norm, and then transform it into a problem of jointly solving the dictionary and the weight parameter to enhance the sparsity of the representation, expressed as:

[0063] ;

[0064] where is the weight for retaining the key subspace;

[0065] In the optimization process, the iterative shrinkage - threshold algorithm is adopted, and the iterative update is as follows:

[0066] ;

[0067] where represents the step size, is used to introduce the non - negativity constraint. By stacking multiple hierarchical speech feature extraction modules, the representation will be gradually compressed and sparsified to extract more discriminative features.

[0068] Preferably, in the pre - training stage of step S300, finally is input into the fully - connected layer FC, Complete the classification task, and optimize the cross-entropy loss function for all parameters through backpropagation for update;

[0069] ;

[0070] Among them, , map the [CLS] token to a logits vector for classification prediction.

[0071] Preferably, in the fine-tuning stage of step S300, initialize the depression screening network framework with the network weights in the pre-training stage, adopt a self-supervised learning strategy, and use the known labels on the labeled speech dataset to fine-tune the model parameters of the depression screening network framework to obtain a trained depression detection model; specifically:

[0072] Extract non-overlapping segments from the same speech signal and , and generate similar input pairs from the same speech signal, where 𝑚 and 𝑛 represent the indices of each segment;

[0073] Construct a network framework consisting of two encoders, a projector, a predictor, and a discriminator; both encoders include a learnable frequency-domain filter bank module and a hierarchical speech feature extraction module, which are used to extract different emotional representations ER and speaker representations SR respectively;

[0074] Input the segments and into the learnable frequency-domain filter bank module to extract features and , where ; further process these features through the hierarchical speech feature extraction module, and use compression and sparsity operations to extract and ; interact and reconstruct the and in each segment to generate and ; then, perform a non-linear transformation on and through the projector to obtain and ; finally, input and into the predictor to obtain and , decouple the emotional representation ER and the speaker representation SR, and initialize the depression screening network framework with the emotional representation ER;

[0075] In the fine-tuning stage, multiple loss functions are combined, including the negative cosine similarity loss function , the discriminative constraint term , the local consistency constraint term and the adversarial loss function ; The fine-tuning stage is carried out by minimizing the following total loss function:

[0076] ;

[0077] wherein, , and are hyperparameters.

[0078] On the other hand, the present invention provides a depression detection system, including:

[0079] A model construction module, configured to construct a trained depression detection model by applying the above-mentioned method for constructing a speech-based depression detection model;

[0080] A depression detection module, configured to collect the speech data of the tester, and after preprocessing, input it into the trained depression detection model, extract the depression-related features therein, and use a classifier to perform depression detection to obtain specific categories;

[0081] A visualization display module, configured to visually display the classification results and the depression biomarkers, that is, the features learned by the depression detection model.

[0082] On the other hand, the present invention provides a non-transitory computer-readable storage medium, on which computer instructions are stored, and the computer instructions cause the computer to execute the above-mentioned method for constructing a speech-based depression detection model.

[0083] On yet another aspect, the present invention provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete mutual communication through the communication bus, and the processor calls the logic instructions in the memory to execute the above-mentioned method for constructing a speech-based depression detection model.

[0084] On still another aspect, the present invention provides a computer program product, the computer program product includes a computer program, the computer program is stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer executes the above-mentioned method for constructing a speech-based depression detection model.

[0085] Advantages: The present invention constructs an efficient deep learning screening network model (Learning interpretable depression representations in speech, DisNet) for detecting depression using speech, which can learn the depression features present in speech end-to-end. At the same time, it realizes the quantification and visual representation of depression biomarkers. In addition, to solve the problem of weak interpretability of the model caused by less clinical data, the present invention also improves the interpretability of features and the generalization performance of the model by using a novel speech self-supervised learning strategy (SLRD).

[0086] Through the synergistic effect of the learnable frequency-domain filter bank module and the hierarchical speech feature extraction module, DisNet can more effectively extract depression-related features in speech, and can achieve better classification performance compared with existing methods, improving the accuracy of depression detection; compared with traditional deep learning methods, DisNet can clearly reveal the acoustic features related to depression in speech, providing intuitive diagnostic basis for clinicians and enhancing the interpretability of the model; the speech self-supervised learning strategy SLRD of the present invention can alleviate the problem of limited labeled data to a certain extent, improve the generalization ability of the model and the recognition ability of depression features, and reduce the dependence on labeled data. Brief Description of the Drawings

[0087] Figure 1 It is the overall flowchart of the method for constructing a depression detection model based on speech of the present invention;

[0088] Figure 2 It is the data processing flowchart of the learnable frequency-domain filter bank module (LFB) in Embodiment 1 of the present invention;

[0089] Figure 3 It is the data processing flowchart of the hierarchical speech feature extraction module (HRE) in Embodiment 1 of the present invention;

[0090] Figure 4 It is the flowchart of the self-supervised learning strategy (SLRD) in Embodiment 1 of the present invention;

[0091] Figure 5 It is the diagram showing the recognition results of differential biomarkers in Embodiment 2 of the present invention;

[0092] Figure 6 It is the visualization diagram of the LFB features of the healthy control group in Embodiment 2 of the present invention;

[0093] Figure 7 It is the visualization diagram of the LFB features of the depression group in Embodiment 2 of the present invention. Detailed Description of the Invention

[0094] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention, and they should not be construed as limiting the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.

[0095] Based on the problems existing in the background technology, we assume that in the SDD task, the acoustic markers in pronunciation are irrelevant to the text content, while the identified local patterns are located in specific frequency bands within a specific time period of the spectral features, and these frequency bands can be used as biomarkers to assist in depression detection.

[0096] The principle on which the present invention is based is: We assume that in the SDD task, the acoustic markers in pronunciation are irrelevant to the text content, while the identified local patterns are located in specific frequency bands within a specific time period of the spectral features, and these frequency bands can be used as biomarkers to assist in depression detection. To study this, an interpretable depression screening network (DisNet) can be designed to lock in the regions containing these depression biomarkers. The following will describe Figures 1 - 7 the method for constructing a voice-based depression detection model and the depression detection system provided by the present invention.

[0097] Embodiment 1: This embodiment provides a method for constructing a voice-based depression detection model, including the following steps:

[0098] S100. Collect voice data and add labels indicating whether it belongs to depressive voice; preprocess the collected voice data to construct a voice data set.

[0099] Normalize the collected voice data, adjust the sampling rate to a unified standard of 16 kHz, and perform non-overlapping sliding window processing on the voice signal, with each segment having a length of 10 s;

[0100] Define the complete voice signal of the i-th subject as , where , represents the total length of the voice, and I is the total number of subjects; After sliding window processing, it is segmented into a set composed of several segments, and the label of each segment is the same as the label The length is , and the label indicates whether the segment belongs to depressive speech, and the predicted label of the speech is obtained by majority voting on all segment labels.

[0101] S200. Construct a depression screening network framework, as Figure 1 shown, which is a jointly optimized framework designed specifically for the SDD (Speech Detection of Depression) task, mainly including the construction and pre-training of the DisNet part and fine-tuning based on self-supervised learning strategies. First, to obtain spectrum features related to depression, the Learnable Frequency-domain Filter Bank (LFB) module can recalibrate the importance of each frequency in the speech. Next, the Hierarchical Speech Representation Extraction (HRE) module compresses and sparsifies the learned features to extract depression-related representations and locate key pronunciation regions. Finally, the self-supervised learning strategy SLRD guides DisNet to focus on emotional representations, thus providing general and interpretable representations for the fine-tuning stage of the SDD task.

[0102] It can be seen from this that DisNet includes two core modules: the Learnable Frequency-domain Filter Bank (LFB) module and the Hierarchical Speech Representation Extraction (HRE) module.

[0103] The processing method of the Learnable Frequency-domain Filter Bank (LFB) module for speech data, as Figure 2 shown: This module selects depression-related frequency band features from the speech by adaptively adjusting the frequency response of the filter. The LFB module includes a learnable frequency-domain filter and a non-linear transformation, which can dynamically adjust the frequency response and update the parameters. The LFB module is designed to extract depression-related spectrum features from the speech signal. Its core idea is to dynamically adjust the frequency response through a learnable filter bank to selectively extract depression-related frequency band information.

[0104] LFB selects depression-related spectrum features from the speech data by adaptively adjusting the frequency response of the filter; the specific method is:

[0105] S210. It is known that the convolution theorem provides time-domain and frequency-domain methods for signal analysis as follows: ;

[0106] Among them, represents the Fourier transform, and ∗ and ⊙ represent convolution and Hadamard product operations. and is the original time - series signal and its corresponding Fourier spectrum. and the impulse response and frequency response of the corresponding filter. In view of this, the existing time - domain filtering methods parameterize the center frequency and bandwidth, and use a convolutional layer with a step size and a kernel size to approximate . However, obtaining the output characteristics of the time - domain filter bank requires more computational cost, and the calculation formula is as follows:

[0107] ;

[0108] It should be noted that the equivalent calculation only requires one Fourier transform in the frequency domain.

[0109] Therefore, in this embodiment, the speech segment is Fourier - transformed into the corresponding Fourier spectrum ;

[0110] Then, in a more efficient way, a learnable frequency - domain filter bank is introduced to perform frequency - domain filtering on , where the center frequency and bandwidth of a single filter are both learnable parameters, which are adaptively adjusted to respond to different speech features, and the feature is obtained, which is expressed as:

[0111] ;

[0112] where is the label of the filter, is the number of filters, ⊙ represents the element - wise multiplication operation; the symbol represents the mean - square absolute - value operation, which is used to emphasize the amplitude information in the filter; is defined as:

[0113] ;

[0114] where 𝑏 represents the sub - band bin subscript; and represent the learnable center frequency and standard deviation respectively;

[0115] All filters are evenly distributed in all sub - band sets of length 𝐵, and the length of each filter is ; in addition, the difference and between the center frequencies of adjacent filters is constrained to be non - negative; the learnable frequency - domain filter bank is formed by stacking a series of Establish; for a given STFT spectrum, the output feature representation of the learnable frequency-domain filter bank can be expressed as:

[0116] ;

[0117] where each serves as an aggregation function of the power spectrum, and by adjusting its importance weight, relevant frequency information is retained or attenuated; represents the event frame;

[0118] S220, Nonlinear transformation: used to enhance the robustness to loudness changes.

[0119] In this embodiment, the nonlinear transformation is used to perform normalization and compression processing. Specifically, each element of the output feature along the filter axis is normalized by using the exponential moving average of past values and is expressed as:

[0120] ;

[0121] The value of

[0122] is updated by the following formula:

[0123] After normalization, the feature dynamic range of is reduced through a compression operation, which is expressed as:

[0124] ;

[0125] where 𝑡∈[0,𝑇] represents the time frame range, is a small constant; , , and are the gain, exponent, offset, and smoothing coefficient respectively, and are parameterized together and participate in optimization and adjustment as the learnable parameters of the LFB.

[0126] The hierarchical speech feature extraction module HRE, as Figure 3 shown, performs block processing on the input spectral features and executes compression and sparsification operations through forward propagation, gradually discarding the features irrelevant to the depression detection task and retaining the key depression-related features;

[0127] S230, Compression operation: project the features onto a low-dimensional subspace to compress the feature dimension.

[0128] The inherent temporal variability of speech and the uniqueness of each segment pose challenges to identifying depressive differences within the spectrum. To address this issue, this embodiment introduces an HRE module to decorrelate the features in order to retain the key information for the SDD task ( ).

[0129] Assume that the speech representations are distributed in the union of low-dimensional subspaces, which are projected through an orthogonal basis set , where , represents the dimension of the token, represents the dimension of the subspace after projection; as Figure 3 shows, the HRE module characterizes the marginal distribution of the token , that is, .

[0130] To ensure the mutual independence between representations and the compressibility within the same representation, HRE maximizes the following objective function:

[0131] ;

[0132] where, represents the mapping function of each layer of the hierarchical speech feature extraction module, represents the parameter space of the mapping function, represents the expected value, and represent the coding rates between representations and within representations, respectively; represents the number of layers of the network, is the tokenized representation obtained by inputting into the patch embedding module based on the fully connected layer. Specifically, is sliced into multiple patch blocks, and the dimension of each patch block is , is the dimension of each patch block after passing through the embedding module. Subsequently, a fully connected patch embedding layer serializes these blocks, where , represents the classification token, represents the 0th layer, that is, at the first input, the 1st to the Pth input tokens;

[0133] HRE establishes a two-stage optimization network structure for the final representation, where each stage contains an incremental operation process that performs a forward mapping based on the distribution of its input ;

[0134] For a given subspace projection matrix , HRE first compresses the tokens into the specified subspace by using the multi-head self-attention mechanism (MSA) in the gradient descent iteration process of minimizing the objective function to obtain the compressed representation , expressed as:

[0135] ;

[0136] where represents the learning rate, and the multi-head self-attention mechanism MSA contains self-attention mechanisms, namely ;

[0137] ;

[0138] ;

[0139] where represents the gradient, is the total number of patch blocks, is a small constant;

[0140] S240. Sparsification operation: The features are sparsified under the action of the global dictionary and the correlation matrix to obtain a sparse and structured representation.

[0141] HRE maximizes through a relaxed proximal gradient descent step to sparsify the token and obtain a sparse and structured output representation . In this embodiment, the sparsification strategy is:

[0142] 1) Introduce a learnable sparse dictionary to achieve a trade-off between representation diversity and sparsity, that is ;

[0143] 2) Introduce the norm constraint to the token representation

[0144] to enhance its sparsity; Based on these two strategies,

[0145] is redefined as:

[0146] where is the regularization term for sparsification;

[0147] For the first term, when the dictionary When it is approximately orthogonal, that is , then:

[0148] ;

[0149] Accordingly, Solve according to the following formula:

[0150] ;

[0151] Different from the static mesh data processing method, our method uses learnable parameters from a global perspective to estimate the temporal correlation between projection subspaces, so as to more effectively separate depression-related representations. Therefore, the above formula is relaxed into a convex optimization problem with a learnable reweighted norm, and then transformed into a problem of jointly solving the dictionary and the weight parameter to enhance the sparsity of the representation, expressed as:

[0152] ;

[0153] where is the weight for retaining the key subspace; during the optimization process, our method focuses on the important subspace by reducing the w_p value of the key subspace and increasing the weight of the non-key subspace.

[0154] In addition, the parameter optimization process adopts the Iterative Shrinkage-Thresholding Algorithm (ISTA), and the iterative update is as follows:

[0155] ;

[0156] where represents the step size, is used to introduce the non-negativity constraint. By stacking multiple HREs, the representation will be gradually compressed and sparsified to extract more discriminative features.

[0157] S300. Perform two-stage training on the constructed depression screening network framework;

[0158] Among them, the first stage is the pre-training stage, which adopts an unsupervised learning strategy. Randomly sample speech samples from the speech dataset, extract non-overlapping segments as the input data of the depression screening network framework. Input the sampled segments into the learnable frequency-domain filter bank module and the hierarchical speech feature extraction module respectively, extract the emotion representation and the speaker representation, and perform interaction and reconstruction operations on the extracted representations. Update the model parameters of the depression screening network framework according to the designed loss function, and save the weights of the depression screening network framework;

[0159] In the pre-training stage, finally is input into the fully connected layer FC, and the classification task is completed. All parameters are optimized by backpropagation to minimize the cross-entropy loss function for updating;

[0160] ;

[0161] Among them, , the [CLS] token is mapped to a logits vector for classification prediction.

[0162] The second stage is the fine-tuning stage. Initialize the depression screening network framework with the network weights of the pre-training stage, adopt a self-supervised learning strategy, and use the known labels on the labeled speech dataset to fine-tune the model parameters of the depression screening network framework to obtain a trained depression detection model.

[0163] To reduce the impact of limited labeled data and improve the reliability of biomarkers, this embodiment designs the SLRD strategy. SLRD is a self-supervised learning strategy through speech representation separation. Specifically, SLRD does not consider the text content and assumes that the language segments extracted from the same speech will approximately exhibit a consistent emotional state in a short period of time. Therefore, the emotion representations (ERs) and speaker representations (SRs) extracted from different segments of the same speech are similar or even the same. We recombine the ER and SR in each segment to promote the interaction and reconstruction of the representations. Finally, by decoupling the ER and SR and initializing the DisNet with the ER, the model can be guided to focus on the emotional regions of interest in advance, thereby improving the interpretability of the learned depression representations. Specifically:

[0164] Extract non-overlapping paragraphs from the same speech signal and , and generate similar input pairs from the same speech signal, where 𝑚 and 𝑛 represent the indices of each segment;

[0165] Construct a network framework consisting of two encoders, a projector, a predictor, and a discriminator; both encoders include a learnable frequency-domain filter bank module and a hierarchical speech feature extraction module, which are used to extract different emotional representations ER and speaker representations SR respectively;

[0166] Input the segments and into the learnable frequency-domain filter bank module to extract features and , where ; further process these features through the hierarchical speech feature extraction module, and use compression and sparsity operations to extract and ; perform interaction and reconstruction on and in each segment to generate and ; then, perform non-linear transformation on and through the projector to obtain and ; finally, input and into the predictor to obtain and , decouple the emotional representation ER and the speaker representation SR, and initialize the depression screening network framework with the emotional representation ER;

[0167] , , , , , , , ,

[0168] , , , , , , , , , , ,

[0169] In the fine-tuning stage, a combination of multiple loss functions is adopted, including the negative cosine similarity loss function , the discriminative constraint term , the local consistency constraint term and the adversarial loss function ; The fine-tuning stage is carried out by minimizing the following total loss function:

[0170] ;

[0171] where, 、 and are hyperparameters.

[0172] First, the negative cosine similarity is selected as the loss function to measure the similarity between representations, expressed as:

[0173] ;

[0174] where represents the stop-gradient operation. Therefore, the similarity between two speech segments is only determined by the same emotion representations (ERs) and speaker representations (SRs) in their latent representations.

[0175] Second, in order to ensure the distinctiveness between speaker representations (SRs) and emotion representations (ERs), a discriminative constraint term is introduced to enforce the speaker representations to satisfy the orthogonality condition, expressed as:

[0176] ;

[0177] where is the square of the Frobenius norm.

[0178] In addition, in order to minimize the difference between emotion representations (ERs) of two speech segments, a local consistency constraint term is introduced to align them within the shared subspace. This constraint term uses the Central Moment Discrepancy (CMD) to measure the difference, and its definition is as follows:

[0179] ;

[0180] where, is the interval distance constant, represents the empirical expectation, is the th-order sample central moment.

[0181] In addition, in order to strengthen the global consistency of emotion representations (ERs), an adversarial loss is introduced to constrain the ER distribution of all samples to approach the Gaussian distribution. Specifically, the discriminator uses the Earth Moving Distance (EMD) to reduce the difference between the ER distribution and the Gaussian distribution The difference between. Under the condition of satisfying Lipschitz continuity, the loss function of the discriminator is defined as follows:

[0182] ;

[0183] where represents the function of the discriminator, and the loss added to the generator DisNet is .

[0184] Example 2: A depression detection system, comprising:

[0185] A model construction module for constructing a trained depression detection model by applying the speech-based depression detection model construction method of Example 1;

[0186] A depression detection module for collecting the speech data of a tester, after preprocessing, inputting the trained depression detection model, extracting the depression-related features therein, and using a classifier to perform depression detection to obtain specific categories;

[0187] In this embodiment, in order to verify the detection effectiveness and accuracy, a healthy control group and a depression group are designed for separate sampling. During sampling, the Depression Anxiety and Stress Scales (DASS-21) and the Insomnia Severity Index (ISI) are used to evaluate emotional health and sleep patterns. Risk behaviors are evaluated by the Ottawa Self-injury 2 Inventory (OSI).

[0188] Subjects exceeding the predefined thresholds: Participants with a depression score (DS)>9, an anxiety score (AS)>7, or a stress score (SS)>14 are marked as being in a sub-healthy mental state, and the 7-item scores of each subscale are multiplied by 2 as the final score. In this embodiment, the depression severity classification based on the DASS-21 score is as follows: mild (10-13), moderate (14-20), severe (21-27), very severe (≥28). For the NCs group, the DASS-21 criteria are defined as follows: depression score ≤9, anxiety score ≤7, stress score ≤14, ISI score ≤7, and no suicidal behavior in the OSI result.

[0189] A visualization display module for visually displaying the classification results and the depression biomarkers, i.e., the features learned by the depression detection model.

[0190] Such as Figure 5As shown, DisNet discovers the specific differences in pronunciation phonemes between the control group and the depression group, and points out the specific ranges of each category. By measuring the specific ranges of the pronunciation phonemes of the subjects, it is possible to identify whether they have depression.

[0191] As Figure 6 , Figure 7 shown, by visualizing the features learned by LFB, it is possible to effectively distinguish whether a patient has depression. The significant features of depressive speech are manifested as discontinuity and the absence of fundamental frequency (F0). In contrast, the features of the normal control group (NCs) show an obvious striped distribution in the F0 region, with uniform energy distribution and stronger continuity.

[0192] The verification results show that the technical solution of the present invention can support the use of an intuitive visualization method to replace the traditional questionnaire survey. Compared with the traditional deep learning method, DisNet can clearly reveal the acoustic features related to depression in speech, providing an intuitive diagnostic basis for clinicians.

[0193] Embodiment 2: This embodiment provides a non-transitory computer-readable storage medium, on which computer instructions are stored, and the computer instructions cause the computer to execute a method for constructing a depression detection model based on speech, and the method includes the following steps:

[0194] S100. Collect speech data and add labels indicating whether it belongs to depressive speech; preprocess the collected speech data to construct a speech data set;

[0195] S200. Construct a depression screening network framework, where the depression screening network framework at least includes a learnable frequency-domain filter bank module for extracting spectrum features related to depression from speech signals, and a hierarchical speech feature extraction module for extracting depression-related representations and locating key pronunciation regions;

[0196] The speech data is input into the depression screening network framework, and the learnable frequency-domain filter bank module adaptively adjusts the frequency response of the filter to select spectrum features related to depression from the speech data;

[0197] The spectrum features extracted by the learnable frequency-domain filter bank module are input into the hierarchical speech feature extraction module;

[0198] The hierarchical speech feature extraction module performs block processing on the input spectrum features, and performs compression and sparsification operations through forward propagation, gradually discarding features irrelevant to the depression detection task and retaining key depression-related features;

[0199] S300. Perform two-stage training on the constructed depression screening network framework;

[0200] Among them, the first stage is the pre-training stage. An unsupervised learning strategy is adopted to randomly sample speech samples from the speech dataset, and extract non-overlapping segments therefrom as the input data of the depression screening network framework. The sampled segments are respectively input into the learnable frequency-domain filter bank module and the hierarchical speech feature extraction module to extract the emotion representation and the speaker representation, and perform interaction and reconstruction operations on the extracted representations. Update the model parameters of the depression screening network framework according to the designed loss function, and save the weights of the depression screening network framework;

[0201] The second stage is the fine-tuning stage. Initialize the depression screening network framework with the network weights of the pre-training stage. Adopt a self-supervised learning strategy. On the labeled speech dataset, use the known labels to fine-tune the model parameters of the depression screening network framework to obtain a trained depression detection model.

[0202] Embodiment 3: This embodiment provides an electronic device, which may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the logical instructions in the memory to execute the method for constructing a depression detection model based on speech, and the method includes the following steps:

[0203] S100. Collect speech data and add labels indicating whether it belongs to depressive speech; preprocess the collected speech data to construct a speech dataset;

[0204] S200. Construct a depression screening network framework, where the depression screening network framework at least includes a learnable frequency-domain filter bank module for extracting depression-related spectral features from speech signals, and a hierarchical speech feature extraction module for extracting depression-related features and locating key pronunciation regions;

[0205] Input the speech data into the depression screening network framework. The learnable frequency-domain filter bank module adaptively adjusts the frequency response of the filter to select depression-related spectral features from the speech data;

[0206] Input the spectral features extracted by the learnable frequency-domain filter bank module into the hierarchical speech feature extraction module;

[0207] The hierarchical speech feature extraction module performs block processing on the input spectral features, and performs compression and sparsification operations through forward propagation, gradually discarding features irrelevant to the depression detection task and retaining key depression-related features;

[0208] S300. Perform two-stage training on the constructed depression screening network framework;

[0209] Among them, the first stage is the pre-training stage. An unsupervised learning strategy is adopted to randomly sample speech samples from the speech dataset, and extract non-overlapping segments therefrom as the input data of the depression screening network framework. The sampled segments are respectively input into the learnable frequency-domain filter bank module and the hierarchical speech feature extraction module to extract the emotion representation and the speaker representation, and perform interaction and reconstruction operations on the extracted representations. The model parameters of the depression screening network framework are updated according to the designed loss function, and the weights of the depression screening network framework are saved.

[0210] The second stage is the fine-tuning stage. The depression screening network framework is initialized with the network weights of the pre-training stage. A self-supervised learning strategy is adopted, and on the labeled speech dataset, the known labels are used to fine-tune the model parameters of the depression screening network framework to obtain a trained depression detection model.

[0211] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0212] Embodiment 4: The present embodiment provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a method for constructing a depression detection model based on speech. The method includes the following steps:

[0213] S100. Collect speech data and add labels indicating whether it belongs to depressive speech; preprocess the collected speech data to construct a speech dataset.

[0214] S200. Construct a depression screening network framework. The depression screening network framework at least includes a learnable frequency-domain filter bank module for extracting spectrum features related to depression from speech signals, and a hierarchical speech feature extraction module for extracting depression-related characteristics and locating key pronunciation regions.

[0215] Speech data input depression screening network framework. The learnable frequency-domain filter bank module adaptively adjusts the frequency response of the filter to select the spectrum features related to depression from the speech data;

[0216] Input the spectrum features extracted by the learnable frequency-domain filter bank module into the hierarchical speech feature extraction module;

[0217] The hierarchical speech feature extraction module performs block processing on the input spectrum features, and executes compression and sparsification operations through forward propagation, gradually discarding the features irrelevant to the depression detection task and retaining the key depression-related features;

[0218] S300. Train the constructed depression screening network framework in two stages;

[0219] The first stage is the pre-training stage. Adopt an unsupervised learning strategy. Randomly sample speech samples from the speech dataset, extract the non-overlapping segments as the input data of the depression screening network framework, input the sampled segments into the learnable frequency-domain filter bank module and the hierarchical speech feature extraction module respectively, extract the emotion representation and the speaker representation, and perform interaction and reconstruction operations on the extracted representations. Update the model parameters of the depression screening network framework according to the designed loss function, and save the weights of the depression screening network framework;

[0220] The second stage is the fine-tuning stage. Initialize the depression screening network framework with the network weights in the pre-training stage. Adopt a self-supervised learning strategy. On the labeled speech dataset, use the known labels to fine-tune the model parameters of the depression screening network framework to obtain the trained depression detection model.

[0221] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0222] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0223] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a voice-based depression detection model, characterized in that, It includes the following steps: S100. Collect voice data and add labels indicating whether it belongs to depressive voice; preprocess the collected voice data to construct a voice dataset; S200. Construct a depressive screening network framework, which at least includes a learnable frequency-domain filter bank module for extracting depressive-related spectral features from voice signals, and a hierarchical voice feature extraction module for extracting depressive-related representations and locating key pronunciation regions; Input the voice data into the depressive screening network framework, and the learnable frequency-domain filter bank module selects depressive-related spectral features from the voice data by adaptively adjusting the frequency response of the filter; Input the spectral features extracted by the learnable frequency-domain filter bank module into the hierarchical voice feature extraction module; The hierarchical voice feature extraction module performs block processing on the input spectral features, and executes compression and sparsification operations through forward propagation, gradually discarding features irrelevant to the depressive detection task and retaining key depressive-related features; S300. Perform two-stage training on the constructed depressive screening network framework; Among them, the first stage is the pre-training stage. Adopt an unsupervised learning strategy, randomly sample voice samples from the voice dataset, extract non-overlapping segments as the input data of the depressive screening network framework, input the sampled segments into the learnable frequency-domain filter bank module and the hierarchical voice feature extraction module respectively, extract emotional representations and speaker representations, and perform interaction and reconstruction operations on the extracted representations. Update the model parameters of the depressive screening network framework according to the designed loss function, and save the weights of the depressive screening network framework; The second stage is the fine-tuning stage. Initialize the depressive screening network framework with the network weights in the pre-training stage, adopt a self-supervised learning strategy, and use the known labels on the labeled voice dataset to fine-tune the model parameters of the depressive screening network framework to obtain a trained depressive detection model.

2. The method for constructing a voice-based depression detection model according to claim 1, wherein The preprocessing of the collected voice data in step S100 to construct a voice dataset is specifically as follows: Normalize the collected voice data, adjust the sampling rate to the unified standard of 16 kHz, and perform non-overlapping sliding window processing on the voice signal, with each segment having a length of 10 s; Define the complete speech signal of the $i$-th subject as , where , represents the total length of the speech, and $I$ is the total number of subjects; After being processed by a sliding window, it is segmented into a set of several segments , and the label of each segment is the same as the label of the original speech , where $J$ represents the total number of segments; the length of each segment is , and the label indicates whether the segment belongs to depressive speech, and the predicted label of the speech is obtained by majority voting on all segment labels.

3. The method for constructing a voice-based depression detection model according to claim 2, wherein The learnable frequency-domain filter bank module in step S200 selects depressive-related spectral features from the voice data by adaptively adjusting the frequency response of the filter, specifically as follows: S210. Perform a Fourier transform on the voice segment to convert it into the corresponding Fourier spectrum ; Introduce a learnable frequency-domain filter bank to perform frequency-domain filtering, where the center frequency and bandwidth of a single filter are both learnable parameters, used to adaptively adjust in response to different speech features to obtain features , denoted as: ; where is the label of the filter, is the number of filters, and ⊙ represents the stationary point multiplication operation; the symbol represents the mean square absolute value operation, which is used to emphasize the amplitude information in the filter; is defined as: ; where 𝑏 represents the sub - band bin index; and represent the learnable center frequency and standard deviation respectively; All filters are evenly distributed among all subbands of length 𝐵, and the length of each filter is ; in addition, the difference between the center frequencies of adjacent filters is constrained to be non - negative; the learnable frequency - domain filter bank is established by stacking a series of ; for a given STFT spectrum, the output feature of the learnable frequency - domain filter bank is expressed as: ; Among them, each serves as an aggregation function of the power spectrum, and retains or attenuates relevant frequency information by adjusting its importance weight; represents an event frame; S220. For each element of the output feature along the filter axis, normalize it by using the exponential moving average of past values as follows: ; The value is updated by the following formula: ; After normalization, the feature dynamic range of is reduced through a compression operation, expressed as: ​ ; where \(t\in[0,T]\) represents the time frame range, is a small constant; , , and are the gain, exponent, offset, and smoothing coefficient respectively, which are parameterized together and participate in the optimization and adjustment as the learnable parameters of the learnable frequency-domain filter bank module.

4. The method for constructing a voice-based depression detection model according to claim 3, wherein The hierarchical voice feature extraction module in step S200 performs block processing on the input spectral features, and executes compression and sparsification operations through forward propagation, gradually discarding features irrelevant to the depressive detection task and retaining key depressive-related features; Specifically: S230. Assume that the speech representations are distributed in the union of low-dimensional subspaces, and these subspaces are projected by an orthogonal basis set , where , represents the dimension of the token, and represents the dimension of the subspace after projection; Define that the hierarchical voice feature extraction module maximizes the following objective function: ; Among them, represents the mapping function of each layer of the hierarchical speech feature extraction module, represents the parameter space of the mapping function, represents the expected value, and represent the coding rates between and within representations respectively; represents the number of layers of the network, is obtained by inputting into a patch embedding module based on a fully connected layer. Specifically, is sliced into multiple patch blocks, and the dimension of each patch block is , is the dimension of each patch block after passing through the embedding module. Subsequently, a fully connected patch embedding layer serializes these blocks, where , represents the classification token, represents the 0th layer, that is, at the first input, the 1st to Pth input tokens; The hierarchical speech feature extraction module establishes a two-stage optimization network structure for the final representation, where each stage contains an incremental operation process that performs a forward mapping based on the distribution of its input and executes a forward mapping ; For a given subspace projection matrix , in the hierarchical speech feature extraction module, first, in the gradient descent iteration process of minimizing the objective function , the multi-head attention mechanism is used to compress the tokens into the specified subspace, so as to obtain the compressed representation , which is expressed as: ; Among them, represents the learning rate, and the multi-head attention mechanism MSA contains self-attention mechanisms, that is ; ; ; Among them, represents the gradient, is the total number of patch blocks, is a small constant; S240. The hierarchical speech feature extraction module maximizes through a relaxed proximal gradient descent step , sparsifies the token to obtain a sparse and structured output representation . The strategy for the sparsification process is: 1) Introduce a learnable sparse dictionary Achieve the trade-off between representational diversity and sparsity, that is ; 2) Introduce a norm constraint on the token representation to enhance its sparsity; Based on these two strategies, redefine as follows: ; Among them, is a sparsity regularization term; For the first term, when the dictionary is approximately orthogonal, i.e., , then: ; Accordingly, Solve according to the following formula: ; Relax the above formula into a convex optimization problem with a learnable reweighted norm, and then transform it into a problem of jointly solving the dictionary and the weight parameters to enhance the sparsity of the representation, expressed as: ; Among them, is the weight for retaining the key subspace; In the optimization process, adopt the iterative shrinkage-threshold algorithm and update iteratively as follows: ; Among them, represents the step size, which is used to introduce non-negativity constraints. By stacking multiple hierarchical speech feature extraction modules, the representation will be gradually compressed and sparsified to extract more discriminative features.

5. The method for constructing a voice-based depression detection model according to claim 4, wherein In the pre-training phase of step S300, finally is input into the fully connected layer FC to complete the classification task, and all parameters are optimized by backpropagation to minimize the cross-entropy loss function for updating. ; Among them, , map the [CLS] token to a logits vector for classification prediction.

6. The method for constructing a voice-based depression detection model according to claim 2, wherein In the fine-tuning stage of step S300, initialize the depression screening network framework with the network weights of the pre-training stage, adopt a self-supervised learning strategy, and use the known labels on the labeled speech dataset to fine-tune the model parameters of the depression screening network framework to obtain a trained depression detection model. Specifically: Extract non - overlapping segments from the same speech signal and , generate similar input pairs from the same speech signal, where \(m\) and \(n\) represent the indices of each segment; Construct a network framework composed of two encoders, a projector, a predictor, and a discriminator. Both encoders include a learnable frequency-domain filter bank module and a hierarchical speech feature extraction module, which are respectively used to extract different emotional representations ER and speaker representations SR. The segments and are input into the learnable frequency-domain filter bank module to extract features and ; these features are further processed by the hierarchical speech feature extraction module, and and and are obtained by using compression and sparsity operations; the and in each segment are interacted and reconstructed to generate and ; then, the and are non-linearly transformed by the projector to obtain and ; finally, the and are input into the predictor to obtain and , decouple the emotion representation ER and the speaker representation SR, and initialize the depression screening network framework with the emotion representation ER; In the fine-tuning stage, multiple loss functions are combined, including the negative cosine similarity loss function , the discriminative constraint term , the local consistency constraint term and the adversarial loss function ; The fine-tuning stage is carried out by minimizing the following total loss function: ; Among them, , and are hyperparameters.

7. A depression detection system, characterized in that, Including: A model construction module for constructing a trained depression detection model by using the speech-based depression detection model construction method described in any one of claims 1-6. A depression detection module for collecting the speech data of the tester, and after preprocessing, inputting it into the trained depression detection model, extracting the depression-related features therein, and using a classifier to perform depression detection to obtain specific categories. A visualization display module for visually displaying the classification results and the depression biomarkers, that is, the features learned by the depression detection model.

8. A non-transitory computer-readable storage medium, characterized in that, It stores computer instructions, and the computer instructions cause the computer to execute the speech-based depression detection model construction method described in any one of claims 1-6.

9. An electronic device, characterized in that, Including: A processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor calls the logical instructions in the memory to execute the speech-based depression detection model construction method described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program, the computer program is stored on a non-transitory computer-readable storage medium, and when the computer program is executed by the processor, the computer executes the speech-based depression detection model construction method described in any one of claims 1-6.

Citation Information

Patent Citations

  • End-to-end sound barrier speech recognition method based on comparative learning

    CN113450777A

  • Depression recognition method, system and equipment based on voice analysis

    CN113633287A

  • Depression detection system based on electroencephalogram emotion nerve feedback signal

    CN115670463A

  • Depression risk detection model training method, depression symptom early warning method and related equipment

    CN115714002A

  • Depression recognition method and system based on voice representation

    CN116965819A

Cited By

  • Depression state prediction method and system based on voice multi-scale time domain perception

    CN120783803A

  • Depression state prediction method and system based on voice multi-scale time domain perception

    CN120783803B

  • Fatigue state recognition method based on learnable filter bank and joint regularization

    CN121265057A

  • Method for cross-domain speech deepfake detection

    US12475896B1