Bipolar Disorder Recognition Method and System Based on Speech Analysis
Through the WavLM model and LSTM network combined with the SE_Res2Blocks module, speech features are extracted and processed, timestamps and band features are obtained, which solves the problem of low accuracy in bipolar disorder recognition and achieves more efficient bipolar disorder recognition.
Patent Information
- Application Number
- CN202510562301.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The prior art has low accuracy in the field of bipolar disorder recognition, especially in the Chinese pronunciation environment, and the accuracy of the recognition of depression, (hypo)manic episodes and stable mood status in patients with bipolar disorder is not high, and there are few studies.
The WavLM model is used to extract speech features, combine long and short-term memory network (LSTM) and multiple SE_Res2Blocks modules to capture context information and extract emotional information from different dimensions. The timestamp and frequency band characteristics are obtained through convolution operations and time-frequency attention modules, and finally the label classifier is used to classify mood states.
It significantly improves the accuracy of the bipolar disorder identification system, reduces the computational complexity and reduces the risk of overfitting, and improves the recognition performance of bipolar disorder mental state.
Smart Images

Figure CN120089165B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular, to a bipolar disorder recognition method and system based on speech analysis. Background Art
[0002] As the acoustic manifestation of human language, speech not only conveys the substantial content of speech, but also carries the emotional information deep in the speaker's heart. For this reason, speech can be used as an important physiological behavior carrier for studying the emotional state of humans and has broad application prospects in the field of bipolar disorder recognition. However, at present, the research on speech recognition of mental diseases mainly focuses on the analysis of speech feature parameters and the research of speech recognition models for depressive disorders, and there are few research reports on bipolar disorder and the quality is limited. In the field of bipolar disorder mood state recognition, the commonly used speech acoustic features are prosody, voice quality and spectral features. At present, researchers have found that using self-supervised models such as Hubert and WavLM for speech feature extraction can significantly improve the performance of speech tasks and have excellent performance in other speech tasks, but there is no relevant research report in the field of bipolar disorder recognition. Although deep learning speech recognition technology has developed vigorously, there are only a few research reports in the field of mental diseases at present: some scholars have classified the (hypo)manic episodes and mood stable states of bipolar disorder patients in publicly available foreign speech datasets by applying four deep learning algorithms, namely Feed-forward Deep Neural Networeks (Feed forward DNNs), Gated Recurrent Neural Networks (GRNNs), Convolutional Neural Networks and Long Short Term Memory Network (LSTM). The accuracy rate ranges from only 60% to 72%, and the unweighted average recall (UAR) ranges from 45% to 65%. Generally speaking, in the above results, the recognition of bipolar disorder mood states does not show good recognition performance, the obtained accuracy rate is low, and the results are not ideal. And the current research mainly comes from English-speaking countries, and the research related to Chinese speech is very limited. There are differences in speech features between different languages, especially between Chinese and English. So far, there has been no literature reporting on the speech recognition research of the three different mood states of the depressive episode state, (hypo)manic episode state and mood stable state of bipolar disorder in the Chinese speech environment.
[0003] Therefore, how to improve the accuracy of the bipolar disorder recognition system has become a technical problem to be solved. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art or related technologies, and discloses a bipolar disorder recognition method and system based on speech analysis, which can effectively improve the accuracy of the bipolar disorder recognition system based on speech.
[0005] Term Explanation:
[0006] LSTM (long short-term memory network) is a special type of RNN that can be used to solve the long-term dependence problem in recurrent neural networks. In many tasks, the recurrent neural network with the LSTM structure performs better than the standard recurrent neural network.
[0007] Res2Net is a new type of multi-scale network architecture that enhances the ability of convolutional layers to process features of different scales through block connections. In many tasks, such as object detection, semantic segmentation, instance segmentation, etc., it shows higher accuracy and scale effectiveness.
[0008] The SE Block is a structure used in deep learning models and can be embedded in other classification or detection models. This module uses global average pooling operations to extract channel information, then generates channel weights through linear layers and activation functions, and finally enhances effective features and weakens ineffective features by weighted adjustment of the feature map.
[0009] The first aspect of the present invention discloses a bipolar disorder recognition method based on speech analysis, including: extracting features from a speech sequence based on the WavLM model to obtain frame-level speech features; processing the frame-level speech features through a long short-term memory network to capture context information and obtain speech features containing context information; using multiple SE_Res2Blocks modules to process the speech features containing context information to extract emotional information of different dimensions, obtaining multiple representations of the speech sequence and fusing them; inputting the fused multiple representations into a convolutional layer, and further refining and purifying the features through convolutional operations to obtain an aggregated representation; inputting the aggregated representation into a time-frequency attention module to obtain timestamp features and frequency band features related to the detection of the bipolar disorder mood state; inputting the timestamp features and frequency band features into a fully connected layer and connecting a function to construct a label classifier, thereby completing the mood state classification.
[0010] According to the bipolar disorder recognition method based on speech analysis disclosed by the present invention, preferably, the SE_Res2Blocks module includes a Res2Net module and an SE Block module. The Res2Net module is used to implement convolutional operations, and the SE Block module is used to perform dynamic channel feature recalibration.
[0011] According to the bipolar disorder recognition method based on speech analysis disclosed in the present invention, preferably, there are three SE_Res2Blocks modules, and the three SE_Res2Blocks modules are connected in sequence. At the same time, the output of the first SE_Res2Blocks module, the output of the second SE_Res2Blocks module, and the output of the third SE_Res2Blocks module are fused and connected to the convolutional layer.
[0012] According to the bipolar disorder recognition method based on speech analysis disclosed in the present invention, preferably, the step of inputting the aggregated representation into the time-frequency attention module to obtain the timestamp features and frequency band features related to the detection of the mood state of bipolar disorder specifically includes:
[0013] Obtain the attention weight vector in the time dimension:
[0014]
[0015] Obtain the attention weight vector in the frequency dimension:
[0016]
[0017] where X is the input representation, , and are two Conv1D operations along the time axis and the frequency axis respectively; and represent the linear transformation in the time dimension; and represent the linear transformation in the frequency dimension; () and () are the ReLU activation function and the Softmax activation function respectively;
[0018] Jointly highlight the timestamps and frequency bands related to bipolar disorder detection:
[0019]
[0020]
[0021] where X TFA is the output of the time-frequency attention module, is the element-wise multiplication, is the matrix multiplication operation, represents the attention matrix of X; is the Softmax function.
[0022] According to the bipolar disorder recognition method based on speech analysis disclosed in the present invention, preferably, the calculation process of the label classifier includes:
[0023]
[0024] Among them, and are the weight matrix and bias vector learned during the training process, represents the input feature, Softmax() represents the activation function, and y represents the attention weight vector.
[0025] According to the bipolar disorder recognition method based on speech analysis disclosed in the present invention, preferably, the speech sequence is the speech of a bipolar disorder patient.
[0026] The second aspect of the present invention discloses a bipolar disorder recognition system based on speech analysis, including: a memory for storing program instructions; a processor for calling the program instructions stored in the memory to implement the bipolar disorder recognition method based on speech analysis according to any one of the above technical solutions.
[0027] The beneficial effects of the present invention at least include: using speech analysis and processing technology to extract speech features, using the WavLM features extracted by the self-supervised model as the basic features for speech mood state recognition to reduce the influence of noise in speech samples; passing the extracted features into the LSTM layer to capture context information, and then passing them into three layers of SE_Res2Blocks to enrich the mood representation; passing the aggregated three-layer representation into the convolutional layer, and further refining and purifying the features through convolutional operations, which helps to reduce the computational complexity and prevent overfitting, and passing them into the TFA module to obtain the timestamps and frequency bands related to mood state detection. It can effectively improve the accuracy of the bipolar disorder recognition system based on speech. Finally, the label classifier is used to predict the final emotional state. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 Shows a schematic flow chart of a bipolar disorder recognition method based on speech analysis according to an embodiment of the present invention.
[0029] Figure 2 Shows a schematic structural diagram of the SE_Res2Blocks module according to an embodiment of the present invention.
[0030] Figure 3 Shows a schematic block diagram of a bipolar disorder recognition system based on speech analysis according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To better understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. In the following description, numerous specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0032] As Figure 1 shown, according to an embodiment of the present invention, a bipolar disorder recognition method based on speech analysis is disclosed, including:
[0033] Step S101, feature extraction: Based on the WavLM model, feature extraction is performed on the speech sequence to obtain frame-level speech features;
[0034] In this step, traditional audio often uses Mel spectrograms as a feature extraction method. However, Mel spectrograms usually only provide local frequency domain information and cannot capture complex patterns. At the same time, Mel spectrograms are sensitive to changes in the input signal, so they may perform poorly in a noisy environment. The self-supervised model WavLM considers various noises and transformations during the training process, which can improve the robustness of the features to environmental interference. In addition, it can utilize a large amount of unlabeled speech data to learn richer feature representations through context information. Therefore, the present invention uses the self-supervised model WavLM to encode the speech sequence to obtain frame-level speech features , to obtain a richer feature representation of the speech sample. Where T represents the number of acoustic frames of the speech, and D in is the unified encoding dimension of the features of the self-supervised model.
[0035] Step S102, feature encoding: The frame-level speech features are processed through a long short-term memory network to capture context information and obtain speech features containing context information;
[0036] In this step, emotional expression is time-varying. The emotional state in the speech signal is often affected by the previous and subsequent speech frames, and the emotional color in the text is also affected by the previous and subsequent words. When the sequence is long enough, it is difficult for traditional recurrent neural networks to transfer information from earlier time steps to subsequent time steps. The LSTM (long short-term memory network) can remember the information from earlier time steps. Therefore, the present invention uses LSTM to encode the input frame-level speech features respectively to obtain encoded features , where T represents the number of acoustic frames of the speech, and D is the unified encoding dimension of the hidden layer features of the LSTM. Relying on the efficient gating mechanism in the LSTM, it can fully capture the dynamic changes and temporal dependencies of the emotional information in the speech sequence.
[0037] Step S103, enriching emotional representation: Using multiple SE_Res2Blocks modules to process the speech features containing context information to extract emotional information in different dimensions, obtaining multiple representations of the speech sequence and fusing them;
[0038] In this step, the unprocessed speech features containing context information often only contain emotional information of a single scale. To further enhance its fine-grained representation, the present invention uses a three-layer SE_Res2Blocks module (squeeze-and-excitation residual second-order module) as shown in Figure 2 to extract emotional information in different dimensions respectively. SE_Res2Blocks is mainly composed of one-dimensional convolution (Conv1d), residual second-order one-dimensional convolution (Res2Conv1d), and squeeze-and-excitation module (SE-Block). One-dimensional convolution and residual second-order one-dimensional convolution implement convolution operations, and the squeeze-and-excitation module performs dynamic channel feature recalibration to improve the representation ability of the network. Residual second-order one-dimensional convolution replaces a group of filters with a smaller filter bank and then connects different filter banks in a hierarchical residual manner, which can process multi-scale features and significantly reduce the number of parameters. The specific process includes:
[0039] As shown in Figure 2 the three-layer SE_Res2Blocks structure shown, the SE-Block has the greatest impact on the model calculation result: First, a channel descriptor z is generated through the average value of the time-domain features, and then the weights of each channel are calculated using formula (1):
[0040] (1)
[0041] where, σ( ) and ( ) represent the sigmod and tanh functions respectively, and are weight matrices learned during the training process, and are bias vectors.
[0042] Finally, the original input features are multiplied by the weights (s) of the channels for reweighting and input into the next layer.
[0043] The dilation rate of the first-layer SE_Res2Blocks module is 2, the dilation rate of the second-layer SE_Res2Blocks module is 3, and the dilation rate of the third-layer SE_Res2Blocks module is 4. The outputs of each block are connected (fused) and input into the convolutional layer to effectively utilize the shallow and deep features expressing different information segments.
[0044] Step S104, Feature Purification: Input the fused multiple representations into the convolutional layer, and further refine and purify the features through convolutional operations to obtain aggregated representations;
[0045] In this step, input the aggregated representations output by the above three-layer SE_Res2Blocks module into the convolutional layer, and further refine and purify the features through convolutional operations, which helps to reduce the computational complexity and can also prevent overfitting.
[0046] Step S105, Time-Frequency Attention: Input the aggregated representations into the time-frequency attention module to obtain timestamp features and frequency band features related to the detection of bipolar disorder mood states;
[0047] In this step, for the input representations , use formulas (2) and (3) to obtain the attention weight vectors in the time dimension and frequency dimension respectively:
[0048] (2)
[0049] (3)
[0050] Where and are two Conv1D operations along the time axis and frequency axis respectively. and are two linear transformations, where r T is the decay ratio in the time dimension. and are two linear transformations, where r D is the decay ratio in the frequency dimension. ( ) and ( ) refer to the activation functions of ReLU and Softmax respectively.
[0051] Formula (4) jointly highlights the timestamps and frequency bands related to depression detection.
[0052] (4)
[0053] Where is element-wise multiplication, and m TFA is obtained using equation (5):
[0054] (5)
[0055] Where is matrix multiplication operation, corresponds to the attention matrix of X. is the Softmax function. Relying on the TFA component, it pays more attention to the timestamps and frequency bands that are discriminative for the bipolar disorder mood detection task.
[0056] Step S106, mood state classification: Input the timestamp features and frequency band features into the fully connected layer and connect a function to construct a label classifier, thereby completing the mood state classification.
[0057] In this step, in order to extract the key features of each basic emotion category from the fused representation and then complete the correct emotion classification, a fully connected layer and a function are used to construct a label classifier. The label classifier predicts emotions by mapping the deep representation from the common space to the emotion space. The calculation method is as shown in formula (6):
[0058] (6)
[0059] where and are the weight matrix and bias vector learned during the training process.
[0060] As Figure 3 shown, according to an embodiment of the present invention, a bipolar disorder recognition system 300 based on speech analysis is disclosed, including: a memory 301 for storing program instructions; a processor 302 for calling the program instructions stored in the memory to implement the bipolar disorder recognition method based on speech analysis as described in the above embodiment.
[0061] In summary, the present invention first uses speech analysis and processing technology to extract speech features, uses the WavLM features extracted by the self-supervised model as the basic features for speech mood state recognition to reduce the influence of noise in the speech samples. Then, the extracted features are passed into the LSTM layer to capture context information, and then passed into three layers of SE_Res2Blocks to enrich the mood representation. The aggregated three-layer representation is passed into the convolutional layer, and the features are further refined and purified through convolutional operations, which helps to reduce the computational complexity and prevent overfitting. Then, it is passed into the TFA module to obtain the timestamps and frequency bands related to the mood state detection. Finally, the label classifier is used to predict the final mood state. It can effectively improve the accuracy of the speech-based bipolar disorder recognition system.
[0062] All or part of the steps in the various methods of the above embodiments can be completed by controlling related hardware through a program, and the program can be stored in a readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disc memories, tape memories, or any other readable medium capable of carrying or storing data.
[0063] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for identifying bipolar disorder based on speech analysis, characterized in that, Including: Performing feature extraction on a speech sequence based on the WavLM model to obtain frame-level speech features; Processing the frame-level speech features through a long short-term memory network to capture context information and obtain speech features containing context information; Processing the speech features containing context information by using multiple SE_Res2Blocks modules to extract emotional information in different dimensions, obtaining multiple representations of the speech sequence and fusing them; Feeding the fused multiple representations into a convolutional layer, and further refining and purifying the features through convolutional operations to obtain an aggregated representation; Inputting the aggregated representation into a time-frequency attention module to obtain timestamp features and frequency band features related to the detection of bipolar disorder mood states; Input the timestamp feature and the frequency band feature into a fully connected layer and connect a function to construct a label classifier, thereby completing the mood state classification.
2. The method for identifying bipolar disorder based on speech analysis according to claim 1, wherein The SE_Res2Blocks module includes a Res2Net module and an SE Block module. The Res2Net module is used to implement convolutional operations, and the SE Block module is used to perform dynamic channel feature recalibration.
3. The method for identifying bipolar disorder based on voice analysis according to claim 1, characterized in that, There are three SE_Res2Blocks modules in total. The three SE_Res2Blocks modules are connected in sequence, and at the same time, the output of the first SE_Res2Blocks module, the output of the second SE_Res2Blocks module, and the output of the third SE_Res2Blocks module are fused and connected to the convolutional layer.
4. The method for identifying bipolar disorder based on speech analysis according to claim 1, characterized in that, The step of inputting the aggregated representation into a time-frequency attention module to obtain timestamp features and frequency band features related to the detection of bipolar disorder mood states specifically includes: Obtaining an attention weight vector in the time dimension: Obtaining an attention weight vector in the frequency dimension: Among them, X is the input representation, , and are two Conv1D operations along the time axis and the frequency axis respectively; and represent linear transformations in the time dimension; and represent linear transformations in the frequency dimension; ( ) and ( ) are the ReLU activation function and the Softmax activation function respectively; Jointly highlighting timestamps and frequency bands related to bipolar disorder detection; Among them, X TFA is the output of the time-frequency attention module, is element-wise multiplication, is matrix multiplication operation, represents the attention matrix of X; is the Softmax function.
5. The method for identifying bipolar disorder based on speech analysis according to claim 1, wherein The calculation process of the label classifier includes: Among them, and are the weight matrix and bias vector learned during the training process, represents the input features, Softmax() represents the activation function, and y represents the attention weight vector.
6. The method for identifying bipolar disorder based on voice analysis according to any one of claims 1 to 5, characterized in that, The speech sequence is the speech of a bipolar disorder patient.
7. A bipolar disorder recognition system based on speech analysis, characterized in that, Including: A memory for storing program instructions; A processor for calling the program instructions stored in the memory to implement the bipolar disorder recognition method based on speech analysis according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice depression automatic detection method and device based on time delay neural network
CN115116475A
Speech emotion recognition method based on beta-attention variation auto-encoder feature decoupling network
CN116758905A