Audio feature extraction method based on cranial nerve model
By combining time convolutional neural network, multi-head attention and bidirectional long and short-term memory network, an adaptive rate smoothing leakage integration-triggered neuronal model is introduced, which solves the problem of audio feature extraction of depression under the influence of noise in the family environment, achieving higher diagnostic accuracy and robustness of depression.
Patent Information
- Application Number
- CN202510912582.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Existing audio data analysis models are susceptible to environmental noise in home environments, resulting in noise amplification and overfitting, making it difficult to effectively extract audio features related to depression.
The RBA-FE model based on the brain neural model is adopted, combining time convolutional neural networks, multi-head attention and bidirectional long and short-term memory networks, and an adaptive rate smooth leakage integration-triggered neuronal model (ARSLIF) is introduced to deal with noise sensitivity and gate overfitting problems through adaptive threshold adjustment and membrane potential leakage mechanism.
It enhances the robustness and generalization ability of the model in a noisy environment, can more accurately identify the speech characteristics of patients with depression, and improves the accuracy and robustness of the diagnosis of depression.
Smart Images

Figure CN120412657A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to an audio feature extraction method based on a cranial nerve model. Background Art
[0002] With the emergence of deep learning, a large number of neural network architectures for audio data analysis have been developed. For example, WavDepressionNet analyzes the original speech signal and uses a characterization and evaluation block to predict the degree of depression. Similarly, SpeechFormer++ classifies depression through paralinguistic features. It uses a unit encoder to simulate the information between voices and fuses blocks to generate features of different granularities. Other models such as bidirectional long short-term memory (Bi-LSTM) and time-distributed convolutional neural network (CNN) focus on capturing time dependencies, where CNN captures the features of the previous sound and Bi-LSTM focuses on the time relationship. However, most of these learning models must be trained on a pre-planned dataset, which limits their applicability and robustness. Moreover, these models use a fixed threshold activation function (such as Sigmoid), which is prone to overfitting in a noisy environment; they do not introduce an adaptive neuron mechanism and cannot dynamically adjust the feature weights, resulting in a problem of noise amplification. Since recording devices such as mobile phones or computers are usually non-professional and are affected by environmental noise, this limitation is particularly severe in a home environment.
[0003] For home depression diagnosis, environmental noise will greatly reduce the quality of audio acquisition. Therefore, there is an urgent need for a model with high noise robustness, specifically for home depression diagnosis. According to the "Frontier Report 2022" of the United Nations Environment Programme, the noise levels in most cities around the world exceed the acceptable limits, and low-level residents are often affected by vehicle and human noise. Due to quality problems, audio acquisition devices such as microphones will also introduce electronic noise. It has been recognized that environmental noise and reverberation may reduce the performance of the depression diagnosis model. Deep Residual Shrinkage Networks (DRSN) are used to handle noise problems in time series anomaly detection. However, these methods cannot be directly used to process noisy audio samples. When directly applied to audio data, thresholding may filter out key features of depression (such as fundamental frequency changes), and there is a lack of a cumulative mechanism for time features and cannot capture the dynamic patterns of speech signals.
[0004] The SNN model based on LIF shows satisfactory performance in dealing with environmental noise in audio recognition and retains some time features of the audio. However, one limitation of the standard LIF model is its rigid reset threshold, which cannot adapt to the heterogeneity of individual signal points in the time domain. Summary of the Invention
[0005] To overcome the above-mentioned disadvantages and deficiencies of the prior art, the purpose of the present invention is to provide an audio feature extraction method based on a brain nerve model.
[0006] Based on these existing problems, the present invention has developed a robust brain-inspired audio feature extractor (RBA-FE model) for audio diagnosis of depression. This RBA-FE model integrates a temporal convolutional neural network, multi-head attention, and a bidirectional long short-term memory network, and utilizes the unique advantages of each element to address complex diagnostic challenges. Further, an adaptive rate smooth leaky integrate-and-fire (ARSLIF) neuron model is proposed, which simulates the "cell selectivity" in the auditory cortex and can handle the problems of gate overfitting and noise sensitivity in the LSTM by using an adaptive threshold.
[0007] The purpose of the present invention is achieved through the following technical solutions: An audio feature extraction method based on a brain nerve model, comprising: Obtaining speech signal data, preprocessing the speech signal data, and segmenting it into N speech segments; Extracting acoustic features representing the emotional state from each speech segment, where the acoustic features include pitch, jitter, CQL cepstrum, mel-frequency cepstral coefficients, first and second order differential entropy of mel-frequency cepstral coefficients, and synthesizing a multi-channel speech spectrogram after normalizing the acoustic features; Inputting the multi-channel speech spectrogram into the RBA-FE model, outputting a BDI score, predicting and classifying the degree of depression according to the BDI score, and adjusting the parameters of each network layer in the RBA-FE model according to the difference between the predicted classification and the actual classification to obtain an optimal RBA-FE model for detecting the degree of depression.
[0008] Further, the RBA-FE model adopts a hierarchical structure, sequentially including a temporal convolutional layer, a multi-head attention layer, and a bidirectional long short-term memory layer; The temporal convolutional layer is used to extract the spatio-temporal features of the multi-channel speech spectrogram; The multi-head attention layer is used to perform weighted processing on the spatio-temporal features and output a global feature representation of the speech spectrum. The multi-head attention layer adds a 75% dropout layer; A bidirectional long short-term memory layer is embedded in the ARSLIF neuron model. It captures past and future contextual information in the global feature representation through independent forward LSTM units and backward LSTM units, extracts the temporal dynamics of complex emotions in speech signals, and achieves a comprehensive understanding of speech sequences.
[0009] Furthermore, the ARSLIF neuron model is obtained by introducing adaptive threshold adjustment into the standard LIF model.
[0010] Furthermore, the ARSLIF neuron model is embedded in the bidirectional long short-term memory layer through adaptive threshold adjustment and membrane potential leakage mechanism to replace the Sigmoid Activation function.
[0011] Furthermore, the membrane potential change mechanism of the ARSLIF neuron model is as follows: (1) (2) in, V mem Initial membrane potential, V th Membrane potential threshold, τ mem Membrane potential constant, V mem_new is the updated membrane potential; In the ARSLIF neuron model, the membrane potential changes with the input signal. Inputs When the threshold is exceeded, the neuron fires a pulse signal and the membrane potential is reset, as follows: (3) (4) Where r(t) is the average emissivity, f adap (t) is the adaptive target emissivity, and the expected emissivity is defined as f , the active firing rate of the ARSLIF neuron model f active and the resting firing rate of the ARSLIF neuron model f rest , i represents the number of neurons, R represents the membrane resistance constant, n is the number of neurons, t is the time, V th is the membrane potential threshold.
[0012] Further, the adaptive threshold adjustment is achieved by calculating the error from the desired emission rate and adaptively adjusting the membrane potential threshold. V th , specifically: In the ARSLIF neuron model, V th is not a fixed constant, but is adaptively adjusted based on the error from the desired emission rate: (5) (6) where, τ adapt is the time constant, α and β are adjustable hyperparameters, r e is the rate error, V th_new is the updated new threshold.
[0013] Further, the temporal convolutional layer includes multiple cascaded convolutional layers and pooling layers.
[0014] Further, the specific process of the temporal convolutional layer is as follows: The acoustic features are convolved through a 7×7 convolutional kernel to extract the comprehensive features of the speech features by expanding the receptive field. After batch normalization, ReLU activation, and average pooling, they are input into a residual network with a 3×3 convolutional kernel to further extract the detailed features of the speech signal.
[0015] Further, the mean absolute error and the mean square error are used to evaluate the difference between the predicted classification and the actual classification.
[0016] Further, the multi-head attention layer includes eight attention "heads" arranged in parallel to learn different acoustic features.
[0017] Compared with the prior art, the present invention has the following advantages and beneficial effects: (1) The model architecture of the present invention is superior, enhancing the audio feature extraction ability. The RBA-FE module is proposed and adopts a hierarchical structure. And for ordinary LSTM or CNN, they can only extract local time features and it is difficult to capture long-term dependencies. The present invention adds multi-head self-attention after T-CNN, which can perform global weight allocation on multiple time points of the input audio, emphasizing key information and ignoring irrelevant noise; Combined with the ARSLIF neuron model simultaneously, the attention mechanism not only focuses on the content but also can adapt to the changes in environmental noise, enhancing the generalization ability of the model in different backgrounds. This mechanism improves the time-correlation modeling ability of audio signals, enabling the model to more accurately identify the speech features of depression patients (such as pitch changes, pause patterns, etc.).
[0018] Regarding the traditional CNN that only focuses on local short-term features and the LSTM that has the problem of gradient vanishing in long-time sequence modeling. The present invention adopts a temporal convolutional neural network (T-CNN) to reduce redundant information in the feature extraction stage and improve the computational efficiency. At the same time, combined with a bidirectional LSTM (Bi-LSTM), it uses bidirectional information flow to capture emotional features with a longer time span, improving the emotion recognition ability of the model, enabling the model to learn both short-term and long-term depression speech patterns simultaneously, and improving the accuracy and robustness of classification; (2) The ARSLIF neuron model has strong anti-noise ability and good stability; (3) The present invention has strong generalization ability and has good implementation effects on multiple data sets. Brief Description of the Drawings
[0019] Figure 1 is the overall structural block diagram of the present invention; Figure 2 is the schematic structural diagram of the RBA-FE model of the present invention; Figure 3 is the schematic diagram of the adaptive and threshold adjustment principle of the ARSLIF neuron model of the present invention; Figures 4(a)-4(f) are respectively the schematic diagrams of the performance comparison of the ARSLIF neuron model of the present invention with other activation functions in three noise environments; Figure 5 is the schematic diagram of the ROC curve of different models on the MODMA data set. Detailed Embodiments
[0020] The following combines the embodiments to further elaborate on the present invention in detail, but the implementation manners of the present invention are not limited thereto.
[0021] Embodiment 1 As Figure 1 shown, an audio feature extraction method based on a cranial nerve model includes a feature extractor and an RBA-FE model. After the original speech sample data is preprocessed and acoustical feature extracted by the feature extractor, it is input into the RBA-FE model for depression degree detection and recognition.
[0022] Specific steps are as follows: S1 Obtain speech signal data, preprocess the speech signal, and segment it into N speech segments; In Example 1, the dataset used is the AVEC 2014 depression dataset. The AVEC 2014 depression dataset consists of a total of 300 video samples recorded by subjects in two tasks, namely text reading and Q&A. The lengths vary from 6 seconds to 4 minutes and are evenly divided into a training set, a development set, and a test set, with each set including 100 audio samples.
[0023] The preprocessing mainly includes: First, use the VideoFileClip library to extract the speech information in the video at 22,050Hz, and apply a -30dB clipping filter to remove irrelevant information. The sampling frequency is 1024Hz, the window length is 1024, and there is no overlap. Then, the speech signal is segmented into 32-frame windows. In Example 1, the samples are segmented into 1.5-millisecond speech segments. Samples shorter than the specified length are repeated to reach the segment length, and the excess part is discarded.
[0024] In S2, use the librosa library to extract features from each speech segment, and extract the acoustic features representing the emotional state. The acoustic features include pitch, jitter, CQL cepstrum, Mel-frequency cepstral coefficients (MFCCs), first and second order differential entropy of MFCCs. After normalizing the acoustic features, a multi-channel speech spectrogram of [batch_size, 128, 64, 6] is synthesized.
[0025] Mel-frequency cepstral coefficients (MFCCs) contain important spectral information of the audio. Its first and second order differential entropy reveals the dynamic information of the speech signal; pitch is closely related to the severity of depression; jitter is used to quantify the irregular changes in speech frequency related to severe depression, and the CQL cepstrum improves the analysis of the harmonic structure to reveal the impact of emotion on speech. After normalizing the acoustic features, a multi-channel speech spectrogram is synthesized.
[0026] In S3, input the multi-channel speech spectrogram into the Robust Brain-inspired Audio Feature Extraction Model, that is, the RBA-FE model, for feature extraction and depression level prediction. According to the difference between the predicted classification and the actual classification, adjust the parameters of the RBA-FE model. The parameters of the RBA-FE model include the parameters of each network layer, and the specific parameters are: the convolution kernel of the temporal convolutional layer, the number of attention heads in the multi-head attention layer, the target firing rate of ARSLIF in the bidirectional long short-term memory layer f and the initial membrane potential threshold V th_0 and the active firing rate of ARSLIF f active and the resting firing rate of ARSLIF f rest and the time constant τadapt Together with the hyperparameters α and β and the training parameters (learning rate and batch size) of the model, the optimal RBA-FE model is obtained for detecting the degree of depression. The RBA-FE model outputs the BDI score, and the degree of depression is predicted and classified according to the BDI score.
[0027] This method uses the Beck Depression Inventory (BDI) score to evaluate the degree of depression of the subjects, organizes the data according to the subject ID, and to reduce the potential fluctuations of the BDI score, calculates the average value of each subject in the training set and the test set to reduce the inaccuracy of the labels.
[0028] In this embodiment, MAE (Mean Absolute Error) and MSE (Mean Squared Error) are used to evaluate the error between the predicted value and the actual value.
[0029] The RBA-FE model is inspired by the human brain's audio information processing mechanism, adopts a hierarchical structure, and integrates a temporal convolutional layer, a multi-head attention layer, and a bidirectional long short-term memory layer. The structure of the RBA-FE model is as shown in the appendix Figure 2 as follows.
[0030] Furthermore, research on the auditory pathway shows that sound processing in the brain goes through multiple hierarchical stages, from the extraction of basic acoustic features in the early auditory regions, such as the cochlear nucleus and the superior olivary complex, to complex temporal processing in the cortex. It shows that these stages are associated with specific layers of the deep neural network (DNN) used for audio processing, supporting the use of a hierarchical architecture for robust feature extraction. Inspired by this brain mechanism, the RBA-FE model uses T-CNN for early feature processing, uses multi-head attention to selectively focus on relevant parts of the audio spectrogram, uses Bi-LSTM to integrate information across frames, and uses a softmax layer for classification or a fully connected layer for regression.
[0031] The temporal convolutional layer is used to analyze and extract the temporal dimension information, i.e., spatio-temporal features, of the multi-channel speech spectrogram. The temporal convolutional layer (T-CNN) includes multiple cascaded convolutional layers and pooling layers to extract the local features of the signal. The 6 acoustic features extracted first go through a 7×7 convolutional kernel for convolution, extract the comprehensive features of the speech features by expanding the receptive field, and after batch normalization, ReLU activation, and average pooling, are input into a residual network with a kernel of 3×3 to further extract the detailed features of the speech signal. The temporal convolutional layer (T-CNN) captures the local and global information of the speech signal through two convolutions, and extracts rich temporal features. The output of the T-CNN is concatenated along the temporal dimension and is the basis for the subsequent feature analysis network, simulating the function of the primary visual cortex of the human brain.
[0032] The multi-head attention layer is used to weight the spatio-temporal features extracted by the temporal convolutional layer and output the global feature representation of the speech spectrum. This output contains key emotional information and spatio-temporal dependencies. It amplifies the sparse features of the input and improves the speed at which the subsequent bidirectional long short-term memory layer (Bi-LSTM) converges to the key features.
[0033] The multi-head attention layer includes 8 attention heads. The 8 attention "heads" learn the input data representations in different subspaces (such as pitch, energy, spectral dynamics), enhancing the model's ability to perceive multi-dimensional information; the parallel arrangement can suppress the interference of noise on a single attention mechanism and enhance the robustness of the model. The design of the 8 attention "heads" is the optimal performance setting that balances computational complexity and model expressiveness verified by experiments. A 75% dropout layer is also added in the multi-head attention layer to prevent overfitting, enrich the model's feature interpretation ability, and reduce overfitting. The multi-head attention layer simulates the mechanism of the human brain's auditory system to selectively focus on and process specific audio information.
[0034] The bidirectional long short-term memory layer captures the past and future context information in the global feature vector representation through independent forward LSTM units and backward LSTM units, extracts the temporal dynamics of complex emotions in the speech signal, and realizes a comprehensive understanding of the speech sequence. In addition, for the environmental noise and reverberation existing in the home environment, the Sigmoid activation function is replaced with the ARSLIF neuron model. The ARSLIF neuron model is achieved by introducing an adaptive threshold adjustment mechanism into the standard LIF model. It uses the adaptive threshold adjustment and smoothing leakage mechanism to simulate the "selective readjustment" of human brain nerve cells, enhances the robustness of the model under noise, and significantly improves the accuracy of depression audio diagnosis.
[0035] The ARSLIF neuron model is inspired by the human brain's audio processing mechanism. It introduces the "cell selective readjustment" mechanism into the LSTM unit, dynamically filters noise through adaptive threshold adjustment, improves the model's robustness to noise, and at the same time enhances the key features and improves the accuracy of depressed speech diagnosis. Replacing the traditional Sigmoid gate function in the LSTM with the ARSLIF neuron model, first redesigned the gradient propagation mechanism to be compatible with the backpropagation algorithm to solve the temporal alignment and information matching problems of the continuous information flow in the LSTM unit and the discrete pulses of the ARSLIF neuron; also designed smoothing leakage and parameter constraints to avoid the gradient explosion or disappearance problems caused by the adaptive threshold adjustment mechanism and increase the stability of the model.
[0036] The working principle of the ARSLIF neuron model is as shown in the appendix Figure 3 as follows: (1) (2) The membrane potential change and pulse emission mechanism of the ARSLIF model are shown in Formulas (1) and (2).
[0037] Among them, V mem The initial membrane potential, V th The membrane potential threshold, τ mem The membrane potential constant, V mem_new is the updated membrane potential. In the ARSLIF model, the membrane potential changes with the input signal Inputs and when it exceeds the threshold, the neuron emits a pulse signal and the membrane potential is reset. In the ARSLIF neuron model, there are: (3) (4) In the formula, r(t) is the average firing rate, f adap (t) is the adaptive target firing rate, and the desired firing rate is defined as f .
[0038] Different from the original LIF model, V th in the ARSLIF neuron model is not a fixed constant, but is adaptively adjusted through the error with the desired firing rate: (5) (6) Among them, τ adapt is the time constant, α and β are adjustable hyperparameters, i represents the neuron number, R represents the membrane resistance constant, n is the number of neurons, t is the time, V th is the membrane potential threshold.
[0039] Formulas (5) and (6) show that the membrane potential threshold of the ARSLIF neuron is adaptively adjusted according to the average firing of the neuron at the current moment, enabling the neuron model to adjust the activity level of the neuron according to different input conditions and improving the adaptability of the spiking neuron.
[0040] The ARSLIF neuron model is embedded in the LSTM unit: The ARSLIF neuron model replaces the traditional sigmoidThe gate function maintains the stability of the pulse distribution under noise, retaining only key features, simulating the processing of speech signals by brain neurons, and effectively preventing overfitting caused by noise sensitivity and noise amplification caused by recursion. The mathematical description of the ARSLIF-BiLSTM unit is as follows: The ARSLIF neuron model is embedded in a Bi-LSTM neural network through dynamic threshold adjustment and membrane potential leakage mechanism, achieving efficient feature filtering and robust time series modeling in noisy environments, providing a bio-inspired interpretability solution for depression audio diagnosis.
[0041] The ARSLIF neuron model has strong anti-noise capability and good stability. The noise power of the ARSLIF neuron model is shown in formula (7): (7) In the formula, η(t) is the external noise. Combining formulas (5) and (6), in order to filter the pulses caused by noise, the adaptive threshold adjustment mechanism of the ARSLIF neuron model makes △Vth(t) increase with the increase of noise, and the noise power Compared with the original LIF, ARSLIF has a higher signal-to-noise ratio and better anti-noise performance.
[0042] The performance of the ARSLIF neuron model compared to other gate activation functions in noisy environments is shown in Figures 4(a)-4(f). Four gate activation functions were substituted into the RBA-FE model: sigmoid, soft threshold, standard LIF, and LIF with a desired firing rate. The performance was compared with ARSLIF. Pink noise was used to simulate low-rate ambient noise generated by traffic and electrical appliances; blue and purple noise simulated high-rate noise such as human voices. The simulated noise levels ranged from -20dB to +20dB.
[0043] Figures 4(a)-4(f) show that the ARSLIF model's root mean square error (RMSE) and mean square error (MAE) remain low at various noise levels, with minimal fluctuations, demonstrating excellent robustness. Compared to other gate activation functions, the ARSLIF model exhibits greater adaptability and stability. In particular, ARSLIF outperforms other functions in blue and purple noise environments, similar to those found in home environments. This demonstrates ARSLIF's superior noise immunity and robustness in these environments.
[0044] In this embodiment, the Adam optimizer with a learning rate of 0.0005 is used, and other parameters follow the Keras default values. The loss function is selected as RMSE, and 16-bit mixed precision is used to accelerate training. The batch size is set to 96 to improve training efficiency. During testing, the prediction situations of each subject are averaged to generate the final BDI prediction. The model performance is evaluated by RMSE and MAE metrics, comparing the predicted scores and the actual scores. The optimal parameter settings of the ARSLIF neuron model are shown in Table 1.
[0045] Table 1 Optimal Parameter Settings of the ARSLIF Neuron Model
[0046] Comparison with Other Depression Speech Recognition Methods: In the present invention, an artificial feature extractor and an RBA-FE module are designed to implement depression speech recognition on the AVEC 2014 dataset. As shown in Table 2, the depression speech recognition method adopted in the present invention has better performance than other methods.
[0047] Table 2 Implementation Effect of the Present Technical Solution on the AVEC 2014 Dataset The root mean square error (RMSE) and mean absolute error (MAE) in Table 2 are used to evaluate the error between the predicted value and the actual value. The smaller the index value, the closer it is to the actual value, so it is suitable for evaluating the performance of various algorithms. Table 2 shows that the RMSE value of RBE-FA is lower, while the MAE value is higher. The results show that RBE-FA is more accurate in predicting data with higher depression scores (BDI) than other scheme models, which can ensure that high-risk patients will not be misdiagnosed as low-risk patients; while in terms of the MAE index, the performance of RBE-FA is slightly weaker, and it is slightly worse than other models using regression algorithms in identifying low-risk patients.
[0048] Table 3 Performance of the RBA-FE Model and Its Variants of the Present Technical Solution on the AVEC 2014 Dataset
[0049] Table 3 shows that each network module in RBA-FE plays a role in effectively extracting the speech features for depression diagnosis, and the Bi-LSTM module has the greatest impact on the overall performance.
[0050] Example 2 The dataset selected in this embodiment is different from that in Example 1: In this embodiment, the MODMA audio dataset is selected from the Second Hospital of Lanzhou University in China. There are a total of 52 subjects, including 23 depression patients (16 males and 7 females) and 29 healthy controls (20 males and 9 females). The experiment includes 3 tasks: emotional interview questions, text reading, and image description. The experimental voice is Mandarin, the audio sampling rate is 44.1 kHz, and the experimental duration is 25 minutes. MODMA is a binary classification dataset, and all speech data is divided into a training set and a dataset at a ratio of 8:2. The sample data is preprocessed by an artificial feature extractor and combined with feature extraction to form a [batch_size, 128, 64, 6] speech spectrum, which is used as the input for the subsequent RBA-FE.
[0051] In this embodiment, the RBA-FE module uses an Adam optimizer with a learning rate of 0.001, and the loss function selects categorical cross-entropy. Other parameters are the same as those in Embodiment 1. The model performance is evaluated by precision, accuracy, recall, and F1-score.
[0052] The best parameter settings of the ARSLIF neuron model are shown in Table 4.
[0053] Table 4 Best Parameter Settings of the ARSLIF Neuron Model
[0054] Comparison with other depression speech recognition methods: As shown in Table 5, the depression speech recognition performance of the technical solution of the present invention on the MODMA dataset is better than that of other methods.
[0055] Table 5 Implementation Effect of the Present Technical Solution on the MODMA Dataset
[0056] Table 5 is evaluated by precision, accuracy, recall, and F1-score. The four indicators of the RBA-FE model of this solution are all better than those of other models.
[0057] Appendix Figure 5 is the ROC curve of the RBA-FE model on the MODMA dataset. The closer the curve is to the upper left corner, the better the classification performance of the model; the area under the curve (AUC) is used to measure the overall performance of the model. The closer the AUC is to 1, the better the performance. The AUC of a perfect classifier is 1. The AUC of RBA-FE is 0.89, indicating good classification performance. Appendix Figure 5The AUC value of the RBA-FE model is shown to be 0.89, which is greater than other classifications, indicating that the RBA-FE has the best classification performance.
[0058] Through the introduction of the ARSLIF neuron model, the neural network of the present invention has stronger adaptability to the input audio signal, can dynamically adjust the threshold according to different changes in the signal, thereby effectively filtering noise and retaining useful audio features. In addition, by combining multi-head self-attention and bidirectional LSTM networks, it can accurately extract spatial and temporal features in the audio, greatly improving the robustness and diagnostic accuracy of the model. Especially when performing self-detection in a home environment, it can provide a more accurate judgment of depression.
[0059] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the described embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An audio feature extraction method based on a cranial nerve model, characterized in that, Including: Obtain voice signal data, preprocess the voice signal data, and segment it into N voice segments; Extract acoustic features representing the emotional state from each voice segment. The acoustic features include pitch, jitter, CQL cepstrum, Mel frequency cepstral coefficients, first-order and second-order differential entropy of Mel frequency cepstral coefficients. After normalizing the acoustic features, synthesize a multi-channel voice spectrogram; Input the multi-channel voice spectrogram into the RBA-FE model, output the BDI score, perform depression degree prediction classification based on the BDI score, and adjust the parameters of each network layer in the RBA-FE model according to the difference between the prediction classification and the actual classification to obtain the optimal RBA-FE model for depression degree detection; 2. The audio feature extraction method according to claim 1, wherein The RBA-FE model adopts a hierarchical structure, sequentially including a temporal convolutional layer, a multi-head attention layer, and a bidirectional long short-term memory layer; The temporal convolutional layer is used to extract the spatio-temporal features of the multi-channel voice spectrogram; The multi-head attention layer is used to perform weighted processing on the spatio-temporal features and output the global feature representation of the voice spectrogram. The multi-head attention layer adds a 75% dropout layer; The bidirectional long short-term memory layer embeds the ARSLIF neuron model, captures the past and future context information in the global feature representation through independent forward LSTM units and backward LSTM units, extracts the time dynamics of complex emotions in the voice signal, and realizes a comprehensive understanding of the voice sequence; 3. The audio feature extraction method according to claim 2, wherein The ARSLIF neuron model is obtained by introducing adaptive threshold adjustment into the standard LIF model; 4. The audio feature extraction method according to claim 2, wherein The ARSLIF neuron model is embedded in a bidirectional long short-term memory layer through an adaptive threshold adjustment and a membrane potential leakage mechanism, and is used to replace the Sigmoid activation function in the bidirectional long short-term memory layer.
5. The audio feature extraction method according to any one of claims 1-4, characterized in that The membrane potential change mechanism of the ARSLIF neuron model is specifically as follows: (1) (2) Among them, V mem the initial membrane potential, V th the membrane potential threshold, τ mem the membrane potential constant, V mem_new is the updated membrane potential; In the ARSLIF neuron model, the membrane potential changes with the input signal Inputs and when it exceeds the threshold, the neuron fires a pulse signal and the membrane potential is reset as follows: (3) (4) where r(t) is the average emission rate, f adap (t) is the adaptive target emission rate, and the desired emission rate is defined as f , the active emission rate of the ARSLIF neuron model f active and the resting emission rate of the ARSLIF neuron model f rest , i represents the neuron number, R represents the membrane resistance constant, n is the number of neurons, t is time, V th is the membrane potential threshold.
6. The audio feature extraction method according to claim 4, wherein The adaptive threshold adjustment is to calculate the error from the desired emissivity and adaptively adjust the membrane potential threshold V th , specifically as follows: in the ARSLIF neuron model V th is not a fixed constant, but is adaptively adjusted by the error from the desired firing rate: (5) (6) Among them, τ adapt is the time constant, α and β are adjustable hyperparameters, r e is the rate error, V th_new is the updated new threshold.
7. The audio feature extraction method according to claim 2, wherein The temporal convolutional layer includes multiple cascaded convolutional layers and pooling layers; 8. The audio feature extraction method according to claim 7, wherein The specific process of the temporal convolutional layer is as follows: The acoustic features are convolved through a 7×7 convolutional kernel, and the comprehensive features of the voice features are extracted by expanding the receptive field. After batch normalization, ReLU activation, and average pooling, they are input into a residual network with a 3×3 convolutional kernel to further extract the detailed features of the voice signal; 9. The audio feature extraction method according to claim 1, wherein The mean absolute error and the mean square error are used to evaluate the difference between the prediction classification and the actual classification; 10. The audio feature extraction method according to claim 2, wherein The multi-head attention layer includes eight attention "heads" arranged in parallel to learn different acoustic features;
Citation Information
Patent Citations
Attention mechanism and convolutional neural network-based voice depression recognition method
CN109599129A
Detection method for discriminating depression based on sound
CN111951824A
Voice depression state recognition method based on Attention and Bi-LSTM
CN113571050A
Epilepsy electroencephalogram detection system based on feed-forward pulse neural network
CN115414054A
Depression state auxiliary detection method based on audio dual-mode fusion type neural network
CN115862684A
Cited By
Self-adaptive AI interview method and device, electronic equipment and storage medium
CN120806742A