Speech emotion recognition method based on small sample learning
The DGNN adapter addresses sensitivity to text prompts and noisy voice data in CLAP models by stabilizing cross-modal similarity and label attention, enhancing performance in small-sample voice emotion recognition tasks.
Patent Information
- Application Number
- CN202510429764.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-04
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
AI Technical Summary
In the field of speech emotion recognition, the CLAP model has high sensitivity to text prompt templates (text outliers problem) and a large impact on low-quality speech samples (voice outliers problem) in small sample learning tasks, resulting in a decline in model performance.
Using dual graph neural network (DGNN) adapters, including instance GNN and prototype GNN, the graph structure is constructed, and the prediction consistency is enhanced by utilizing the dot product and bidirectional KL divergence of node feature embedding, mitigating the influence of speech and text outliers, and enhancing the migration ability of the model.
The performance of the CLAP model in small sample speech emotion recognition tasks is improved, the attention to category labels is enhanced, and the overall performance and prediction accuracy of the model are improved.
Smart Images

Figure CN120279948A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech emotion recognition, and in particular relates to a speech emotion recognition method based on few-shot learning. Background Technique
[0002] Emotion is the attitude that a person shows towards external events or conversation activities. Usually, human emotions can be classified into categories such as happiness, anger, sadness, fear, and surprise. By analyzing the collected signals, a machine can judge a person's emotional state, and this process is called emotion recognition. Emotion recognition usually relies on two types of signals: one is physiological signals, such as breathing, heart rate, and body temperature; the other is behavioral signals, including facial expressions, speech, and postures. Among them, speech is often used for emotion recognition because of its convenient collection. Speech emotion recognition has been widely applied in many fields and played an important role. For example, by analyzing the speech of a driver to detect their mental state and reminding them of emotion management to ensure safe driving; by collecting the speech signals of patients through wearable devices to monitor abnormal emotional states in real time and improve the treatment effect; by combining speech emotion information and automatic translation results to help speakers of different languages communicate more smoothly.
[0003] Deep learning methods construct feature learning through multi-layer neural networks. Compared with traditional machine learning methods, they can generate a more advantageous representation space. However, due to the large number of neural network parameters, their effective training usually requires a large amount of data support. In some application scenarios, it may face great challenges to obtain sufficient training data and accurately label it. Therefore, the dependence of deep learning models on a large number of samples, especially labeled samples, has gradually become a bottleneck problem for their wide application. In response to this dilemma, the few-shot learning paradigm has emerged and has been widely applied in classification tasks.
[0004] The core idea of few-shot learning is to use a large dataset that is category-agnostic but has sufficient data volume. By combining experts' understanding of the task, it extracts category-agnostic or partially shareable prior knowledge from it, models it, and then transfers this knowledge to few-shot learning tasks. In addition, few-shot learning shows certain advantages in inferring and generalizing new concepts, laying a solid foundation for the gradual transformation of artificial intelligence technology from a mode that relies on big data to strong artificial intelligence.
[0005] CLAP (Contrastive Language-Audio Pretraining) is a contrastive language-audio pre-training model open-sourced by LAION. This model draws on the idea of CLIP (Contrastive Language-Image Pretraining) and pre-trains with a large amount of audio-text pair data to learn the joint representation of audio and language. The emergence of CLAP has brought a new paradigm to audio understanding tasks and provided strong support for downstream tasks such as audio classification, retrieval, and generation.
[0006] However, when applying the CLAP model to few-shot learning, the following problems still exist:
[0007] (1) Using a single word to replace the language description will affect the transfer ability of few-shot learning. At the same time, the language descriptions generated by different prompt templates will also affect the transfer performance of CLAP. That is to say, CLAP is highly sensitive to the prompt template, and the attention of CLAP to the label will vary due to the text prompt template. This phenomenon is called the "text outlier" problem.
[0008] (2) Few-shot learning tasks in the field of speech emotion recognition rely on high-quality speech samples. Low-cost speech acquisition methods often result in noisy speech data sets and may carry the risk of mislabeling. This noisy speech will have an adverse effect on the transfer ability of CLAP, leading to a decline in model performance. This is called the "speech outlier" problem.
[0009] Graph neural networks (GNNs) are a type of neural network specifically designed to process graph-structured data composed of nodes and edges. The construction of a graph is usually achieved by mapping the samples of a data set to nodes and defining the distance metric between samples as edges. In the scenario of few-shot learning, the number of samples for each class is limited, and it is difficult to fully optimize the model by directly using the samples for training. Graph neural networks achieve information diffusion through the connections between nodes, spreading the information of known samples to the most similar unknown samples to alleviate the problem of insufficient samples in few-shot learning.
[0010] For few-shot learning based on the "n-way k-shot" problem (in the training set of the few-shot learning task, there are many categories, and there are multiple samples in each category. During the training phase, n categories are randomly selected from the training set, and k samples in each category (a total of n*k data) are used to construct a meta-task, which is input as the support set of the model; then a batch of samples are drawn from the remaining data in these n categories as the prediction targets (batch set) of the model. That is, the model is required to learn how to distinguish these n categories from n*k data, and such a task is called the "n-way k-shot" problem), the input data includes the support set and the query set. By optimizing the distance metric between the query set and the support set, the model can be generalized to tasks containing unknown categories. Summary of the Invention
[0011] In view of this, the present invention aims to overcome the deficiencies of the above problems in the prior art, and proposes a voice emotion recognition method based on few-shot learning to solve the "text outliers" and "voice outliers" problems of CLAP in voice emotion classification, reduce the impact of low-quality voice samples and changing text prompt templates on the model, and enhance the performance of CLAP in few-shot learning tasks.
[0012] To achieve the above object, the technical solution of the present invention is realized as follows:
[0013] The first aspect of the present invention provides a voice emotion recognition method based on few-shot learning, including the following steps:
[0014] Step 1: Construct a dual graph neural network adapter, including two parts: instance GNN and prototype GNN. Instance GNN is used to calculate the similarity between voices, and prototype GNN is used to calculate the similarity between voices and texts;
[0015] Step 2: Construct a graph structure based on the few-shot dataset, and input the graph structure into instance GNN and prototype GNN in the DGNN adapter;
[0016] Step 3: Train the dual graph neural network adapter. During the training process, the dot product of two node feature embeddings is used as the edge prediction value, and the node feature embeddings are updated to gradually approach the preset value;
[0017] Step 4: Convert the edge prediction into a class prediction, and use the bidirectional KL divergence to enhance the prediction consistency of the two GNNs during the training process;
[0018] Step 5: In the testing phase, the final prediction result is obtained by calculating the average of the two prediction results.
[0019] Further, in Step 2, a set of input support sets and query sets can be expressed as S = {(x i , y i )|i = 1, …, n×k}, Q = {(x i , y i )|i = 1, …, n×q}, where S is the support set, Q is the query set, q represents the number of samples in each class of the query set, x i represents the feature (voice or text) embedding, y i represents the corresponding label. Based on the above data samples, a graph structure is constructed. Each instance in the form of (x i , y i ) in the data is mapped to a node in the graph. The edge value between nodes i and j is defined as follows:
[0020]
[0021] Further, in Step 3, during the training process, the dot product of the two node feature embeddings is used as the edge prediction value, and the node feature embeddings are updated to gradually approach the preset values. The edge prediction outputs of instance GNN and prototype GNN are in the following forms:
[0022]
[0023]
[0024] Among them, and respectively represent the transfer functions of instance GNN and prototype GNN parameterized by parameters θ and φ. x i represents the voice embedding of the query sample, while x j in the formulas and respectively represent the voice embedding and text embedding of the support samples of IGNN and PGNN;
[0025] The similarity between the query sample embedding and the support sample embedding is processed as the following matrix:
[0026]
[0027] Among them, and are continuous values, and the value ranges are in [0, 1].
[0028] Further, in step 4, the edge prediction is converted into a class prediction through the following formula:
[0029] P I =E I C
[0030] P P =E P C
[0031] where represents a matrix composed of one-hot encodings of the support set sample labels, and the values of E I and E P range from P I and P P range from where c represents the total number of classes in the entire dataset, and each row in P P and P I represents the probability that a sample in the query set belongs to c classes;
[0032] The loss functions corresponding to instance GNN and prototype GNN are defined as follows:
[0033]
[0034]
[0035] Further, in step 4, the bidirectional KL divergence is calculated using the class probability distributions obtained from instance GNN and prototype GNN. The formula is defined as:
[0036]
[0037] The total network loss is defined as:
[0038]
[0039] where the three terms α, β, and γ are the loss function coefficients for instance GNN, prototype GNN, and prediction consistency loss respectively, and are dynamically adjusted during the training process of DGNN. α, β, and γ are dynamically updated according to the following rules:
[0040] (1) Initialization: α (0) =β (0) =γ (0) =1
[0041] (2) Update every training cycle t: Calculate the normalization factor and the weights where η is the temperature coefficient.
[0042] The second aspect of the present invention provides a voice emotion recognition device based on few-shot learning, including:
[0043] A first processing unit for constructing a dual graph neural network adapter, including two parts: an instance GNN and a prototype GNN. The instance GNN is used to calculate the similarity between voices, and the prototype GNN is used to calculate the similarity between voice and text;
[0044] A second processing unit for constructing a graph structure based on a few-shot dataset and inputting the graph structure into the instance GNN and the prototype GNN in the DGNN adapter;
[0045] A third processing unit for training the dual graph neural network adapter. During the training process, the dot product of two node feature embeddings is used as the edge prediction value, and the node feature embeddings are updated to gradually approach the preset value;
[0046] A fourth processing unit for converting edge predictions into class predictions and using bidirectional KL divergence to enhance the prediction consistency of the two GNNs during the training process;
[0047] A fifth processing unit for obtaining the final prediction result by calculating the average of the two prediction results during the test phase.
[0048] The third aspect of the present invention provides an electronic device, including
[0049] At least one processor, and
[0050] At least one memory communicatively connected to the processor, wherein:
[0051] The memory stores program instructions executable by the processor, and the processor can execute the above-mentioned voice emotion recognition method based on few-shot learning by calling the program instructions.
[0052] The fourth aspect of the present invention provides a non-volatile computer-readable storage medium, which, when the computer-executable instructions are executed by one or more processors, enables the processors to execute the above-mentioned voice emotion recognition method based on few-shot learning.
[0053] Compared with the prior art, the voice emotion recognition method based on few-shot learning according to the present invention has the following advantages:
[0054] The present invention proposes a DGNN adapter based on GNN, which efficiently completes the few-shot speech emotion recognition task in combination with CLAP. The dual GNN structure effectively solves the "text outlier" and "speech outlier" problems of CLAP in speech emotion recognition. Specifically, by using a random prompt template in PGNN, the sensitivity of the model to the prompt is reduced, the attention of CLAP to the class label is enhanced, and the KL divergence is used to improve the consistency of the prediction results between the two GNNs, thereby improving the overall performance of the model. DGNN reduces the impact of low-quality speech samples and varying text prompt templates on the model, and enhances the performance of CLAP in the few-shot speech emotion recognition task. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0056] Figure 1 It is a framework diagram of the DGNN adapter of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0058] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0059] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0060] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0061] Embodiment 1:
[0062] The present invention proposes a speech emotion recognition method based on few-shot learning, and designs a dual graph neural network (DGNN) adapter to apply the contrastive language-audio pre-trained model (CLAP) to few-shot learning tasks, so as to reduce the impact of low-quality speech samples and changing text prompt templates on the model, and enhance the performance of CLAP in few-shot speech emotion recognition tasks.
[0063] The DGNN in the present invention consists of two parts: an instance GNN (IGNN) and a prototype GNN (PGNN). Its overall structure takes the support set and the query set as inputs. The support set contains a small number of labeled samples, which are used to help the model learn category features; the query set is the sample to be classified. The IGNN focuses on learning the similarity between speech samples, maps the speech embeddings of the support set and the query set as nodes, constructs a graph structure through an adjacency matrix based on category labels, and uses multi-layer graph convolution to capture the relationships between nodes. Finally, the category probability distribution of the query sample is generated through edge prediction. The PGNN learns the similarity between the speech embeddings of the query set and the text embeddings of the support set, where the text is generated by combining a prompt (randomly selected from 10 different prompt templates) and a speech label. The random prompt template can ensure the stability of cross-modal similarity calculation under different prompt conditions and enhance CLAP's attention to category labels. First, the speech and text embeddings are projected into a unified feature space, then the class prototypes are generated by aggregating the nodes belonging to the same category in the query set, and the edge similarity is calculated using these prototypes and the support samples, and the category probability distribution is output. In order to enhance the consistency of the predictions of the two networks, the DGNN aligns the two distributions through the bidirectional Kullback–Leibler (KL) divergence, so as to achieve the efficient matching of speech and text cross-modal data, promote the consistency of the two-branch predictions and improve the overall performance of the model.
[0064] As Figure 1 shown, the speech emotion recognition method based on few-shot learning of the present invention specifically includes the following steps:
[0065] Step 1:
[0066] As Figure 1 shown, the present invention performs edge prediction by constructing a few-shot learning framework of "n-way k-shot", where n represents the number of categories and k represents the number of samples per category in the support set. A set of input support set and query set can be expressed as S = {(x i , y i ) | i = 1, …, n×k}, Q = {(x i , y i ) | i = 1, …, n×q}, where S is the support set, Q is the query set, q represents the number of samples per class in the query set, x i represents the feature (voice or text) embedding, and y i represents the corresponding label.
[0067] Based on the above data samples, a graph structure is constructed. Each instance in the form of (x i , y i ) in the data is mapped to a node in the graph, and the edge value between nodes i and j is defined as follows:
[0068]
[0069] In other words, the value of the edge represents the probability that two adjacent points belong to the same category.
[0070] Step 2:
[0071] Based on the above graph construction method, the graph structure is input into two GNNs in the DGNN adapter. For the support set S = {(x i , y i ) | i = 1, …, n×k}, the input x i of the IGNN is the voice embedding, while in the PGNN it is the text embedding. This is because the learning tasks of the two GNNs are different. The task of the IGNN is to learn the similarity of the voice embeddings of the samples in the support set S and the query set Q, while the PGNN learns the similarity between the text embedding and the voice embedding. To make the PGNN branch robust to label hints and focus more on the propagation of voice labels, we randomly assign a hint from the hint set D to each sample in the dataset to generate its language description.
[0072] Step 3:
[0073] During the training process, the dot product of the two node feature embeddings is used as the edge prediction value, and the node feature embeddings are updated to gradually approach the preset value. The edge prediction outputs of the two GNNs are in the following form:
[0074]
[0075] Among them, and respectively represent the transfer functions of the IGNN and PGNN parameterized by the parameters θ and φ, and x i represents the speech embedding of the query sample, while x j in the formula and respectively represent the speech embedding and text embedding of the support samples of the IGNN and PGNN. At this time, the similarity between the query sample embedding and the support sample embedding has been calculated, and these similarities are processed into the following matrix:
[0076]
[0077] Among them, and are continuous values, and the value range is [0, 1].
[0078] Step 4:
[0079] Our ultimate goal is to determine the class labels of the samples in the query set. The two GNN branches are responsible for learning the similarities between the query set samples and the support set samples (i.e., the edges in the graph structure) respectively. The edge prediction is converted into class prediction through the following formula:
[0080] P I = E I C
[0081] P P = E P C
[0082] Among them, represents the matrix composed of the one-hot encoding of the support set sample labels. The value ranges of E I and E P are P I and P P are where c represents the total number of classes in the entire dataset. Each row in P P and P I represents the probability that a sample in the query set belongs to c classes. So far, we have obtained the class prediction results of the two GNNs. The loss functions corresponding to the two GNNs can be defined as follows:
[0083]
[0084] In addition, the bidirectional KL divergence is calculated using the class probability distributions obtained from the two GNNs. The formula is defined as:
[0085]
[0086] The total network loss is defined as:
[0087]
[0088] Among them, the three terms α, β, and γ are the loss function coefficients of the instance GNN, prototype GNN, and prediction consistency loss respectively, and are dynamically adjusted during the training process of the DGNN. α, β, and γ are dynamically updated according to the following rules:
[0089] (1) Initialization: α (0) = β (0) = γ (0) = 1;
[0090] (2) Update per training cycle t: Calculate the normalization factor and the weight where η is the temperature coefficient.
[0091] Step Five:
[0092] In the test phase, the final prediction result is obtained by calculating the average of the two prediction results.
[0093] Embodiment Two:
[0094] A speech emotion recognition device based on few-shot learning, comprising:
[0095] A first processing unit for constructing a dual graph neural network adapter, including two parts: an instance GNN and a prototype GNN. The instance GNN is used to calculate the similarity between voices, and the prototype GNN is used to calculate the similarity between voice and text;
[0096] A second processing unit for constructing a graph structure based on a few-shot data set and inputting the graph structure into the instance GNN and prototype GNN in the DGNN adapter;
[0097] A third processing unit for training the dual graph neural network adapter. During the training process, the dot product of two node feature embeddings is used as the edge prediction value, and the node feature embeddings are updated to gradually approach a preset value;
[0098] A fourth processing unit for converting the edge prediction into a class prediction and using the bidirectional KL divergence to enhance the prediction consistency of the instance GNN and prototype GNN during the training process;
[0099] A fifth processing unit, configured to obtain a final prediction result by calculating an average value of two prediction results during a test phase.
[0100] Embodiment III:
[0101] An electronic device, comprising
[0102] at least one processor, and
[0103] at least one memory communicatively connected to the processor, wherein:
[0104] The memory stores program instructions executable by the processor, and the processor can execute the above-mentioned voice emotion recognition method based on few-shot learning by invoking the program instructions.
[0105] Embodiment IV:
[0106] A non-volatile computer-readable storage medium, when the computer-executable instructions are executed by one or more processors, enables the processors to execute the above-mentioned voice emotion recognition method based on few-shot learning.
[0107] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for speech emotion recognition based on few-shot learning, characterized in that: It includes the following steps: Step 1: Construct a dual graph neural network adapter, which includes two parts: instance GNN and prototype GNN. Instance GNN is used to calculate the similarity between voices, and prototype GNN is used to calculate the similarity between voice and text; Step 2: Construct a graph structure based on a small sample dataset, and input the graph structure into instance GNN and prototype GNN in the DGNN adapter; Step 3: Train the dual graph neural network adapter. During the training process, the dot product of two node feature embeddings is used as the edge prediction value, and the node feature embeddings are updated to gradually approach the preset value; Step 4: Convert the edge prediction to a class prediction, and use the bidirectional KL divergence to enhance the prediction consistency of instance GNN and prototype GNN during the training process; Step 5: In the test phase, obtain the final prediction result by calculating the average of the two prediction results.
2. The method for speech emotion recognition based on few-shot learning according to claim 1, wherein: In step 2, a set of input support sets and query sets can be expressed as S = {(x i , y i ) | i = 1, …, n×k}, Q = {(x i , y i ) | i = 1, …, n×q}, where S is the support set, Q is the query set, q represents the number of samples in each class of the query set, x i represents the feature embedding, y i represents the corresponding label. Based on the above data samples, a graph structure is constructed. Each instance in the data in the form of (x i , y i ) is mapped to a node in the graph. The edge value between nodes i and j is defined as follows:
3. The method for speech emotion recognition based on few-shot learning according to claim 1, wherein: In Step 3, during the training process, the dot product of two node feature embeddings is used as the edge prediction value, and the node feature embeddings are updated to gradually approach the preset value. The edge prediction output forms of instance GNN and prototype GNN are as follows: Among them, and represent the transfer functions of the instance GNN and the prototype GNN parameterized by the parameters θ and φ, respectively. x i represents the speech embedding of the query sample, while x j in the formulas and represent the speech embedding and text embedding of the support samples of the instance GNN and the prototype GNN, respectively. The text is composed of a prompt template randomly selected from 10 different prompt templates and a speech label; The similarity between the query sample embedding and the support sample embedding is processed into the following matrix: Among them, and are continuous values, and the value range is [0, 1].
4. A method for speech emotion recognition based on few-shot learning according to claim 3, characterized in that: In Step 4, the edge prediction is converted to a class prediction through the following formula: P I = E I C P P = E P C Among them, represents a matrix composed of one-hot encodings of the support set sample labels. E I and E P The value range of P I and P P The value range of where c represents the total number of categories in the entire dataset, and each row in P P and P I represents the probability that a sample in the query set belongs to one of the c categories; The loss functions corresponding to instance GNN and prototype GNN are defined as follows:
5. A method for speech emotion recognition based on few-shot learning according to claim 4, characterized in that: In Step 4, the bidirectional KL divergence is calculated using the class probability distributions obtained from instance GNN and prototype GNN. The formula is defined as: The total network loss is defined as: Among them, the three terms α, β, and γ are respectively the loss function coefficients of instance GNN, prototype GNN, and prediction consistency loss, and are dynamically adjusted during the training process of DGNN. α, β, and γ are dynamically updated according to the following rules: (1) Initialization: α (0) = β (0) = γ (0) = 1 (2) Update per training cycle t: Calculate the normalization factor and weight Where η is the temperature coefficient.
6. A voice emotion recognition device based on few-shot learning, characterized in that: It includes: The first processing unit is used to construct a dual graph neural network adapter, which includes two parts: instance GNN and prototype GNN. Instance GNN is used to calculate the similarity between voices, and prototype GNN is used to calculate the similarity between voice and text; The second processing unit is used to construct a graph structure based on a small sample dataset, and input the graph structure into instance GNN and prototype GNN in the DGNN adapter; The third processing unit is used to train the dual graph neural network adapter. During the training process, the dot product of two node feature embeddings is used as the edge prediction value, and the node feature embeddings are updated to gradually approach the preset value; A fourth processing unit, configured to convert edge prediction into class prediction, and enhance the prediction consistency between the instance GNN and the prototype GNN using bidirectional KL divergence during the training process; A fifth processing unit, configured to obtain a final prediction result by calculating an average of two prediction results during the test phase.
7. An electronic device, characterized in that: Comprising At least one processor, and At least one memory communicatively connected to the processor, wherein: The memory stores program instructions executable by the processor, and the processor can execute a method for speech emotion recognition based on few-shot learning according to any one of claims 1-5 by invoking the program instructions.
8. A non-volatile computer-readable storage medium, when the computer-executable instructions are executed by one or more processors, enabling the processors to execute a method for speech emotion recognition based on few-shot learning according to any one of claims 1-5.