Multitask speech emotion recognition method based on parallel processing hybrid expert network
By parallel processing hybrid expert networks, the problems of insufficient feature processing capabilities, training efficiency and task adaptability of the speech emotion recognition model are solved, and more efficient multi-task joint optimization and recognition performance improvement are achieved.
Patent Information
- Application Number
- CN202510980709.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-17
AI Technical Summary
The existing speech emotion recognition model under the multi-task learning framework has deficiencies in feature processing capabilities, model training efficiency and task adaptability, resulting in low recognition performance and accuracy.
A parallel processing hybrid expert network is adopted to perform sentiment classification and speaker recognition in parallel by sharing features. The speech recognition task is combined with the dynamic routing mechanism and multi-task joint optimization of the hybrid expert network to optimize the model parameters.
It improves the model's feature processing capabilities and training efficiency, enhances its adaptability to different tasks, and significantly improves the performance and accuracy of speech emotion recognition.
Smart Images

Figure CN120808822A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of speech recognition, and more particularly relates to a multi-task speech emotion recognition method based on a parallel processing hybrid expert network. BACKGROUND
[0002] At present, in the field of speech emotion recognition, a related technical project adopts a multi-task learning framework MTL, takes speech emotion recognition SER as a main task, and takes automatic speech recognition ASR and speaker recognition SR as auxiliary tasks. After inputting original audio, shared features are first extracted, and then input into SER, ASR and SR task networks for processing, and finally model parameters are optimized through a joint loss function.
[0003] The existing speech emotion model based on multi-task learning can better realize emotion analysis, but still faces multi-dimensional technical bottlenecks in actual deployment, such as reduced training convergence speed caused by multi-task gradient conflict, and the like, which are specifically embodied in the following aspects: limited feature processing capability, low model training efficiency and insufficient task adaptability, thus resulting in poor recognition performance and low accuracy of speech emotion.
[0004] Therefore, how to improve the recognition performance and accuracy of speech emotion is a technical problem to be solved at present. SUMMARY
[0005] In view of the defects of the prior art, the purpose of the present application is to provide a multi-task speech emotion recognition method based on a parallel processing hybrid expert network, aiming at solving the problems of poor recognition performance and low accuracy of speech emotion.
[0006] To achieve the above-mentioned purpose, in a first aspect, the present application provides a multi-task speech emotion recognition method based on a parallel processing hybrid expert network, comprising: performing feature extraction on input original speech to obtain shared features; constructing a parallel processing hybrid expert network, inputting the shared features into the parallel processing hybrid expert network, performing emotion classification and speaker recognition on the shared features in parallel, and outputting predicted emotion values and predicted speakers respectively; inputting the shared features into a full connection layer for speech recognition processing to obtain text prediction values; respectively according to the predicted emotion values, predicted speakers and text prediction values, obtaining loss functions combined with corresponding labels, and optimizing model parameters according to the loss functions to obtain an optimized parallel processing hybrid expert network after training.
[0007] Optionally, the specific process of emotion classification comprises: inputting the shared feature into a sentiment classification head of the parallel processing hybrid expert network; selecting a first expert network from a sentiment expert pool according to the shared feature to activate; inputting the shared feature into the selected first expert network to extract a first feature representation related to sentiment; generating a first gating weight of the selected expert network by using a noise injection and a Top-k selection mechanism; performing weighted summation on the first feature representation and the corresponding first gating weight to obtain the predicted sentiment value as a final output.
[0008] Optionally, the specific process of speaker recognition includes: inputting the shared feature into a speaker recognition classification head of the parallel processing hybrid expert network; selecting a second expert network from a speaker recognition expert pool according to the shared feature to activate; inputting the shared feature into the selected second expert network to extract a second feature representation related to speaker; generating a second gating weight of the selected expert network by using a noise injection and a Top-k selection mechanism; performing weighted summation on the second feature representation and the corresponding second gating weight to obtain a hierarchical decoupled voiceprint feature, and completing speaker recognition based on the voiceprint feature.
[0009] Optionally, the first expert network and / or the first expert network are obtained by screening through a gating mechanism of the parallel processing hybrid expert network.
[0010] Optionally, the implementation process of the gating mechanism includes: performing linear projection on the input shared feature to obtain an initial logical value; injecting Gaussian noise and noise dimension segmentation into the initial logical value, and constraining by a softplus function to obtain a final noise logical value; performing a softmax operation on the noise logical value to obtain an expert probability distribution; selecting Top-k expert networks and their corresponding gating weights in the expert probability distribution to generate a sparse gating matrix.
[0011] Optionally, the parallel processing hybrid expert network realizes expert load balancing through a coefficient of variation loss, an expert utilization rate loss, and a numerical stability loss: The implementation process of the expert load balancing includes: determining a gating value sum, a gating value variance and a gating value mean based on the first gating weight and the second gating weight, and constructing the coefficient of variation loss in combination with a set minimum constant; determine the expert network selection probability according to the probability distribution of each expert network, determine the network use frequency according to the number of times each expert network is selected, and construct the expert utilization loss in combination with the total number of expert networks; construct the numerical stability loss according to the noise logic value corresponding to each expert network; Combine the coefficient of variation loss, the expert utilization loss, and the numerical stability loss with the predefined weight coefficient to obtain the final auxiliary loss for model training.
[0012] Optionally, the first expert network and the second expert network each include a first linear layer, a GELU activation function layer, and a second linear layer. The first linear layer is configured to independently linearly transform the input shared features and output hidden layer features. The GELU activation function layer is configured to apply a Gaussian error linear unit activation to the hidden layer features. The second linear layer is configured to map the activated features back to an output space matching the dimension of the input shared features.
[0013] In a second aspect, the present application also provides a multi-task speech emotion recognition system based on a parallel processing hybrid expert network, comprising: The feature extraction module is configured to extract features from input raw speech to obtain shared features. The parallel processing module constructs a parallel processing hybrid expert network, inputs the shared features into the parallel processing hybrid expert network, performs emotion classification and speaker recognition on the shared features in parallel, and outputs predicted emotion values and predicted speakers, respectively. The speech recognition module is configured to input the shared features into a fully connected layer for speech recognition processing to obtain text prediction values. The optimization module is configured to obtain a loss function according to the predicted emotion values, predicted speakers, and text prediction values in combination with corresponding labels, and optimize model parameters according to the loss function to obtain an optimized parallel processing hybrid expert network.
[0014] In a third aspect, the present application provides an electronic device, comprising at least one memory for storing a program, and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation manner of the first aspect.
[0016] In a fifth aspect, the present application provides a computer program product, which, when running on a processor, causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0017] It can be understood that the beneficial effects of the above-mentioned second aspect to the fifth aspect can be referred to the related description in the first aspect, which will not be repeated here.
[0018] Overall, compared with the prior art, the above technical solutions conceived by the present application have the following beneficial effects: (1) The present application constructs a parallel processing hybrid expert network, which inputs shared features into emotion classification and speaker recognition tasks at the same time, effectively solves the limitations of single model in feature processing capability, so that the model can more comprehensively capture multi-dimensional information in the speech signal; at the same time, through the shared feature input and parallel processing mechanism, the efficiency of model training is significantly improved, and repeated calculation and parameter redundancy are reduced; further, the emotion classification, speaker recognition and speech recognition tasks are organically combined, the adaptability of the model to different tasks is enhanced, and through multi-task joint optimization, the performance of each task is synergistically improved, and finally the performance and accuracy of speech emotion recognition are significantly improved.
[0019] (2) The present application realizes efficient use of computing resources through a multi-task joint optimization mechanism. The emotion classification, speaker recognition and speech recognition three tasks share the bottom feature extraction network, which reduces the parameter quantity and computing cost. The loss functions of each task are synergistically optimized, so that the model can complete multi-task learning through one forward propagation, which significantly improves the training efficiency.
[0020] (3) The hybrid expert network of the present application can flexibly adjust the number and structure of experts according to different task requirements, such as setting different hidden layer dimensions for emotion classification and speaker recognition. Through the customizable architecture, not only can it adapt to diversified speech processing requirements, but also can automatically optimize task allocation through the competition mechanism between experts, so as to maintain stable recognition performance in complex scenarios, providing a general and efficient solution for multi-task speech analysis. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 is a schematic diagram of the overall framework of multi-task learning of the prior art; Figure 2 is a flowchart of speech recognition of the prior art; Figure 3 is one of the flowcharts of the multi-task speech emotion recognition method based on parallel processing hybrid expert network provided by the embodiment of the present application; Figure 4Figure 2 is a flowchart of a method for multi-task speech emotion recognition based on a parallel processing hybrid expert network according to an embodiment of the present application; Figure 5 Figure 2 is a flowchart of a method for multi-task speech emotion recognition based on a parallel processing hybrid expert network according to an embodiment of the present application; Figure 6 Figure 3 is a structural diagram of a device for multi-task speech emotion recognition based on a parallel processing hybrid expert network according to an embodiment of the present application; Figure 7 Figure 4 is a structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0023] The term "and / or" used herein is used to describe an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone. The symbol " / " in this paper represents the relationship of or, for example, A / B represents A or B.
[0024] The terms "first" and "second" and the like in the specification and claims herein are used to distinguish different objects, and are not used to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, and are not used to describe a specific order of the response messages.
[0025] In the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0026] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more, for example, a plurality of processing units means two or more processing units, and the like; a plurality of elements means two or more elements, and the like.
[0027] Currently, the related projects of speech emotion recognition use the multi-task learning framework MTL, taking speech emotion recognition (SER) as the main task and automatic speech recognition (ASR) and speaker recognition (SR) as auxiliary tasks. After inputting the original audio, shared features are first extracted, and then input into the SER, ASR, and SR task networks for processing. Finally, the model parameters are optimized through a joint loss function.
[0028] The specific process of the task execution flow is as follows Figure 1
[0029] SER task: Average pooling is performed on the shared features to reduce the dimensionality of the acoustic features, and then a fully connected layer is used for emotion label classification. Finally, the predicted emotion label is obtained through the argmax layer.
[0030] SR task: Similar to SER, the SR task also uses an average pooling layer to process the shared features, and then maps them to the speaker label sample space through a fully connected layer. Finally, the predicted speaker label is obtained through the argmax layer, which supplements the emotion information that the SER task failed to extract.
[0031] Specifically, the following processes are involved: Average Pooling: The input signal first passes through the average pooling layer, which is usually used to reduce the dimensionality of the data and extract global information.
[0032] Fully Connected Layers (FC1 and FC2): The data after average pooling enters two fully connected layers (FC1 and FC2). These two fully connected layers may be used to learn more complex feature representations.
[0033] Argmax Layer: The argmax layer is used to select the maximum value, which may be used to determine the most likely class or feature.
[0034] Emotion and Speaker Tasks: Emotion Task: This part is responsible for identifying the emotional state of the input signal, such as "happy", "angry", "sad", or "neutral". This usually involves a classification task, where each emotion corresponds to a class.
[0035] Speaker Recognition Task (SR Task): This part is responsible for identifying the source speaker of the signal, such as "speaker 1" and "speaker 2" in the example shown in the figure.
[0036] Output: The final output is divided into two parts, one is the result of emotion recognition, and the other is the result of speaker recognition. Each task has corresponding output categories, emotion recognition has four categories (happy, angry, sad, neutral), while speaker recognition may have multiple categories, depending on how many known speakers there are.
[0037] ASR Task: The goal is to convert the input audio into English text. Using the advantages of pre-trained wav2vec2.0 in the ASR field, shared features are input into a fully connected layer, and then the final text prediction value is obtained through the Argmax layer, providing text information support for emotion recognition. The specific process is as follows Figure 2 as shown.
[0038] FC3 continues to process data through a series of stacked fully connected layers, gradually learning and extracting high-level features.
[0039] Argmax Layer: After these fully connected layers is a layer labeled "Argmax". The Argmax layer is used to select the element with the highest probability in the output.
[0040] Text Output: After the Argmax layer is the text output part, which shows "Yeah, It is a good idea!". This indicates that the model has successfully converted the input audio signal into text form.
[0041] However, the current multi-task learning-based speech emotion model, although it can better achieve emotion analysis, still faces multi-dimensional technical bottlenecks in actual deployment, such as multi-task gradient conflict leading to reduced training convergence speed, etc. The following are the main shortcomings: (1) Limited feature processing capability: In the original technology, the speaker recognition and speech emotion recognition task network uses ordinary fully connected layers, which have limited ability to process complex speech features. Speech signals contain a wealth of information, and fully connected layers are difficult to fully exploit deep and diverse feature relationships. For example, when dealing with speech from different speakers with similar emotions, it may not be able to accurately distinguish the speaker's identity and recognize emotions, resulting in limited accuracy.
[0042] (2) Low model training efficiency: In terms of training time, the original technology takes more than 24 hours to run 100 epochs, which is a high time cost. This is because the parameter structure of the fully connected layer is relatively fixed and cannot adapt to the complex distribution of speech data during the learning process, resulting in slow model convergence speed and the need for more training time to optimize parameters to achieve better performance.
[0043] (3) Insufficient task adaptability: The ordinary fully connected layer has single adaptability for different tasks and cannot perform targeted feature processing and learning according to the unique needs of speaker recognition and speech emotion recognition tasks. Each task has its special information processing focus, and the fully connected layer cannot effectively highlight these differences, affecting the performance of the model on each task.
[0044] The embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0045] Referring to Figure 3 , Figure 3 is a general flowchart of a multi-task speech emotion recognition method based on a parallel processing hybrid expert network provided by the present application, which includes: S101. Extracting features from the input original speech to obtain shared features; S102. Constructing a parallel processing hybrid expert network, inputting the shared features into the parallel processing hybrid expert network, and performing emotion classification and speaker recognition on the shared features in parallel to output predicted emotion values and predicted speakers, respectively; S103. Inputting the shared features into the fully connected layer for speech recognition processing to obtain text prediction values; S104. Obtaining loss functions according to the predicted emotion values, predicted speakers, and text prediction values, respectively, in combination with the corresponding labels, and optimizing the model parameters according to the loss functions to obtain an optimized parallel processing hybrid expert network after training.
[0046] The application first extracts features from the input original speech through S101 to obtain shared features, avoids redundant calculation of repeated feature extraction for different tasks, and improves feature utilization efficiency. In S102, a parallel processing hybrid expert network is constructed, and the shared features are input into the emotion classification and speaker recognition modules at the same time. The dynamic routing mechanism of the hybrid expert network enables the model to adaptively select the most suitable expert to process different tasks, thereby enhancing the feature processing capability and avoiding the problem of insufficient adaptability of a single model to different tasks. S103 further inputs the shared features into a fully connected layer for speech recognition processing, realizes multi-task joint learning, and enables the model to optimize emotion classification, speaker recognition and speech recognition tasks simultaneously during the training process, thereby improving training efficiency and promoting feature sharing. Finally, S104 adjusts the model parameters by calculating the loss function of each part respectively and combining the joint optimization strategy, so that different tasks promote each other and improve the accuracy and robustness of speech emotion recognition.
[0047] It should be noted that there is no execution sequence between step S102 and step S103 in the present application, i.e., the emotion recognition, speaker recognition and speech recognition are parallel implementation processes.
[0048] Specifically, first, through step S101, the input original speech signal is feature extracted to generate shared bottom-layer speech representation. These shared features contain both semantic information of the speech for speech recognition and acoustic characteristics related to speaker identity and emotion, providing a unified input representation for subsequent multi-task processing, avoiding repeated calculation and improving feature utilization efficiency.
[0049] It should be noted that the present application adopts a Wav2Vec2 self-supervised pre-training model as a front-end feature extractor, which extracts speech common features with strong representation capability through multi-layer convolution downsampling and Transformer encoding on the original speech waveform. Wav2Vec2 models the general acoustic characteristics of speech through a contrastive learning task in the pre-training phase, and the output context-related features contain multiple information such as speech content, speaker identity and emotion.
[0050] Then, through step S102, the shared features are input into a parallel hybrid expert network (MoE), which is composed of two independent expert modules focusing on emotion classification and speaker recognition. The emotion classification expert models emotion-related nonlinear features through two layers of linear transformation and GELU activation function; the speaker recognition expert extracts speaker identity features using a similar structure. The expert modules are activated through a gating mechanism to ensure task-specific processing. Finally, the two branches output emotion prediction results (such as happy, angry, etc.) and speaker prediction results (such as speaker ID) in parallel, realizing efficient collaborative reasoning of multi-task.
[0051] Further, the shared features in step S103 are input to a fully connected layer for speech recognition task processing. This module decodes the time-series features through stacked fully connected layers or recurrent neural networks (RNN) to map the acoustic features to a character probability distribution (e.g., phonemes or text), and outputs a character prediction sequence. This branch shares the underlying features with the emotion and speaker tasks, but implements a dedicated decoding of semantic information through an independent task head, forming a multi-task joint learning framework.
[0052] Finally, through step S104, the model is optimized end-to-end by minimizing the weighted loss function of emotion classification, speaker recognition, and speech recognition based on the prediction results and corresponding labels. The design of the loss function balances the importance of different tasks, avoiding the dominance of a single task in gradient updates. During backpropagation, the parameters of the shared feature extraction network and the expert modules are adjusted simultaneously, prompting the model to learn representations that balance generality and task-specificity. This joint training mechanism not only improves the performance of each task, but also significantly improves training efficiency through parameter sharing and parallel computing, adapting to the needs of complex speech analysis scenarios.
[0053] Referring to Figure 4 , Figure 4 is a basic framework diagram of the embodiments of the present application, including the following steps: The input raw speech is first subjected to feature extraction to obtain shared features.
[0054] Parallel processing of hybrid expert network: The shared features are input into two independent ParaK-MoE modules, which work in parallel. The first ParaK-MoE module is responsible for emotion classification and outputs emotion prediction values. The second ParaK-MoE module is responsible for speaker recognition and outputs speaker prediction values.
[0055] Emotion classification: The emotion classification module accepts shared features, processes them through ParaK-MoE, and outputs predicted emotions. The loss between the actual emotion label and the predicted emotion is calculated Lcls .
[0056] Speaker recognition: The speaker recognition module also accepts shared features, processes them through ParaK-MoE, and outputs predicted speakers. The loss between the actual speaker label and the predicted speaker is calculated Lspc .
[0057] Speech recognition: After average pooling, the shared features are input into a fully connected layer (FC) for speech recognition processing. The fully connected layer outputs text prediction values. The loss between the text prediction values and the true text is calculated Lctc .
[0058] Referring to Figure 5 ,Figure 5 is a structural schematic diagram of parallel processing by the application using shared features, comprising: performing Wav2vec feature extraction on the original speech to obtain shared features, which are used for a CTC classification head, an emotion MoE classification head, and a speaker MoE recognition head, respectively.
[0059] Optionally, the specific process of the emotion classification comprises: inputting the shared features into an emotion classification head of the parallel processing hybrid expert network; selecting a first expert network from an emotion expert pool according to the shared features to activate; inputting the shared features into the selected first expert network to extract emotion-related first feature representations; generating first gating weights of the selected expert network using a noise injection and Top-k selection mechanism; performing weighted summation on the first feature representations and the corresponding first gating weights to obtain the predicted emotion value as the final output.
[0060] Optionally, the specific process of the speaker recognition comprises: inputting the shared features into a speaker recognition classification head of the parallel processing hybrid expert network; selecting a second expert network from a speaker recognition expert pool according to the shared features to activate; inputting the shared features into the selected second expert network to extract speaker-related second feature representations; generating second gating weights of the selected expert network using a noise injection and Top-k selection mechanism; performing weighted summation on the second feature representations and the corresponding second gating weights to obtain a hierarchical decoupled voiceprint feature, and completing speaker recognition based on the voiceprint feature.
[0061] On the one hand, in the emotion classification branch, the shared features are first input into the emotion classification head, and the most relevant first expert network is dynamically screened from the emotion expert pool through a learnable routing mechanism. The selected expert network performs nonlinear transformation on the shared features to extract emotion-specific feature representations (such as prosody and intonation-related features). To enhance robustness, controllable Gaussian noise is introduced in the gating weight calculation, and a Top-k sparsification strategy is adopted, retaining only the top-k experts for calculation. Finally, the output features of each activated expert are aggregated through weighted summation to generate emotion label prediction. This design effectively alleviates overfitting through expert selection and noise injection, while the Top-k mechanism ensures computational efficiency, enabling the model to adaptively focus on the most discriminative emotion features.
[0062] The emotion classification task is used to capture fine-grained emotion patterns. In view of the high-dimensional sparse characteristics of the features for the emotion classification task, the system adopts a dynamic expert hybrid network architecture to replace the traditional fully connected layer. Through the synergistic mechanisms such as expert number adaptation, double expert dynamic activation and gate noise injection, fine-grained emotion feature decoupling is realized. The core parameter design logic is shown in Table 1. Table 1 Parameter design of emotion classification head
[0063] On the other hand, the speaker recognition branch adopts a similar expert routing architecture, but the expert pool is dedicated to voiceprint feature extraction. After the shared features are routed to the second expert network through the speaker classification head, the expert network extracts hierarchical voiceprint features (such as speaker embedding vectors). The gating weights are also generated through noise injection and Top-k mechanism to improve generalization. Unlike emotion classification, this branch outputs decoupled voiceprint features, which can be directly used for speaker verification or clustering. This design takes advantage of the fine-grained feature decoupling capability of the expert network to realize separate modeling of speaker identity and emotion features in a unified framework, avoiding interference between tasks. Meanwhile, the noise and sparsification strategies enhance the discriminability and stability of voiceprint features.
[0064] Compared with the 4 experts of emotion classification, speaker recognition needs more experts to cover different frequency bands or acoustic features. The speaker recognition task needs to capture fine-grained voiceprint features such as fundamental frequency, formant trajectory, pronunciation habit, etc. These features have characteristics such as multi-scale, high frequency sensitivity and high cross-channel robustness. In view of the wide frequency domain distribution and significant individual differences of voiceprint features, the system adopts a dynamic expert hybrid network architecture to replace the traditional fully connected layer. Through the synergistic strategies such as expert division of labor, dynamic gating weight and noise-resistant gating mechanism, hierarchical voiceprint feature decoupling is realized. The core parameter design logic is shown in Table 2. Table 2 Speaker recognition classification head
[0065] Optionally, the first expert network and / or the first expert network are obtained by screening through the gating mechanism of the parallel processing hybrid expert network.
[0066] The implementation process of the gating mechanism includes: Linearly projecting the input shared features to obtain initial logical values; Injecting Gaussian noise and noise dimension segmentation into the initial logical values, and constraining them through a softplus function to obtain final noise logical values; Performing a softmax operation on the noise logical values to obtain an expert probability distribution; The Top-k expert networks in the expert probability distribution and their corresponding gating weights are selected to generate a sparse gating matrix.
[0067] Specifically, the MoE model processes input by dynamically selecting expert networks, and the core idea is to assign input data to multiple expert networks and combine expert outputs through a gating mechanism.
[0068] The gating mechanism is as follows: Calculate which experts are selected for each sample, which involves the implementation of noise gating, which increases exploration by adding noise. Calculate logits, apply softmax to get probability distribution, then select top-k experts and calculate the corresponding gating values.
[0069] (1) Calculate the initial logic value logits linear projection, apply the gating matrix to the input feature, and set the bias to false. The formula is as follows:
[0070] wherein, is the gating matrix, containing a noise parameter; (2) Noise injection, the formula is as follows:
[0071]
[0072] Split into two tensors along the last dimension, clean_logits represents the expert selection logic value without noise in the first N dimensions, represents the last N dimensions, represents the noise logic value. represents the noise standard deviation, and a small value is added to avoid zero standard deviation, is the noise term, with the same shape as clean_logits, ; (3) Probability distribution, apply softmax operation to the final logits to get probability distribution probs; (4) Top-k selection: keep the top-k largest probability values and the corresponding expert index;
[0073] wherein, are the top-k largest probability values, are the top-k largest probability values corresponding to the expert index; k is the number of selections; (5) Generate a sparse gating matrix, which has non-zero values only in the first k positions, which is used for weighted summation with the output of the expert in the subsequent process.
[0074] Optionally, the parallel processing hybrid expert network realizes expert load balancing through a coefficient of variation loss, an expert utilization rate loss, and a numerical stability loss: The implementation process of the expert load balancing includes: Determine the sum of the gating values, the variance of the gating values, and the mean of the gating values based on the first gating weight and the second gating weight, and construct the coefficient of variation loss in combination with a set minimum constant; Determine the expert network selection probability according to the probability distribution of each expert network, determine the network usage frequency according to the number of times each expert network is selected, and construct the expert utilization rate loss in combination with the total number of expert networks; Construct the numerical stability loss according to the noise logic value corresponding to each expert network; Combine the coefficient of variation loss, the expert utilization rate loss, and the numerical stability loss with the predefined weight coefficient to obtain the final auxiliary loss for model training.
[0075] Specifically, to prevent expert utilization imbalance and improve model robustness, the following three loss functions are used to realize load balancing.
[0076] (1) Coefficient of variation loss (CV Loss) Balance the gating value distribution of each expert to prevent some experts from being excessively activated. Perform L1 normalization on the accumulated gating values. Calculate the square ratio of the variance and the mean of the normalized distribution.
[0077] The formula is as follows:
[0078] Where g is the sum of the gating values of each expert, and ε is a small constant to prevent division by zero error.
[0079] (2) Expert utilization rate loss (Switch Loss) Constrain the expert selection distribution, align the usage frequency of the expert with its cumulative probability distribution, and realize load balancing. The formula is as follows:
[0080] Where prob is the probability distribution of each expert, freq is the number of times each expert is selected, is the expert selection probability, represents the usage frequency.
[0081] (3) Numerical stability loss (Z Loss) Stabilize the numerical distribution of the gating logits to avoid high-frequency experts being suppressed due to gradient sparsity. The formula is as follows:
[0082] The weighted sum of the three losses constitutes a final auxiliary loss for model training:
[0083] wherein, , , is a predefined weight coefficient.
[0084] Optionally, the first expert network and the second expert network each comprise a first linear layer, a GELU activation function layer, and a second linear layer. The first linear layer is configured to independently linearly transform the input shared feature and output a hidden layer feature. The GELU activation function layer is configured to apply a Gaussian Error Linear Unit activation to the hidden layer feature. The second linear layer is configured to map the activated feature back to an output space matching the dimension of the input shared feature.
[0085] Specifically, in the embodiment, each expert is composed of two linear layers with a GELU activation function in between. According to the expert index selected by the gating network, the indexed expert can collect the corresponding sample from the extracted shared feature, perform expert input transformation, apply a GELU activation function in the middle, and then obtain the output value of each expert through expert output transformation. Finally, the original input batch and sequence dimensions are restored by weighted summation with the gating weight. The detailed process is shown below: (1) Screening expert input According to the batch_index (original sample index) output by the gating mechanism, the selected sample is screened from the input:
[0086] wherein, is the input shared feature, is the selected sample input, is the original sample index.
[0087] (2) Expert network forward calculation Each expert is an independent two-layer linear transformation with an activation function: 1. First layer linear transformation (input → hidden layer): implemented by (ParallelExperts instance), the output dimension is The emotion classification head is 4 and the speaker classification head is 10, and the formula is:
[0088] wherein, is the number of experts, The output dimension of the expert is M, and M is the total number of selected samples. ParallelExperts is calculated in parallel by a custom ParallelLinear function, and the formula is:
[0089] wherein, is the first layer weight of the e-th expert, is the bias.
[0090] 2*Activation function: apply the activation function to the hidden layer features, and use the GELU activation here:
[0091] wherein, is the output of the activation function.
[0092] 3*Second layer linear transformation (hidden layer->output): realized by output_experts (another ParallelExperts instance), the output dimension is consistent with the input dimension head_size:
[0093] The formula is:
[0094] wherein, is the second layer weight of the e-th expert.
[0095] (3) Weight scaling and output aggregation Weight scaling: weight the expert output with the gating weight batch_gates:
[0096] Aggregated output: map the weighted expert output back to the original sample dimension through index_add, and finally output , which is consistent with the input dimension.
[0097] wherein, is the weighted expert output, is the weighted calculation, is the gating weight.
[0098] In summary, the application has certain improvement in weighted accuracy (WA), unweighted accuracy (UA), and total running time of the same training round, among which the performance improvement is particularly significant.
[0099] In terms of performance improvement, there are the following reasons: (1) Sparse computation and dynamic routing Selective activation: Each sample in MoE is only activated by top-k (e.g. k = 2) expert networks, instead of full parameter calculation of traditional fully connected layers. For example: the sentiment classification head uses 4-to-2, and the calculation amount is reduced to 50%. Lightweight of gating network: The gating network is only a lightweight linear layer (the parameter amount is much smaller than that of the expert network).
[0100] (2) Parallelization advantage Expert parallel calculation: Parallelization of cross-expert matrix multiplication is realized through ParallelLinear, and GPU utilization is improved. Memory access optimization: index_add and other operations reduce memory fragmentation and improve memory bandwidth utilization. In terms of accuracy improvement, in order to make each expert specialized in processing tasks, a sentiment classification expert and a speaker recognition expert are constructed respectively. The sentiment classification network has 4 experts, and different experts can focus on different patterns such as prosody and spectral features. The speaker recognition network includes 6 experts, which learn different features such as timbre and pronunciation habits.
[0101] Load balancing optimization is also realized between experts. zloss is used to prevent the variance of gating logits from being too large, keep the training stable, and switchloss is used to promote the division of labor among experts. Experiments show that it can improve the coverage rate of experts.
[0102] Reference Figure 6 The application also provides a multi-task speech emotion recognition system based on parallel processing hybrid expert network, comprising: A feature extraction module 610 is configured to extract features from input raw speech to obtain shared features. A parallel processing module 620 is configured to construct a parallel processing hybrid expert network, input the shared features into the parallel processing hybrid expert network, perform parallel sentiment classification and speaker recognition on the shared features, and output predicted emotion values and predicted speakers respectively. A speech recognition module 630 is configured to input the shared features into a fully connected layer for speech recognition processing to obtain text prediction values. An optimization module 540 is configured to obtain a loss function according to the predicted emotion values, predicted speakers, and text prediction values, respectively, in combination with corresponding labels, and optimize model parameters according to the loss function to obtain an optimized parallel processing hybrid expert network.
[0103] Optionally, the specific process of sentiment classification comprises: Input the shared features into the sentiment classification head of the parallel processing hybrid expert network. Select a first expert network from the emotion expert pool according to the shared features to activate. inputting the shared feature into the selected first expert network to extract a first feature representation related to sentiment; adopting a noise injection and Top-k selection mechanism to generate first gating weights of the selected expert network; performing weighted summation on the first feature representation and corresponding first gating weights to obtain the predicted sentiment value of the final output.
[0104] Optionally, the specific process of speaker recognition comprises: inputting the shared feature into a speaker recognition classification head of the parallel processing hybrid expert network; selecting a second expert network from a speaker recognition expert pool according to the shared feature to activate; inputting the shared feature into the selected second expert network to extract a second feature representation related to a speaker; adopting a noise injection and Top-k selection mechanism to generate second gating weights of the selected expert network; performing weighted summation on the second feature representation and corresponding second gating weights to obtain hierarchical decoupled voiceprint features, and completing speaker recognition based on the voiceprint features.
[0105] Optionally, the first expert network and / or the first expert network is obtained through a gating mechanism of the parallel processing hybrid expert network.
[0106] Optionally, the implementation process of the gating mechanism comprises: performing linear projection on the input shared feature to obtain an initial logical value; injecting Gaussian noise and noise dimension segmentation into the initial logical value, and constraining through a softplus function to obtain a final noise logical value; performing a softmax operation on the noise logical value to obtain an expert probability distribution; selecting Top-k expert networks and corresponding gating weights in the expert probability distribution to generate a sparse gating matrix.
[0107] Optionally, the parallel processing hybrid expert network realizes expert load balancing through a coefficient of variation loss, an expert utilization rate loss, and a numerical stability loss: The implementation process of the expert load balancing comprises: determining a gating value sum, a gating value variance, and a gating value mean based on the first gating weights and the second gating weights, and combining a set minimum constant to construct the coefficient of variation loss; determining an expert network selection probability according to the probability distribution of each expert network, determining a network usage frequency according to the number of times each expert network is selected, and combining the total number of expert networks to construct the expert utilization rate loss; According to the noise logic value corresponding to each expert network, the numerical stability loss is constructed; The coefficient of variation loss, the expert utilization rate loss, and the numerical stability loss are combined with a pre-defined weight coefficient to obtain a final auxiliary loss for model training.
[0108] Optionally, the first expert network and the second expert network each include a first linear layer, a GELU activation function layer, and a second linear layer. The first linear layer is configured to perform independent linear transformation on the input shared feature and output a hidden layer feature. The GELU activation function layer is configured to apply a Gaussian error linear unit activation to the hidden layer feature. The second linear layer is configured to map the activated feature back to an output space matching the dimension of the input shared feature.
[0109] It should be understood that the above device is used to execute the method in the above embodiments, and the corresponding program modules in the device have similar implementation principles and technical effects to those described in the above method. The working process of the device can refer to the corresponding process in the above method, which will not be described here.
[0110] Reference Figure 7 Based on the method in the above embodiments, the embodiments of the present application provide an electronic device, which can include a processor (Processor) 710, a communication interface (Communications Interface) 720, a memory (Memory) 730, and a communication bus 740. The processor 710, the communication interface 720, and the memory 730 can communicate with each other through the communication bus 740. The processor 710 can invoke the logic instructions in the memory 730 to execute the method in the above embodiments.
[0111] In addition, the logic instructions in the memory 730 described above can be implemented in the form of a software functional unit and sold or used as an independent product. When used, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application.
[0112] Based on the method in the above embodiments, the embodiments of the present application provide a computer readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiments.
[0113] Based on the method in the above embodiments, the embodiments of the present application provide a computer program product, which, when running on a processor, causes the processor to perform the method in the above embodiments.
[0114] It can be understood that the processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.
[0115] The method steps in the embodiments of the present application can be implemented in the form of hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC.
[0116] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in or transmitted by a computer readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0117] It can be understood that various numerical numbers involved in the embodiments of the present application are only distinguished for convenience of description, and are not used to limit the scope of the embodiments of the present application.
[0118] Those skilled in the art easily understand that the above only describes the preferred embodiments of the present application and is not used to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A multi-task speech emotion recognition method based on parallel processing hybrid expert network, characterized in that: Extract features from the input original speech to obtain shared features; Constructing a parallel processing hybrid expert network, inputting the shared features into the parallel processing hybrid expert network, performing sentiment classification and speaker recognition on the shared features in parallel, and outputting a predicted sentiment value and a predicted speaker respectively; Inputting the shared features into a fully connected layer for speech recognition processing to obtain a text prediction value; A loss function is obtained according to the predicted emotion value, the predicted speaker and the text prediction value in combination with corresponding labels, and the model parameters are optimized according to the loss function to obtain a parallel processing hybrid expert network after optimization training.
2. The multi-task speech emotion recognition method based on parallel processing hybrid expert network according to claim 1 is characterized in that: The specific process of sentiment classification includes: Inputting the shared features into the sentiment classification head of the parallel processing hybrid expert network; Selecting a first expert network from the emotion expert pool according to the shared features for activation; Inputting the shared features into the selected first expert network to extract a first feature representation related to emotion; Noise injection and Top-k selection mechanism are used to generate the first gating weight of the selected expert network; The first feature representation and the corresponding first gating weight are weightedly summed to obtain the predicted emotion value that is finally output.
3. The multi-task speech emotion recognition method based on parallel processing hybrid expert network according to claim 1 is characterized in that: The specific process of speaker recognition includes: Inputting the shared features into a speaker recognition classification head of a parallel processing hybrid expert network; selecting a second expert network from a speaker recognition expert pool according to the shared features for activation; Inputting the shared features into the selected second expert network to extract a second feature representation related to the speaker; Noise injection and Top-k selection mechanism are used to generate the second gating weight of the selected expert network; A weighted sum is performed on the second feature representation and the corresponding second gating weight to obtain a hierarchically decoupled voiceprint feature, and speaker recognition is performed based on the voiceprint feature.
4. The multi-task speech emotion recognition method based on parallel processing hybrid expert network according to claim 3 is characterized in that: The first expert network and / or the first expert network are obtained by screening through a gating mechanism of parallel processing hybrid expert networks.
5. The multi-task speech emotion recognition method based on parallel processing hybrid expert network according to claim 4 is characterized in that: The implementation process of the gating mechanism includes: Perform linear projection on the shared features of the input to obtain the initial logical value; Injecting Gaussian noise and noise dimension splitting into the initial logic value, and constraining it through a softplus function to obtain a final noise logic value; Performing a softmax operation on the noise logic value to obtain an expert probability distribution; The top-k expert networks and their corresponding gating weights in the expert probability distribution are selected to generate a sparse gating matrix.
6. The multi-task speech emotion recognition method based on parallel processing hybrid expert network according to claim 4 is characterized in that: The parallel processing hybrid expert network achieves expert load balancing through coefficient of variation loss, expert utilization loss, and numerical stability loss: The implementation process of the expert load balancing includes: Determining a gating value sum based on the first gating weight and the second gating weight, determining a gating value variance and a gating value mean, and constructing the coefficient of variation loss in combination with a set minimum constant; Determine the probability of selecting an expert network based on the probability distribution of each expert network, determine the frequency of network usage based on the number of times the expert network is selected, and construct the expert utilization loss based on the total number of expert networks; Constructing the numerical stability loss according to the noise logic value corresponding to each expert network; The coefficient of variation loss, expert utilization loss, and numerical stability loss are combined with predefined weight coefficients to obtain the final auxiliary loss for model training.
7. The multi-task speech emotion recognition method based on parallel processing hybrid expert network according to claim 4 is characterized in that: The first expert network and the second expert network both include a first linear layer, a GELU activation function layer and a second linear layer; The first linear layer is used to perform independent linear transformation on the input shared features and output hidden layer features; GELU activation function layer, used to apply Gaussian error linear unit activation to hidden layer features; The second linear layer is used to map the activated features back to the output space that matches the shared feature dimensions of the input.
8. A multi-task speech emotion recognition system based on parallel processing hybrid expert network, characterized in that: include: The feature extraction module is used to extract features from the input original speech to obtain shared features; A parallel processing module constructs a parallel processing hybrid expert network, inputs the shared features into the parallel processing hybrid expert network, performs sentiment classification and speaker recognition on the shared features in parallel, and outputs a predicted sentiment value and a predicted speaker respectively; A speech recognition module is used to input the shared features into a fully connected layer for speech recognition processing to obtain a text prediction value; The optimization module is used to obtain a loss function based on the predicted emotion value, the predicted speaker and the text prediction value in combination with the corresponding labels, and optimize the model parameters according to the loss function to obtain a parallel processing hybrid expert network after optimization training.
9. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed on a processor, the processor is caused to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Voice extraction method and system in robot interaction process, medium and robot
CN122314008A