Speech emotion recognition method and system based on multi-modal feature fusion
By employing a multimodal feature fusion method, the problems of insufficient efficiency and robustness in high-dimensional feature processing in speech emotion recognition are solved, achieving compact and discriminative feature selection, and improving the accuracy and real-time performance of speech emotion recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN XIAOYU ZHIHE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-04-21
AI Technical Summary
Existing speech emotion recognition technologies struggle to balance feature representation capabilities and processing efficiency under limited computing power and strict latency constraints when processing high-dimensional features. This leads to the loss of key time-frequency details and emotion-sensitive features, resulting in decreased overall classification accuracy and insufficient robustness.
Feature selection and fusion optimization are achieved through multimodal feature fusion methods, including quality screening of speech emotion datasets, unified encoding of scenes and labels, construction of frame-level acoustic and spectral features, generation of submodal emotion feature vector sets, multimodal deep encoding and gated collaborative attention mechanism, random forest weight initialization and two-order mutated gray wolf optimization.
While suppressing unnecessary dimensionality reduction, it provides compact and discriminative input features, reduces redundant information interference, improves the stability and robustness of feature selection, and enhances the accuracy and real-time performance of speech emotion recognition.
Smart Images

Figure CN121565211B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech emotion recognition technology, specifically to a speech emotion recognition method and system based on multimodal feature fusion. Background Technology
[0002] With the widespread application of voice emotion recognition technology in scenarios such as real-time voice monitoring in emergency centers, driver status and emotion monitoring, intelligent translation and voice assistant interaction, and online customer service, higher requirements have been placed on the accuracy, response latency, and robustness of emotion recognition. While ensuring the integrity of information, high-dimensional voice emotion features also significantly increase model complexity and real-time computing burden. Existing technologies struggle to balance feature representation ability and processing efficiency under limited computing power and strict latency constraints, which can easily lead to the discarding of some key time-frequency details and emotionally sensitive features, distorting the distribution structure of the original emotion feature space.
[0003] For example, the invention patent with announcement number CN115312033B discloses a speech emotion recognition method, device, equipment, and medium based on artificial intelligence, including: performing frame segmentation and windowing processing on the speech information to be recognized to obtain a speech frame sequence, extracting the speech feature tensor and text feature tensor of the speech information to be recognized, aligning the speech feature tensor and text feature tensor and performing feature extraction to obtain multimodal features, using local windows to perform average pooling and global max pooling processing on the speech frame sequence to obtain enhanced speech features, fusing the enhanced speech features with the multimodal features, determining the fusion result, and obtaining the emotion recognition result based on the fusion result. By obtaining low-level enhanced speech features through pooling processing and fusing the enhanced speech features with the multimodal features, the degradation problem of deep networks is avoided, and the accuracy of speech emotion recognition is effectively improved while improving the generalization ability of the model.
[0004] For example, invention patent CN111583964B discloses a natural speech emotion recognition method based on multimodal deep feature learning, including the following steps: S1, generating appropriate multimodal representations: generating three appropriate audio representations from the original one-dimensional speech signal for use as input to different CNN models; S2, learning multimodal features using a multi-deep convolutional neural network model; S3, integrating the classification results of different CNN models using a score-level fusion method to output the final speech emotion recognition result. This invention significantly improves emotion classification performance by fusing and learning complementary deep multimodal features through multi-deep convolutional neural networks, providing features with good discriminative power for natural speech emotion recognition.
[0005] In existing technologies, existing systems typically perform noise cancellation, feature extraction, and normalization preprocessing operations sequentially on the acquired speech signals, and then directly input the high-dimensional features into the emotion classification model. When facing scenarios such as emergency command recognition and real-time driver emotion monitoring, in order to meet the requirements of online inference speed, feature dimensions are often coarsely reduced by downsampling and feature compression, which leads to bias in emotion recognition results, decreased overall classification accuracy, and insufficient robustness to noise and scene changes.
[0006] Therefore, in order to address the above problems, there is an urgent need for speech emotion recognition methods and systems based on multimodal feature fusion. Summary of the Invention
[0007] Technical problems to be solved
[0008] To address the shortcomings of existing technologies, this invention provides a speech emotion recognition method and system based on multimodal feature fusion, which solves the problems of existing speech emotion feature selection being single-stage and single-index, making it difficult to balance emotion preservation and real-time performance while reducing dimensionality, and limiting overall performance.
[0009] Technical solution
[0010] To achieve the above objectives, this invention provides the following technical solution: a speech emotion recognition method and system based on multimodal feature fusion, comprising: S1, collecting a speech emotion dataset, performing quality screening, scene and label unified encoding and normalization on the speech emotion dataset to form a standardized speech emotion sample library; S2, performing frame-level acoustic analysis and constructing spectral features on the standardized speech emotion sample library, and constructing prosodic and phonological emotion features and labeling confidence to obtain a submodal emotion feature vector set; S3, a speech emotion representation generation method based on the submodal emotion feature vector set, which integrates multimodal deep coding and gated collaborative attention mechanism; S4, performing pre-screening and random forest weight initialization on the speech emotion representation vector set, and performing two-order mutation gray wolf and mapping evaluation to obtain a speech emotion feature subset; S5, performing speech emotion recognition and confidence decision on the speech emotion feature subset, and performing emotion recognition service orchestration and adaptive update based on operational feedback.
[0011] Furthermore, the second aspect of the present invention provides a speech emotion recognition system based on multimodal feature fusion, applied to a speech emotion recognition method based on multimodal feature fusion, comprising: a speech data acquisition and emotion annotation module, used to acquire a speech emotion dataset, perform quality screening, scene and label unified encoding and normalization processing on the speech emotion dataset to form a standardized speech emotion sample library; an emotion feature construction module, used to perform frame-level acoustic and spectral feature construction on the standardized speech emotion sample library, and to construct prosodic and phonological emotion features and perform confidence annotation to obtain a submodal emotion feature vector set; a feature fusion and learning module, used to generate a speech emotion representation based on the submodal emotion feature vector set by fusing multimodal deep coding and gated collaborative attention mechanism; an emotion feature selection and adaptive optimization module, used to pre-screen and initialize random forest weights on the speech emotion representation vector set, and perform two-order mutation gray wolf and mapping evaluation to obtain a subset of speech emotion features; and an emotion recognition decision and output module, used to perform speech emotion recognition and confidence decision on the subset of speech emotion features, and to perform emotion recognition service orchestration and operation feedback adaptive update.
[0012] The present invention has the following beneficial effects:
[0013] (1) This invention helps to suppress unnecessary dimensionality reduction operations from the source by performing emotional feature recognition and real-time performance evaluation on speech signals before feature selection and fusion optimization; and provides more compact and discriminative input features for subsequent emotion classification and feature extraction without destroying effective emotional information, thereby reducing the interference of redundant information caused by high-dimensional features.
[0014] (2) In this invention, a multi-factor comprehensive score is constructed by combining the feature weights obtained by the ReliefF algorithm with the feature importance weights output by the random forest, which avoids the problem of a certain indicator masking the true contribution of multi-dimensional features; upgrading the single-dimensional judgment to a multi-dimensional comprehensive evaluation helps to reduce the mis-screening and omission of key emotional features in the feature selection process, and improves the discriminability and reliability of the final feature subset.
[0015] (3) In this invention, by introducing submodal emotional influence factors and noise interference factors in the ReliefF pre-screening and random forest weight initialization process, the importance score is weighted and differentiated, which helps to filter out redundant features with large noise interference in speech features, concentrate on core emotional features, reduce the search space, reduce the convergence difficulty of feature selection, and improve the stability and robustness of the feature optimization process.
[0016] (4) This invention constructs a nonlinear quality assessment mechanism for feature subset quality scores, and dynamically scores each candidate feature subset in the global search and local refinement stages of the two-order mutated gray wolf optimization. This helps to accurately identify performance bottlenecks caused by feature selection, noise mismatch or abnormal optimization state, and achieves a unified improvement in feature dimensionality reduction efficiency and speech emotion recognition accuracy and real-time performance.
[0017] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0018] Figure 1 This is a flowchart of the speech emotion recognition method based on multimodal feature fusion according to the present invention;
[0019] Figure 2 This is a framework diagram of the speech emotion recognition system based on multimodal feature fusion according to the present invention;
[0020] Figure 3 This is a scatter matrix diagram showing the relationship between the submodal sentiment index and the benchmark factor of this invention.
[0021] Figure 4 This is a diagram illustrating the overall framework of the feature selection algorithm of this invention.
[0022] Figure 5 This is a flowchart of the adaptive update method for speech emotion recognition service orchestration and operation feedback in this invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Please see Figures 1-5This invention provides a technical solution: a speech emotion recognition method and system based on multimodal feature fusion, comprising: S1, collecting a speech emotion dataset, performing quality screening, scene and label unified encoding and normalization on the speech emotion dataset to form a standardized speech emotion sample library; S2, performing frame-level acoustic analysis and constructing spectral features on the standardized speech emotion sample library, and constructing prosodic and phonological emotion features and assigning confidence to obtain a submodal emotion feature vector set; S3, a speech emotion representation generation method based on the submodal emotion feature vector set, which integrates multimodal deep coding and gated collaborative attention mechanism; S4, performing pre-screening and random forest weight initialization on the speech emotion representation vector set, and performing two-order mutated gray wolf and mapping evaluation to obtain a speech emotion feature subset; S5, performing speech emotion recognition and confidence decision on the speech emotion feature subset, and performing emotion recognition service orchestration and adaptive update based on operational feedback.
[0025] Specifically, the process of collecting speech emotion datasets, performing quality screening, scene and label unified encoding, and normalization on these datasets to form a standardized speech emotion sample library is as follows: Speech emotion datasets are collected collaboratively from multiple sources. These datasets include: raw speech waveform datasets, emotion label and scene marker datasets, speaker attribute and conversation structure datasets, noise environment and channel state datasets, and sample partitioning and version management datasets. Multiple source scenarios include real-time speech monitoring in emergency centers, driver status and emotion monitoring, intelligent translation and voice assistant interaction, online customer service and hotline services, and publicly available speech emotion corpora. The raw speech waveform dataset includes speech waveform sequences and their sampling rate, quantization accuracy, number of channels, and acquisition time markers; these are raw speech signals acquired through call center recording equipment, vehicle microphones, or terminal acquisition devices. Emotion label and scene marker data... The dataset includes emotion categories, emotion intensity levels, urgency levels, and business scenario type labels. It is obtained by semi-automatic emotion annotation of speech samples, uniformly encoding emotion states, intensity levels, emergency commands, driving monitoring, and daily dialogue business scenarios. The speaker attribute and conversation structure dataset includes speaker identifiers, gender, age group, dialect accent, speaker role, as well as conversation identifiers, turn numbers, and the start and end times of speech segments. It is speaker and conversation structure information collected or derived from upper-layer business systems. The noise environment and channel state dataset includes environmental noise type, signal-to-noise ratio estimation, recording location type, channel coding method, packet loss rate, and distortion labels. It is obtained by analyzing the energy and spectral characteristics of speech segments and combining them with the status logs of the acquisition equipment. The sample partitioning and version management dataset includes sample numbers, partitioning labels, data version numbers, and activation labels. It is generated through stratified sampling and a unified numbering system.
[0026] The stored speech emotion data sequences already completed in the same speech emotion recognition business scenario are used as historical operational data. Historical operational data includes: historical speech waveform sequences from different time periods, historical emotion tag and scene label sequences, historical speaker attribute and conversation structure records, historical noise environment, channel state records, and corresponding sample division and version information. The collected speech emotion datasets are classified and managed according to data type labeling and quality attribute classification. They are categorized and labeled as raw speech, emotion tag, speaker and conversation structure, noise and channel state, and sample management. In the access cache queue, the records of the raw speech and emotion tag categories undergo integrity and consistency checks. Each speech record is tested for duration, silence ratio, and distortion, and a basic signal-to-noise ratio is evaluated. Quality anomalies are marked and entered into the review queue. Samples lacking emotion tags or scene labels are marked with missing labels and entered into the supplementary labeling queue. Uploaded voice segments are formatted and completed using a unified voice acquisition standard. For devices supporting automatic metadata acquisition, sampling rate, quantization accuracy, and time reference are periodically synchronized. For devices not supporting automatic synchronization, device parameters and time deviation information are recorded through configuration files. Preliminary sample alignment is performed based on the original voice waveform and acquisition time. The original voice waveform and sampling parameters are uniformly encoded, and the recording device identifier, session identifier, and time field are written into the voice metadata table to obtain the processed voice emotion dataset. At the acquisition nodes, a global voice reception queue sorted by acquisition time and a local voice reception queue divided by business scenario type are maintained. Samples with abnormal quality and missing labels are marked and managed. The processed voice emotion dataset is encapsulated with a unified data structure to construct voice emotion acquisition sample frames. These frames include: sample number, session identifier, voice start and end time, sampling rate, quantization accuracy, voice duration, business scenario type label, emotion category label, emotion intensity level, urgency level label, speaker attribute field, noise and channel state field, sample quality label, and dataset partitioning label. Feature extraction status field and feature selection label field are reserved. The fields in the collected sample frames are subjected to z-score dimensionless normalization and encoding preprocessing to obtain a standardized speech emotion sample library: z-score dimensionless normalization is only applied to the original numerical fields in the collected sample frames; the emotion intensity and urgency scale labels are mapped to a unified numerical range, and the discrete labels of speaker attributes and business scenario types are converted into a standardized input form using one-hot encoding.
[0027] In this implementation plan, a standardized voice emotion sample library is obtained by screening the voice emotion dataset for quality, uniformly encoding the scenes and labels, and normalizing the data. This reduces the interference of differences in multi-source data and low-quality samples on modeling, and significantly improves the stability, accuracy, and traceability of subsequent emotion feature construction and classification model training.
[0028] Specifically, the process of performing frame-level acoustic and spectral feature construction on a standardized speech emotion sample library is as follows: Input speech emotion collection sample frames, perform frame-level acoustic and spectral feature construction to generate speech emotion spectral submodal features; Read speech emotion collection sample frames according to sample number, session identifier, and dataset division tags, extract speech segments according to the speech start and end time fields, perform frame segmentation and windowing according to uniform frame length and frame shift parameters to obtain a speech frame sequence; Perform short-time Fourier transform on each speech frame, calculate the power spectrum and amplitude spectrum, construct a Mel filter bank to extract Mel frequency energy coefficients, and calculate the spectral correlation features of MFCC and first- and second-order differences, spectral centroid, spectral bandwidth, spectral flatness, and spectral roll-off point, and concatenate them in sequence to form a frame-level acoustic and spectral feature vector; Index and manage all frame-level acoustic and spectral feature vectors of the same speech segment and store them in a structured manner, establish a correspondence between frame number and sample number, session identifier, and speech start and end time, and output a frame-level acoustic and spectral feature dataset.
[0029] In this implementation scheme, by performing frame-by-frame windowing with unified parameters and joint extraction of multiple spectral features on a standardized speech emotion sample library, the original speech signal is transformed into a time-consistent frame-level acoustic and spectral feature dataset with rich time-frequency details, providing a complete and structurally sound basic representation for submodal emotion modeling.
[0030] Specifically, the process of constructing prosodic and phonological emotional features and labeling confidence levels to obtain submodal emotional feature vector sets is as follows: Input frame-level acoustic and spectral feature datasets, construct prosodic and phonological emotional features, encapsulate them to form submodal emotional feature vectors, align frame-level acoustic and spectral feature sequences using speech start and end time fields and session identifiers, perform fundamental frequency estimation and rhythm analysis on the aligned speech segments, calculate fundamental frequency trajectory, short-time energy, zero-crossing rate parameters frame by frame, statistically analyze fundamental frequency representative values, fluctuation amplitude, extreme value distribution, fundamental frequency change rate, and fundamental frequency trajectory stability index at the segment level, and construct prosodic features reflecting pitch fluctuations and emotional excitement; combine speaker attributes and round numbers recorded in the session structure field with speech segment division information to construct rhythm features related to tension and urgency; calculate phonological features of jitter, tremolo, and harmonic noise ratio, and attach noise-sensitive labels to phonological features by referring to the environmental noise type in the noise environment field, the recording location type, and the packet loss rate and distortion markers in the channel state field. The sentiment sensitivity index is obtained by normalizing the gradient contribution and inter-class separability statistics between submodal features and sentiment labels in historical operational data; the quality index is obtained by fusing and mapping the sample quality label, the signal-to-noise ratio estimate in the noise environment field, and the packet loss rate and distortion label in the channel state field; the confidence score is attached to the submodal feature vector and encapsulated together with the frame-level acoustic and spectral feature dataset mapping relationship into a submodal sentiment feature vector set. The quality difference term is obtained by squaring the difference between the quality reference threshold and the submodal quality index. The quality difference term is obtained by multiplying the square of the m-th submodal sentiment sensitivity index (minus the quality penalty weight) by the quality difference term. The quality adjustment term is then multiplied by the confidence sharpening coefficient. The reciprocal of the dynamic benchmark factor is taken as the sentiment-related term. The quality adjustment term is then negatively evaluated and exponentially calculated to obtain the index term. The sentiment-related term is multiplied by the index term to obtain the penalty term. A fixed value is added to the penalty term to obtain the scoring denominator. Finally, the fixed value is divided by the scoring denominator to obtain the confidence score of the m-th submodal. The specific calculation formula for the confidence score of the m-th submodal is as follows:
[0031] ;
[0032] ;
[0033] In the formula, This represents the confidence score of the m-th submodality, used to quantify the degree of matching between the sensitivity and reliability of the submodality on the current sample; The quality adjustment term is obtained by squaring the difference between the quality reference threshold and the submodal quality index. The quality penalty weight is multiplied by the quality difference term to obtain the penalty term. The square of the emotional sensitivity index of the m-th submodal is subtracted from the penalty term to obtain the comprehensive quality adjustment term. The comprehensive quality adjustment term is multiplied by the confidence sharpening coefficient. It is used to quantify the impact of the deviation of the quality level of the m-th submodal from the reference threshold on the overall performance of the system. The sentiment sensitivity index of the m-th submodality is obtained by normalizing the gradient contribution and inter-class separability statistics between submodal features and sentiment labels in historical data. The larger the value, the more sensitive the submodality is to distinguish sentiment categories. The quality index of the m-th sub-mode is obtained by fusing and mapping the sample quality label, the signal-to-noise ratio estimate in the noise environment field, and the packet loss rate and distortion label in the channel state field. The larger the value, the higher the reliability of the sub-mode on the current sample. The quality reference threshold is obtained by taking into account the quality level range of emotion recognition accuracy and robustness in historical operating data, and by using the method that minimizes the false positive rate from the range. It is used as the ideal quality center for the penalty term. The confidence sharpening coefficient is determined by sweeping different values across historical data, comparing the consistency between confidence-driven feature selection results and emotion recognition accuracy and robustness indicators, and using cross-validation to select the value that best matches the target performance. The value range is a real number greater than zero, and it is used to adjust the steepness of the change in the confidence score response emotion sensitivity and quality deviation. The quality penalty weight is selected by analyzing the correspondence between quality indicators and emotion recognition performance under different noise and distortion conditions, and by minimizing the confidence false positive rate. The value range is a real number greater than zero, which is used to control the non-linear penalty intensity of confidence score when the quality deviates from the reference threshold. This represents a dynamic benchmark factor. By statistically analyzing the correspondence between confidence distribution and actual performance in different scenarios, it adaptively selects representative values that ensure the confidence score distribution falls within the target working range. These values are used to adjust the numerical scale and upper bound of the confidence score.
[0034] In this implementation scheme, by constructing multi-modal emotional features based on fundamental frequency, rhythm, and sound quality parameters and introducing a confidence scoring mechanism that considers emotional sensitivity and sample quality, the sub-modal features participating in fusion and feature selection are jointly constrained in terms of information content, reliability, and noise robustness, thereby providing more reliable sub-modal inputs for multi-modal emotional representation and optimization.
[0035] Specifically, the process of the speech emotion representation generation method based on the submodal emotion feature vector set, which integrates multi-submodal deep coding and gated collaborative attention mechanism, is as follows: Inputting the submodal emotion feature vector set and the corresponding submodal confidence scores, the method performs deep coding and gated collaborative attention fusion on the spectral, prosodic, and phonological submodal features carried in each speech emotion acquisition sample frame to generate a unified speech emotion representation vector, and outputs the submodal emotion influence factor and noise interference factor; for the spectral submodal features, a spectral coding subnetwork is constructed, taking the frame-level acoustic and spectral feature dataset as input, and extracting local time-frequency patterns and long-range dependencies through two-dimensional convolution and temporal modeling units, and statistically aggregating them in the time dimension to obtain the spectral submodal embedding vector; for the prosodic submodal features, a prosodic coding subnetwork is constructed, taking the submodal emotion feature vector set as input, and extracting local time-frequency patterns and long-range dependencies through one-dimensional convolution and gated recurrent... The unit extracts the rhythmic beat structure and overall rhythmic evolution pattern to obtain the prosodic submodal embedding vector. For the sound quality submodal features, a sound quality coding subnetwork is constructed, using jitter, tremolo, harmonic-to-noise ratio, statistics, and rate of change as inputs. High-dimensional sound quality embedding vectors related to vocal cord stability and tension are extracted through multi-layer nonlinear mapping. Gating suppression is applied to feature channels affected by strong noise based on noise-sensitive labels, ensuring the subnetwork remains responsive to noise-dominant patterns. After normalizing the quality index and emotional sensitivity index to a unified interval, a monotonic nonlinear mapping aggregation is used to obtain the submodal confidence score. Based on the submodal confidence score, the channel activation and dropout rates within each subnetwork are dynamically adjusted. When the submodal confidence is within the first-level threshold range, intermediate feature channels and representation dimensions are retained; when the submodal confidence is within the second-level threshold range, the dropout intensity of the submodal branch is increased, and the output amplitude is limited.
[0036] The noise inconsistency index is obtained by statistically analyzing the distribution differences of sub-modal embeddings under different noise labels, channel states, and sample quality levels in multiple noise scenarios and multiple channel conditions, and compressing and mapping the cross-scenario output distribution deviation. A comprehensive index characterization value for sub-modal representation is constructed based on the emotion sensitivity index and the noise inconsistency index. The sensitivity amplification factor is multiplied by the emotion sensitivity index of the k-th sub-modal to obtain the amplification term. The amplification term is then exponentially calculated and added to the dynamic benchmark factor of emotion sensitivity to obtain the emotion sensitivity term. A logarithmic operation is performed on the emotion sensitivity term to obtain the sensitivity term. The difference between the noise inconsistency index of the k-th sub-modal and the noise inconsistency reference threshold is obtained to obtain the deviation term. The deviation term is multiplied by the noise penalty amplification factor and exponentially calculated to obtain the penalty term. The penalty term is added to the dynamic benchmark factor of noise inconsistency to obtain the penalty consistency term. A logarithmic operation is performed on the penalty consistency term to obtain the noise term. The comprehensive index characterization value of the k-th sub-modal is obtained by subtracting the noise term from the sensitivity term. The specific calculation formula for the comprehensive index characterization value is as follows:
[0037] ;
[0038] ;
[0039] ;
[0040] In the formula, The comprehensive index representation value of the k-th submodality is used to simultaneously reflect the emotional sensitivity enhancement effect and noise inconsistency penalty effect of the submodality within a unified index space. This represents the amplification term, which is obtained by multiplying the sensitivity amplification factor by the emotional sensitivity index of the k-th submodality. It is used to amplify the numerical influence of the index, making the emotional sensitivity characteristics of the submodality more prominent in performance analysis and state assessment. The penalty term is obtained by subtracting the noise inconsistency index of the k-th submode from the noise inconsistency reference threshold. The deviation term is then multiplied by the noise penalty amplification factor and exponentially calculated. This term is used to quantify the penalty for noise inconsistency in the k-th submode exceeding the reference threshold. The greater the deviation of the noise inconsistency index from the reference threshold, the more significant the penalty will be after being amplified by the noise penalty amplification factor and exponential calculation. This helps to suppress the interference of submodes with excessive noise inconsistency on the overall evaluation results and ensure the reliability of the evaluation. The dynamic benchmark factor for emotion sensitivity is represented by the distribution of emotion recognition accuracy and recall corresponding to different sensitivity levels in historical operating data. It adaptively selects a representative benchmark value that makes the high-sensitivity submodal and the general submodal distinguishable in the exponential space. The sensitivity amplification factor is determined by sweeping through historical data to examine the impact of different values on emotion recognition performance and robustness indicators. Cross-validation is then used to select the value that best matches the target performance. The value ranges from 0.5 to 2.0 and is used to adjust the amplification of emotion sensitivity changes in the exponential space. This represents a dynamic benchmark factor for noise inconsistency. By comparing noise inconsistency indicators and performance degradation under different noise scenarios, an index benchmark that can distinguish between acceptable noise fluctuations and unacceptable noise mismatches is selected to achieve an adaptive characterization of noise levels. The noise penalty amplification factor ranges from 1.0 to 3.0. It is obtained by fitting the relationship between the degree of noise inconsistency deviation and the decline in emotion recognition performance on historical operating data and minimizing the prediction error. It is used to adjust the penalty intensity of the index representation when the noise inconsistency deviates from the reference threshold. The noise inconsistency reference threshold is represented by the noise inconsistency index and the model performance stability index under multiple scenarios and multiple device conditions. The noise level that still maintains stable performance is selected from the index distribution as a representative reference value to distinguish between acceptable noise deviation and noise deviation that needs to be severely penalized. The emotional sensitivity index of the k-th submodality is used to characterize the strength of the submodality's ability to distinguish the current speech emotion category; The noise inconsistency index for the k-th submodality is used to characterize the stability and noise interference level of the output mode of this submodality under different noise environments and channel conditions. By statistically analyzing the distribution differences of the k-th submodality embedding under different noise labels, channel states, and sample quality levels in multiple noise scenarios and multiple channel states, the cross-scenario output distribution deviation and performance degradation trend are compressed and mapped into dimensionless values. A gated collaborative attention fusion mechanism is constructed, converting the comprehensive index representation value into submodality attention weights and influence factors through a monotonically increasing sigmoid function. The comprehensive index representation value is also converted into submodality fusion weight coefficients through a monotonically increasing sigmoid function. Simultaneously, using the comprehensive index representation value as the independent variable and the noise interference factor as a sample pair, a monotonically increasing mapping function is trained using a nonlinear regression method. The submodality emotion influence factor is obtained through the mapping function. Using the obtained submodality fusion weight coefficients, the spectral submodality embedding, prosodic submodality embedding, and phonological submodality embedding are gated and fused to form a unified speech emotion representation vector. Meanwhile, the sentiment influence factor and noise interference factor of each submodal are retained in the bypass branch and associated with the sample number, session identifier, and original feature dimension index and written into the feature metadata table.
[0041] As shown in Table 1, the comprehensive index representation data of the speech emotion representation generation submodal is as follows: In a quiet environment, under the spectral submodal, the emotion sensitivity index quality score is 0.9; the noise inconsistency index quality score is 0.22; the inconsistency reference threshold quality score is 0.4; the comprehensive index representation value quality score is 0.287; the emotion sensitivity dynamic benchmark factor is 0.95; and the noise inconsistency dynamic benchmark factor is 0.15. Under the prosody submodal, the emotion sensitivity index quality score is 0.82; the noise inconsistency index quality score is 0.28; the comprehensive index representation value quality score is 0.192; and the emotion sensitivity dynamic benchmark factor is 0.88; the noise inconsistency dynamic benchmark factor is 0.22. Under the tone quality submodal, the emotion sensitivity index quality score is 0.78; the noise inconsistency index quality score is 0.25; the comprehensive index representation value quality score is 0.164; the emotion sensitivity dynamic benchmark factor is 0.85; and the noise inconsistency dynamic benchmark factor is 0.20. The conversation quality score S_005 is in a slightly noisy environment. In the spectral sub-modality, the emotional sensitivity index quality score is 0.82, the noise inconsistency index quality score is 0.45, the comprehensive index characterization value quality score is 0.073, and the emotional sensitivity dynamic benchmark factor is 0.8; the noise inconsistency dynamic benchmark factor is 0.48. In the prosody sub-modality, the emotional sensitivity index quality score is 0.76, the noise inconsistency index quality score is 0.50, the comprehensive index characterization value quality score is 0.021, and the emotional sensitivity dynamic benchmark factor is 0.7; the noise inconsistency dynamic benchmark factor is 0.55. In the sound quality sub-modality, the emotional sensitivity index quality score is 0.60, the noise inconsistency index quality score is 0.58, the comprehensive index characterization value quality score is -0.105, and the emotional sensitivity dynamic benchmark factor is 0.55; the noise inconsistency dynamic benchmark factor is 0.7.
[0042] Table 1. Characterization data of the comprehensive index of submodal speech emotion representation generation.
[0043]
[0044] like Figure 3The scatter plot of the relationship between submodal sentiment indices and benchmark factors, along with the histogram at the diagonal, reveals the numerical clustering patterns of each parameter: the sentiment sensitivity index mainly ranges from 0.6 to 0.8, the noise inconsistency index mainly ranges from 0.3 to 0.5, the comprehensive index value is concentrated between -0.2 and 0.2, the dynamic benchmark factor for sentiment sensitivity ranges from 0.6 to 0.8, and the dynamic benchmark factor for noise inconsistency ranges from 0.25 to 0.75. The sentiment sensitivity index and the comprehensive index value show a significant positive correlation, while the sentiment sensitivity index and the dynamic benchmark factor for sentiment sensitivity show a weak positive correlation. The noise inconsistency index and the comprehensive index value show a significant negative correlation, while the noise inconsistency index and the dynamic benchmark factor for noise inconsistency show a weak positive correlation. The comprehensive index value and the dynamic benchmark factor for sentiment sensitivity show a positive correlation, while the comprehensive index value and the dynamic benchmark factor for noise inconsistency show a negative correlation. It presents the distribution characteristics of each core parameter and verifies the correlation between parameters, providing data support for submodal performance optimization and parameter threshold adjustment.
[0045] In this implementation scheme, by performing deep encoding, confidence scoring and comprehensive index representation calculation on spectrum, prosody and phonological submodal, and combining gated collaborative attention to complete weighted fusion, the unified speech emotion representation achieves an adjustable balance between emotion sensitivity and noise robustness, providing a structurally stable and weighted interpretable high-quality emotion representation foundation for feature selection and emotion classification.
[0046] Specifically, the process of pre-screening the speech emotion representation vector set and initializing the random forest weights is as follows: Input the speech emotion representation vector set and the corresponding submodal emotion influence factors and noise interference factors; perform ReliefF importance pre-screening and random forest weight initialization on the feature space; construct a candidate subset of emotion features and an initial feature selection vector; select a training subset from the standardized speech emotion sample library according to the sample number and dataset partition label; use the speech emotion representation vector as the input feature; use the emotion category label, emotion intensity level, and urgency label as the supervision label to form an emotion feature sample set; based on the submodal emotion influence factors and noise interference factors, add the corresponding submodal label and submodal factor weight to each dimension of the feature. The ReliefF candidate feature set is obtained by performing importance scoring and irrelevant feature removal on sentiment features. For each training sample, a feature space metric is constructed based on the sentiment representation vector, and submodal sentiment influence factor and noise interference factor are introduced as weight coefficients in the metric. The improved ReliefF algorithm is used to calculate the effect of the difference of each feature dimension on the classification boundary in the neighborhood of samples of the same class and samples of different classes. The degree of maintaining similarity in samples of the same class and increasing the distance in samples of different classes is mapped to the feature importance weight. Features with negative importance weights are directly judged as irrelevant or features that destroy sentiment discrimination and are removed. Features with importance weights close to zero and noise interference factors higher than the average level are set as weakly relevant features and retained with low priority. Features with importance weights greater than zero are sorted from largest to smallest to form a preliminary candidate feature set and importance ranking list. Importance assessment and initial selection vector construction are performed. Using the ReliefF candidate feature set as input, a random forest sentiment classification model containing multiple decision trees is constructed. During training, the purity gain or error rate reduction resulting from each feature's participation in splitting on different trees is recorded. The gains across all trees are accumulated and normalized to obtain the random forest feature importance index. Combining the ReliefF weights and the random forest importance index, a dual-source fusion importance score is constructed. Each feature in the ReliefF candidate feature set is mapped to one bit in the initial binary selection vector according to its dual-source fusion importance score. Feature encoding and fitness evaluation interfaces are generated. The ReliefF candidate feature set is sorted by importance to determine the feature index order. Each initial feature selection vector is used as the initial position set for the wolf pack, recording the dual-source importance score, submodal sentiment influence factor, and noise interference factor. The submodal attention weight directly determines the submodal's participation in feature selection. Submodalities with weights higher than the weight threshold are retained for feature selection, while those lower than the weight threshold are eliminated or suppressed.
[0047] In this implementation scheme, by introducing ReliefF pre-screening that combines submodal sentiment influence factors and noise interference factors, random forest dual-source importance fusion, and participation control of submodal attention weights before feature selection, we can focus more on feature dimensions with high contribution and low interference. This reduces the search dimensions and optimization difficulty, while providing an initial feature foundation for feature subset optimization and sentiment classification performance improvement.
[0048] Specifically, the process of obtaining a subset of speech emotion features through two-order mutation of gray wolves and mapping evaluation is as follows: Inputting a candidate feature set, an initial feature selection vector, and a joint fitness evaluation interface, a two-order mutation of gray wolves optimization algorithm is used to perform global search and local refinement in the candidate feature space, constructing a mapping evaluation mechanism. In the first stage of the gray wolf global search, the initial feature selection vector is used as the initial position of the gray wolf group, and the binary encoding vector of each gray wolf is randomly perturbed to form an initial feature subset. Based on the joint fitness evaluation interface, a lightweight emotion classifier is trained on the feature subset corresponding to each gray wolf, calculating the classification error rate and feature dimension ratio. Simultaneously, combining the submodal emotion influence factor and noise interference factor, the proportion of emotion-influencing submodals and noise interference submodals in the currently selected features is statistically analyzed to form a comprehensive fitness value. In the first stage of iteration, the gray wolf position update parameters are set to favor the exploration mode. By increasing the search step size of the globally optimal individual and the current individual, a mutation operator targeting the selected feature position is introduced to encourage the algorithm to explore the subset space containing new feature combinations. In the second stage of gray wolf optimization, the initial feature subset obtained in the first stage undergoes local refinement and factor-driven structural adjustment: a feature subset with fitness values in the top several percentiles from the first stage iteration history is selected as a seed solution to form the initial position of the gray wolf population in the second stage; in the second stage iteration, the exploration intensity of the gray wolf position update parameters is reduced, and a local search strategy centered on the difference vector between seed solutions and feature positions that appear frequently in historical seed solutions is introduced, and the feature subset is finely adjusted through range position flipping and pattern recombination; gray wolf optimization influence factors and classifier performance influence factors are introduced into the fitness calculation; the classifier performance influence factor is formed by comprehensively considering the accuracy, recall, and stability indicators of the current feature subset on historical running data.
[0049] The classification modifier is obtained by multiplying the performance change value by the hyperbolic tangent and the performance parameter; the sentiment modifier is obtained by multiplying the sentiment deviation value by the influence parameter; the interference modifier is obtained by multiplying the noise deviation value by the interference parameter; the influence modifier is obtained by multiplying the gray wolf influence value by the optimization parameter; the performance modifier and the sentiment modifier are added together, and the noise modifier and the influence modifier are subtracted to obtain the power; the power is then exponentially calculated to obtain the update index; the quality score value of the feature subset in the nth evaluation is obtained by multiplying the quality score value from the previous evaluation by the update index; the specific calculation formula for the quality score value is as follows:
[0050]
[0051] ;
[0052] ;
[0053] ;
[0054] ;
[0055] In the formula, This represents the quality score of the feature subset at the nth evaluation, used to quantify the overall quality level of the feature subset in the current round; The performance change value is obtained by dividing the difference between the classifier's performance value in the current round and the previous round by the sum of the classifier's performance factor parameter and the numerical stability term. It is used to quantify the relative change in the classifier's performance between two adjacent rounds. The emotional deviation value is obtained by dividing the difference between the current emotional influence factor and the emotional threshold by the sum of the emotional influence factor parameters and the numerical stability term. It is used to quantify the relative deviation between the current emotional influence factor and the emotional threshold, and to provide data support for parameter adaptation and interference suppression of emotional-related modules. The noise deviation value is obtained by dividing the difference between the current noise interference factor and the interference threshold by the sum of the noise interference factor parameter and the numerical stability term. It is used to quantify the relative degree to which the current noise interference factor exceeds the interference threshold, and provides a quantitative reference for the selection of anti-interference strategies and the adjustment of noise suppression algorithm parameters. The gray wolf influence value is obtained by dividing the difference between the current gray wolf optimization influence factor and the influence threshold by the sum of the gray wolf optimization influence factor parameters and the numerical stability term. It is used to quantify the relative deviation between the current gray wolf optimization influence factor and the influence threshold, providing data basis for algorithm parameter adjustment and optimization strategy adaptation, and ensuring the optimization efficiency and stability of the gray wolf optimization algorithm in the target scenario. This represents the quality score from the previous evaluation round, used to achieve multiplicative iterative updates of the score. The classifier performance value for the current round is obtained by normalizing the accuracy, F1 and latency metrics, and is used to reflect the performance level of the classifier in the current round. This represents the classifier performance value from the previous round, which is obtained by normalizing the accuracy, F1 score, and latency metrics, and is used to compare the performance with that of the current round classifier. The sentiment influence factor of the current feature subset is calculated by performing sentiment classification performance analysis on the feature subset in historical running data, combined with the improvement effect of the feature on the classifier performance. It is used to reflect the strength of the positive effect of the current feature subset on the sentiment dimension. The noise interference factor representing the current feature subset is calculated by analyzing the false alarm rate, false negative rate, and performance degradation value of the feature subset under different noise scenarios and channel conditions. It is used to characterize the degree of noise interference to the current feature subset. The gray wolf optimization influence factor in the current iteration stage is calculated by the population fitness changes, the stability of feature selection results and distribution differences during the gray wolf optimization process. It is used to reflect the state deviation of the current feature subset in the gray wolf optimization process. This represents the emotional threshold, used to measure the degree of deviation from a baseline. This represents the interference threshold, used to measure the degree of deviation from the baseline; This represents the gray wolf influence threshold, used to measure the degree of deviation from the baseline; This represents the classifier performance factor parameter, which is obtained by normalizing the maximum and minimum values of the classifier performance over time. The value range is 0.1-0.5, and it is used to normalize the difference values of the classifier performance. The emotional impact factor parameter is calculated by measuring the range of variation of the emotional impact factor with different submodalities. The value range is 0.05-0.3, and it is used to normalize the difference between the emotional impact factor and the reference threshold. The noise interference factor parameter is obtained by normalizing the maximum and minimum values of the noise interference factor as it varies with different scenarios and channel conditions. The value range is 0.1-0.4, and it is used to normalize the difference between the noise interference factor and the reference threshold. The parameter representing the gray wolf optimization impact factor is calculated by measuring the range of change of the gray wolf optimization impact factor during the feature subset optimization process. The value range is 0.08-0.35, and it is used to normalize the difference between the gray wolf optimization impact factor and the reference threshold. This represents a performance parameter, obtained through sensitivity analysis of classifier performance, with a value range of 0.5-2.0, used to adjust the intensity of the impact of classifier performance differences on quality score updates; This represents the influence parameter, which evaluates the contribution of the emotion influence factor to the effect in the emotion recognition process. The value ranges from 0.3 to 1.8 and is used to adjust the influence strength of the emotion influence factor on the quality score update. The interference parameter is obtained by sweeping different values across multiple noise scenarios and combining them with cross-validation for noise robustness. The value range is 0.6-2.2, and it is used to adjust the influence of the noise interference factor on the quality score update. The influencing parameter is obtained by sweeping different values in the Grey Wolf optimization iterative experiment, combined with convergence speed and cross-validation. The value range is 0.4-1.9, which is used to adjust the influence strength of the Grey Wolf optimization influencing factor on the quality score update. This represents a numerically stable term, obtained by selecting a very small positive constant with a value range of 10−6-10−3, used to prevent the denominator from being zero during the calculation process.
[0056] When the quality score is higher than the upper threshold, it is determined that the current feature subset has reached a balance between preserving emotionally sensitive features and suppressing noise. The feature subset is marked as a candidate optimal solution and the search intensity is reduced. When the quality score falls between the upper and lower thresholds, the current search direction remains unchanged, with only a moderate increase in local mutation operations. When the quality score is lower than the lower threshold or the noise interference factor increases, a restart strategy is triggered, re-entering the first stage or strengthening the global search component. When the iteration termination condition is met, the feature subset with the highest comprehensive quality score is output as the speech emotion feature subset. The termination conditions include: the number of iterations in the Grey Wolf optimization reaches the maximum iteration threshold, the improvement in the comprehensive quality score is lower than the convergence threshold within several consecutive iterations, and the classifier performance index on historical running data shows no improvement within several consecutive iterations. At least one of the three conditions is met. The corresponding binary selection vector is written into the feature selection label field to record the factor evaluation results associated with the feature subset, such as... Figure 5The diagram shows the overall framework of the fusion feature selection algorithm. The framework can be divided into three parts from top to bottom: In the feature extraction stage, the original dataset is used as input, and the original feature vector is obtained through the feature extraction module. All features undergo uniform scaling and encoding processing to form a standardized feature matrix. In the filtering stage, the standardized dataset is divided into training and test sets using a ten-fold cross-validation method. On one hand, the training set is fed into the ReliefF algorithm to evaluate and rank the importance of each feature dimension, filtering out candidate feature subsets and constructing a dataset-feature vector. On the other hand, both the standardized complete feature vector and the candidate feature subsets can be used with the base classifier. The packaging method stage initializes the population based on the candidate feature subset output by ReliefF, encoding each gray wolf individual into a binary feature selection vector. The population position is continuously updated in the two-order mutation gray wolf optimization module. In each generation, a corresponding dataset-feature vector is constructed based on the features selected by the current individual, which is then fed into the classifier to obtain the classification accuracy. The fitness function value is then calculated by combining the feature dimension information, and it is determined whether the stopping criterion is met. When the stopping criterion is met, the feature subset with the highest comprehensive quality score is output from the evolutionary history, and the corresponding final dataset-feature vector is generated, realizing the fusion feature selection process.
[0057] In this implementation scheme, by combining two-order mutant gray wolf optimization with multi-factor quality score mapping, the retention of emotion-sensitive features, suppression of noise interference and improvement of classification performance are dynamically balanced during the feature selection process. This achieves adaptive optimization and stable convergence of feature subset quality, providing a feature combination with high discriminative power, low redundancy and high performance matching for speech emotion recognition.
[0058] Specifically, the process of performing speech emotion recognition and confidence decision on a subset of speech emotion features is as follows: Constructing a task-based emotion classification model: Inputting the subset of speech emotion features and factor evaluation results, performing emotion recognition and confidence decision on a standardized speech emotion sample library, outputting emotion category, emotion intensity, and decision confidence results; reading the currently effective feature subset binary selection vector and corresponding feature index from the feature selection label field; filtering the speech emotion representation vector according to the feature selection vector, retaining only the selected feature dimensions to form a simplified feature vector; combining the factor evaluation results, adding corresponding feature subset quality scores, emotion influence factors, noise interference factors, and classifier performance influence factors to each sample. Using the simplified feature vector as input, and emotion category labels, emotion intensity levels, and urgency level labels as multi-task outputs, an emotion classification network is constructed; the shared feature extraction layer can adopt a multilayer perceptron structure; the multi-head output layer outputs the emotion category probability distribution, emotion intensity regression value or grading result, urgency level, and internal consistency score respectively; in the loss function design, the emotion category cross-entropy loss, emotion intensity regression loss, and urgency level grading loss are combined through task weights. A sample-level weighting and regularization constraint mechanism is introduced. The quality score of the feature subset is used as the sample weight. Samples with quality scores above the average level are given higher weight in the loss calculation, while samples with quality scores below the average level and noise interference factors above the average level are given lower weight. This reduces the negative impact of noise-dominant samples on the parameters of the task sentiment classification model. The submodal sentiment influence factor is used to construct a regularization constraint term. When a submodality is retained for a long time in the feature selection stage and its weights show abnormal decay in the training of the current task sentiment classification model, a lightweight regularization that maintains the contribution is added to the weights of the submodality. The classifier performance influence factor is combined with the real-time inference latency or feature dimension ratio to automatically search or prune network structure parameters (such as hidden layer width and output head complexity). Confidence decision and quality labeling of emotion recognition results based on the output of the task emotion classification model and the quality information of feature subsets: Combine the corresponding feature subset quality score, submodal noise interference factor and sharpness or entropy value of the output probability distribution to calculate the overall decision confidence and give the result quality label; Mark the results with decision confidence below the business scenario threshold and send them into the delayed decision process; Trigger the rapid linkage strategy in emergency rescue scenario and driving safety scenario for results with confidence above the business scenario threshold.
[0059] In this implementation scheme, a sample-level weighting and regularization constraint mechanism based on feature subset quality scores, submodal factors, and classifier performance influence factors is introduced into the task emotion classification model. This enables the joint recognition process to adaptively distinguish between high-quality and noise-dominated samples, reducing interference from invalid samples while improving the accuracy, stability, and decision confidence of multi-task emotion recognition results.
[0060] Specifically, the specific process of emotion recognition service orchestration and operation feedback adaptive update is as follows: As Figure 5 shown in the flowchart of the method for emotion recognition service orchestration and operation feedback adaptive update of voice, the emotion recognition result, decision confidence, and feature subset quality score value are input. When facing different business scenarios, emotion recognition services are orchestrated and linked, and the data during the operation process is fed back to achieve continuous adaptive optimization. An independent emotion service configuration file is established for each business scenario. The configuration file includes: a set of available emotion categories, emotion intensity and urgency thresholds, business action strategies corresponding to different emotion states, and the maximum allowable in-chip inference delay and the minimum decision confidence lower limit. According to the business scenario type label in the voice emotion acquisition sample frame, the corresponding emotion service configuration file is automatically selected to perform scenario-based interpretation and action mapping on the emotion category, emotion intensity, and urgency. During the emotion recognition service orchestration process, the linkage control and feedback record between the emotion recognition result and the upper-layer business system are realized. For example, in the real-time voice monitoring scenario of the emergency center, when the recognition result shows that the caller is in a tense or panicked state and the urgency is higher than the emergency threshold, the result is pushed to the emergency dispatch assistance system to联动 adjust the dispatch priority, remind the operator to use specific soothing words or quickly confirm key information. Record the emotion recognition result, decision confidence, corresponding feature subset version number, operation delay, and business action execution result for each online inference request to form an emotion recognition operation log. Summarize the manually reviewed samples, emotion recognition results with explicit user feedback errors, and cases marked as false positives or missed detections in the business system to construct a difficult sample pool. During regular offline training, the difficult sample pool is merged with the training set and historical operation data in layers to update the factor mapping threshold and gray wolf optimization strategy. And trigger the update of the task emotion classification model and feature strategy, and update the parameters of the task emotion classification model, feature selection strategy, sub-modal weight, and optimization algorithm. The fed-back data undergoes quality screening to remove mislabeled and abnormal samples, ensuring that the fed-back data has a positive promotion effect on the optimization process and avoiding frequent oscillations and performance fluctuations caused by invalid data. When the cumulative number of running samples exceeds the sample threshold or new acquisition scenarios and devices cause changes in noise characteristics, an offline retraining process is automatically triggered to re-execute feature subset optimization and update the task emotion classification model. After the update is completed and passed the historical operation data and small-scale tests, the new version of the feature selection configuration and model parameters are marked as available states and gradually released in a gray manner in each scenario. At the same time, the new version number and effective time are written into the voice emotion acquisition sample frame and the feature metadata table to achieve continuous iteration and stable operation under different scenarios and noise conditions.
[0061] In this implementation plan, by orchestrating emotion recognition services according to business scenarios and combining the feedback update mechanism of operation logs and difficult sample pools, the emotion service strategy, task emotion classification model and feature selection configuration can be continuously iterated and dynamically optimized in practical applications, maintaining the long-term stability, adaptability and business linkage effect of the speech emotion recognition system under multiple scenarios and multiple noise conditions.
[0062] Specifically, the second aspect of this invention provides a speech emotion recognition system based on multimodal feature fusion, applied to a speech emotion recognition method based on multimodal feature fusion, comprising: a speech data acquisition and emotion annotation module, used to acquire speech emotion datasets, perform quality screening, scene and label unified encoding and normalization processing on the speech emotion datasets, and maintain sample numbers, version numbers and dataset partitioning labels during the acquisition and annotation stages, and perform unified temporal labeling and hierarchical management of multi-source scene speech samples to form a standardized speech emotion sample library; an emotion feature construction module, used to perform frame-level acoustic and spectral feature construction on the standardized speech emotion sample library, and organize feature channels according to three submodalities: spectrum, prosody and timbre, calculate the corresponding submodal confidence scores and write them into the feature metadata table to obtain the submodal emotion feature vector set; and a feature fusion and learning module, used to perform submodal emotion feature fusion and learning based on the submodal emotion feature vectors. A speech emotion representation generation method is proposed, which integrates multi-submodal deep coding and gated collaborative attention mechanism. During the unified speech emotion representation generation process, submodal attention weights, gating states, and intermediate representation indices are recorded, and the fused speech emotion representation is associated and stored with the corresponding submodal weights. An emotion feature selection and adaptive optimization module is used to pre-screen the speech emotion representation vector set and initialize random forest weights. Two-order mutation gray wolf and mapping evaluation are performed to obtain a subset of speech emotion features, which are registered in the feature selection label field in the form of binary selection vectors and factor evaluation results. An emotion recognition decision and output module is used to perform speech emotion recognition and confidence decision on the subset of speech emotion features, and to perform emotion recognition service orchestration and adaptive updates based on runtime feedback. According to the business scenario configuration file, the emotion recognition results, decision confidence, and feature subset version information are written to the runtime log and feedback queue.
[0063] In this implementation plan, an end-to-end closed-loop multimodal speech emotion recognition system is constructed by integrating multi-scenario speech acquisition and annotation, submodal feature construction and deep fusion, quality-driven feature selection and adaptive optimization, and an integrated architecture for emotion recognition and service orchestration oriented towards business scenarios. This system achieves collaborative linkage and traceable management of the entire process from data acquisition and feature representation to decision output and online feedback updates.
[0064] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0065] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A speech emotion recognition method based on multimodal feature fusion, characterized in that, Includes the following steps: S1. Collect a voice emotion dataset, perform quality screening, scene and label unified encoding and normalization on the voice emotion dataset to form a standardized voice emotion sample library. The standardized voice emotion sample library includes voice emotion collection sample frames. S2, perform frame-level acoustic analysis and construct spectral features on the standardized speech emotion sample library, and construct prosodic and phonological emotion features and add confidence labels to obtain a submodal emotion feature vector set; The specific process of performing frame-level acoustic analysis and constructing spectral features on a standardized speech emotion sample library is as follows: Input a sample frame of voice emotion acquisition, perform frame-level acoustic and construct spectral features to generate voice emotion spectral submodal features; concatenate them to form frame-level acoustic and spectral feature vectors; index and manage all frame-level acoustic and spectral feature vectors of the same speech segment and store them in a structured manner to output a frame-level acoustic and spectral feature dataset; The specific process of constructing and labeling prosody and timbre-based emotional features to obtain the submodal emotional feature vector set is as follows: Input frame-level acoustic and spectral feature datasets, construct prosodic and tone quality emotion features, encapsulate them into submodal emotion feature vectors, and obtain the emotion sensitivity index by normalizing the gradient contribution and inter-class separability statistics between submodal features and emotion labels in historical running data; obtain the quality index by fusing and mapping the sample quality label, the signal-to-noise ratio estimate in the noise environment field, and the packet loss rate and distortion label in the channel state field. The quality difference term is obtained by squaring the difference between the quality reference threshold and the submodal quality index. The quality difference term is obtained by multiplying the square of the m-th submodal sentiment sensitivity index minus the quality penalty weight by the quality difference term. The quality adjustment term is then multiplied by the confidence sharpening coefficient. The reciprocal of the dynamic benchmark factor is taken as the sentiment-related term. The quality adjustment term is then negatively evaluated and exponentially calculated to obtain the exponential term. The sentiment-related term is multiplied by the exponential term to obtain the penalty comprehensive term. A fixed value is added to the penalty comprehensive term to obtain the scoring denominator. Finally, the fixed value is divided by the scoring denominator to obtain the confidence score of the m-th submodal. The confidence score is then appended to the submodal feature vector and encapsulated to form the submodal sentiment feature vector set. S3, a speech emotion representation generation method based on submodal emotion feature vector sets that integrates multimodal deep coding and gated collaborative attention mechanism; The specific process of the speech emotion representation generation method based on submodal emotion feature vector sets, which integrates multimodal deep coding and gated collaborative attention mechanism, is as follows: The system takes a set of submodal emotion feature vectors and the corresponding confidence scores of the submodal as input. It performs deep encoding and gated collaborative attention fusion on the spectral, prosodic, and phonological submodal features carried in each voice emotion collection sample frame to generate a unified voice emotion representation vector and outputs the submodal emotion influence factor and noise interference factor. The noise inconsistency index is obtained by statistically analyzing the distribution differences of submodal embeddings under different noise labels, channel states, and sample quality levels in multiple noise scenarios and under multiple channel conditions, and compressing and mapping the cross-scenario output distribution deviation. A submodal comprehensive index representation value is constructed based on the emotional sensitivity index and the noise inconsistency index. The sensitivity amplification coefficient is multiplied by the emotional sensitivity index of the k-th submodal to obtain the amplification term. After exponential operation on the amplification term, it is added to the dynamic benchmark factor of emotional sensitivity to obtain the emotional sensitivity term. Logarithmic operation is performed on the emotional sensitivity term to obtain the sensitivity term. The difference between the noise inconsistency index of the k-th submodal and the noise inconsistency reference threshold is obtained to obtain the deviation term. The deviation term is multiplied by the noise penalty amplification coefficient and then exponentially operated to obtain the penalty term. The penalty term is added to the dynamic benchmark factor of noise inconsistency to obtain the penalty consistency term. Logarithmic operation is performed on the penalty consistency term to obtain the noise term. The comprehensive index representation value of the k-th submodal is obtained by subtracting the noise term from the sensitivity term. A gated collaborative attention fusion mechanism is constructed to convert the comprehensive index representation value into submodal attention weights and influence factors through a monotonic sigmoid function to form a unified speech emotion representation vector. S4, pre-screen the speech emotion representation vector set and initialize the random forest weights, then perform two-order mutation gray wolf and mapping evaluation to obtain a subset of speech emotion features; The specific process of pre-screening the speech emotion representation vector set and initializing the random forest weights is as follows: The input speech emotion representation vector set and the corresponding submodal emotion influence factors and noise interference factors are used to perform ReliefF importance pre-screening and random forest weight initialization on the feature space to construct a candidate subset of emotion features and an initial feature selection vector. The importance of emotion features is scored and irrelevant features are removed to obtain a ReliefF candidate feature set, which is then evaluated for importance and the initial selection vector is constructed. Feature encoding and fitness evaluation interfaces are generated. The submodal attention weights directly determine the participation of submodalities in feature selection. Submodalities with weights higher than the weight threshold will be retained for feature selection, while submodalities with weights lower than the weight threshold will be removed or suppressed. S5 performs speech emotion recognition and confidence decision on a subset of speech emotion features, and performs emotion recognition service orchestration and adaptive updates based on operational feedback.
2. The speech emotion recognition method based on multimodal feature fusion according to claim 1, characterized in that: The specific process of collecting the voice emotion dataset and performing quality screening, scene and label unified encoding, and normalization on the voice emotion dataset to form a standardized voice emotion sample library is as follows: A speech emotion dataset is collected, which includes: original speech waveform dataset, emotion tag and scene label dataset, speaker attribute and conversation structure dataset, noise environment and channel state dataset, and sample partitioning and version management dataset. Speech emotion data sequences already stored in the same speech emotion recognition business scenario are used as historical operational data. The collected speech emotion dataset is labeled with data types and classified according to quality attributes. Uploaded speech segments are formatted and completed using a unified speech acquisition standard to obtain a processed speech emotion dataset. This processed dataset is then encapsulated with a unified data structure to construct speech emotion acquisition sample frames. Finally, z-score dimensionless normalization and encoding preprocessing are performed on the fields in the acquired sample frames to obtain a standardized speech emotion sample library.
3. The speech emotion recognition method based on multimodal feature fusion according to claim 1, characterized in that: The specific process of obtaining the subset of speech emotion features through two-order mutation gray wolf and mapping evaluation is as follows: Input the candidate feature set, initial feature selection vector, and joint fitness evaluation interface. Use the two-order mutation gray wolf optimization algorithm to perform global search and local refinement in the candidate feature space to construct a mapping evaluation mechanism. Multiply the hyperbolic tangent of the performance change value by the performance parameter to obtain the classification adjustment term. The emotional modulation term is obtained by multiplying the hyperbolic tangent of the emotional deviation value by the influence parameter. The noise deviation value is multiplied by the interference parameter to obtain the interference adjustment term; the gray wolf influence value is multiplied by the optimization parameter to obtain the influence adjustment term; the classification adjustment term and the sentiment adjustment term are added together, and the interference adjustment term and the influence adjustment term are subtracted to obtain the power term; the power term is exponentially calculated to obtain the update index; the quality score value of the feature subset in the previous evaluation is multiplied by the update index to obtain the quality score value of the feature subset in the nth evaluation. When the quality score is higher than the upper threshold, it is determined that the current feature subset has reached a balance between the preservation of emotionally sensitive features and the ability to suppress noise; when the quality score falls between the upper and lower thresholds, the current search direction remains unchanged; when the quality score is lower than the lower threshold or the noise interference factor increases, the restart strategy is triggered; and the speech emotion feature subset is output.
4. The speech emotion recognition method based on multimodal feature fusion according to claim 1, characterized in that: The specific process of performing voice emotion recognition and confidence decision on a subset of voice emotion features is as follows: Construct a task-based emotion classification model: Input a subset of speech emotion features and factor evaluation results, perform emotion recognition and confidence decision on a standardized speech emotion sample library, and output emotion category, emotion intensity and decision confidence results. Introduce a sample-level weighting and regularization constraint mechanism, and perform confidence decision and quality labeling of emotion recognition results based on the output of the task-based emotion classification model and the quality information of the feature subset.
5. The speech emotion recognition method based on multimodal feature fusion according to claim 4, characterized in that: The specific process of orchestrating and adaptively updating the emotion recognition service based on operational feedback is as follows: Input emotion recognition results, decision confidence, and quality scores of feature subsets; orchestrate and link emotion recognition services in different business scenarios; and feed back data during operation to achieve continuous adaptive optimization. During the orchestration of the emotion recognition service, the linkage control and feedback recording of emotion recognition results with the upper-layer business system are realized, and the task emotion classification model and feature strategy are updated to achieve continuous iteration and stable operation under different scenarios and noise conditions.
6. A speech emotion recognition system based on multimodal feature fusion, employing the speech emotion recognition method based on multimodal feature fusion as described in any one of claims 1-5, characterized in that, include: The speech data acquisition and emotion annotation module is used to collect speech emotion datasets, perform quality screening, scene and label unified encoding and normalization on the speech emotion datasets, and form a standardized speech emotion sample library. The emotion feature construction module is used to perform frame-level acoustic and spectral feature construction on a standardized speech emotion sample library, and to construct prosodic and phonological emotion features and add confidence labels to obtain a submodal emotion feature vector set; The feature fusion and learning module is used for a speech emotion representation generation method based on the fusion of multi-submodal deep coding and gated collaborative attention mechanism using submodal emotion feature vector sets; The emotion feature selection and adaptive optimization module is used to pre-screen the speech emotion representation vector set and initialize the random forest weights, and to obtain a subset of speech emotion features by performing two-order mutation gray wolf and mapping evaluation. The emotion recognition decision and output module is used to perform speech emotion recognition and confidence decision on a subset of speech emotion features, and to perform emotion recognition service orchestration and adaptive updates based on operational feedback.
Citation Information
Patent Citations
A Natural Speech Emotion Recognition Method Based on Multimodal Deep Feature Learning
CN111583964B
Artificial intelligence-based speech emotion recognition method, device, equipment and medium
CN115312033B
Voice emotion recognition method and device based on artificial intelligence, equipment and medium
CN116844573A
Call center dialogue sentiment analysis method and system fusing voice and text
CN121217864A