Dialogue type speech emotion automatic labeling method based on graph language model

Through the dialogue-based speech-based automatic labeling method of speech-based emotions based on the graph language model, combined with large-scale pre-trained language model and adaptive technology, the accuracy and cost of emotional recognition in government service hotline calls are solved, and efficient and accurate emotional labeling and robust recognition are achieved.

CN120279947APending Publication Date: 2025-07-08HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510383616.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing voice emotion recognition methods are difficult to capture subtle emotional changes in government service hotline calls, and traditional manual labeling is expensive, difficult to meet the needs of large-scale data analysis, and poor cross-situation adaptability.

Method used

The dialogue-type speech emotion automatic labeling method based on the graph language model is adopted. Through the combination of acquisition, preprocessing, text and speech emotion recognition models, large-scale pre-trained language model (LLM) and multi-task learning modules are used, and the adaptive labeling technology is combined, accent adaptive modules are designed for different accents to optimize the emotional prediction results.

Benefits of technology

It realizes efficient and accurate emotional annotation, reduces manual annotation costs, improves labeling efficiency and accuracy, enhances robustness and adaptability, and can accurately capture subtle emotional changes in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279947A_ABST
    Figure CN120279947A_ABST
Patent Text Reader

Abstract

The invention relates to a dialogue type voice emotion automatic labeling method and system based on a graph language model, and aims to solve the problems of low efficiency, high subjectivity and high cost of artificial voice emotion labeling in government affair service hotlines. The method comprises the following steps: collecting diversified government affair service hotline dialogue voice data; preprocessing the collected data, and converting the collected data into voice-text bimodal data; extracting text emotion features by using a large language model and optimizing a classification result through multi-task learning; inputting the map sequence into a voice emotion recognition model, and generating a voice emotion predicted value through a backbone network and an emotion relationship mining module; and carrying out confidence score evaluation on the text and speech emotion prediction result, if the confidence score is higher than a preset threshold, constructing a high-quality emotion annotation data set, otherwise, iteratively optimizing the model. The text and voice multi-modal fusion is realized, the emotion labeling automation and accuracy are remarkably improved, the cost and subjectivity are reduced, and powerful support is provided for large-scale emotion analysis of dialogue scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of adaptive annotation, relates to the task of speech emotion recognition, and particularly relates to a system for efficiently recognizing emotions in speech signals using large language models and language emotion recognition technology. Background Art

[0002] In the scenario of government service hotline calls, accurately identifying the speech emotion information of citizens has become an important topic for improving service quality and optimizing user experience. However, existing speech emotion recognition methods mainly rely on traditional manual annotation or models based on paragraph-level emotion labels. These methods are difficult to capture the subtle emotion changes in the speech signals during the calls, resulting in insufficient accuracy and real-time performance of emotion recognition. With the continuous growth of the amount of government service hotline call data, the traditional manual annotation method is not only costly but also difficult to meet the needs of large-scale data analysis, becoming a key bottleneck restricting the development of emotion recognition technology. Therefore, researching and developing a method for automatically annotating dialogue-based speech emotions based on a graph language model is of great significance for real-time perception of citizens' emotional states, optimizing the service response strategies of operators, and improving the quality of government services. When dealing with the complexity and diversity of contexts in government calls, existing methods often show limitations such as insufficient capture of emotion intensity and poor cross-situational adaptability, further highlighting the necessity of innovative technologies. Summary of the Invention

[0003] The present invention aims to solve the technical problems existing in existing speech emotion recognition methods, such as difficult data annotation and low accuracy of emotion classification, and proposes a method for automatically annotating dialogue-based speech emotions based on a graph language model. This method can efficiently and accurately complete the annotation of speech emotion data, significantly reduce the subjectivity in the manual annotation process, avoid the uncertainty of emotion categories, and at the same time effectively improve the accuracy of speech emotion recognition. In addition, by introducing adaptive annotation technology, the present invention not only significantly reduces the manual annotation cost but also greatly improves the annotation efficiency and quality. It is worth mentioning that, aiming at the problem of decreased accuracy of emotion recognition caused by accent differences between the two parties in the dialogue, this method designs an accent adaptation module based on a language model, which can intelligently identify and adapt to different accent features, thereby further improving the robustness and accuracy of speech emotion recognition.

[0004] The technical solution adopted by the present invention to solve its technical problems is: a method for automatically annotating dialogue-based speech emotions based on a graph language model, including the following steps:

[0005] Step 1: Through the government service hotline system, comprehensively collect the dialogue speech data between the operator and the citizen, ensure the authenticity and diversity of the data source, and write it into the speech data set D voice , providing a reliable data basis for subsequent emotion analysis;

[0006] Step 2: Preprocess the collected dialogue voice data D voice including operations such as noise reduction, speech segmentation, and spectrum processing. Meanwhile, convert the voice data into corresponding text data T to form bimodal data D containing voice signal information and text content bimodal ;

[0007] Step 3: Input the preprocessed text data T into a large-scale pre-trained language model (LLM) to perform sentiment analysis on the text information. The model first extracts the sentiment features F in the text based on the prompt words text , then optimizes and enhances these features through a multi-task learning module, and finally generates the text sentiment prediction result R through the sentiment classification layer text , and calculate the loss value L with the true value Y text ; text ;

[0008] Step 4: Input the transformed spectrogram sequence I spec into a voice sentiment recognition model to perform sentiment analysis on the voice information. The model first extracts the sentiment features in the voice signal through the backbone network, and these features will then be input into the sentiment relationship mining module to capture more subtle sentiment changes, and finally generate the voice sentiment prediction result R through a multi-layer perceptron audio , and calculate the loss value L with the true value Y audio ; audio ;

[0009] Step 5: Compare the text sentiment prediction result R text and the voice sentiment prediction result R audio generated in Step 3 and Step 4. If the prediction results of the two are the same, it is considered that the sentiment prediction result is credible and directly perform sentiment annotation; otherwise, re-input the data with different predictions into the text information understanding module and the voice sentiment recognition module to iteratively optimize the prediction performance of the model until the required confidence level is reached. By repeatedly performing the annotation, evaluation, and optimization processes, a high-quality sentiment annotation dataset is finally constructed;

[0010] Collect the dialogue data between citizens and operators through the government service hotline system. The collection process includes capturing the operator's voice signal and collecting the citizen's voice signal. The specific steps are as follows:

[0011] Step 1.1: Capture the operator's voice signal. Use a high-sensitivity microphone to capture the operator's voice signal in real time and directly convert it into a digital signal. The signal collected by the microphone is recorded by the collection system to form clear digital voice data. The signal conversion process is shown as follows:

[0012] S′ agent =f mic (Sagent ), (1)

[0013] Among them, S' agent is the processed voice digital signal of the operator, and f mic is the microphone signal processing function.

[0014] Step 1.2: Collection of citizen voice signals. The voice signals of citizens are transmitted through the telephone network and automatically converted into digital signals by the telephone system. The digital signals are then recorded by the collection system to generate the voice data of citizens. The signal conversion process is shown as follows:

[0015] S' citizen = f tel (S citizen ), (2)

[0016] Among them, S' citizen is the processed voice digital signal of citizens, and f tel is the signal conversion function of the telephone system.

[0017] A series of preprocessing operations are performed on the collected voice data, including noise removal, speech segmentation, emotion alignment, and feature extraction, etc. The purpose is to improve the quality of the voice data and convert it into a standardized input suitable for emotion annotation and analysis. Through these preprocessing steps, it is ensured that the data can provide high-quality input features for the subsequent emotion recognition model. The specific steps are as follows:

[0018] Step 2.1: Noise removal. Use a noise suppression algorithm to remove background noise and eliminate static and dynamic noise parts. For complex background noise, an adaptive filter is used to further process the voice signal. The denoising process is shown as follows:

[0019] S clan (t) = f WienerFilter (S noisy (t)), (3)

[0020] Among them, S clean (t) represents the clear voice signal after denoising, f WienerFilter is the used noise suppression algorithm, and S noisy (t) is the original noisy voice signal.

[0021] Step 2.2: Speech signal segmentation. Perform frame-level segmentation on the clear voice signal [d1, d2, …, d n , and each frame of the signal is windowed through a window function to form short-time signal segments [w1, w2, …, w n . Then, for each frame of the signal [w1, w2, …, w nPerform short-time Fourier transform to generate corresponding time-frequency segments. The specific time-frequency analysis formula is:

[0022]

[0023] where S stft (f, t) represents the speech signal after time-frequency analysis, and f STFT is the short-time Fourier transform operation.

[0024] Step 2.3: Speech transcription. Process the processed speech signal [w1, w2, …, w n frame by frame, and convert each frame of the signal into corresponding text information [t1, t2, …, t n . For unclear speech segments, combine the context information to perform word prediction, and finally obtain the complete transcribed text. The transcription process can be expressed as:

[0025] T = f transcription (W), (5)

[0026] where T = [t1, t2, …, t n represents the transcribed text sequence, W = [w1, w2, …, w n is the speech signal sequence, and f transcription is the conversion function from speech to text.

[0027] Step 2.4: Spectrogram conversion. Extract the Mel-frequency cepstral coefficients (MFCCs) from the processed speech signal [w1, w2, …, w n , and use the Mel filter bank to perform feature conversion on the signal to obtain the spectrogram sequence [F1, F2, …, F k . The specific spectrogram conversion process can be expressed as:

[0028] F mfcc = f MFCC (W), (6)

[0029] where F mfcc = [f1, F2, …, F k is the extracted spectrogram feature sequence, and f MFCC is the Mel-frequency cepstral coefficient extraction function.

[0030] Perform sentiment analysis on the transformed text data through the text information understanding module. Extract text sentiment features through a large-scale pre-trained language model (LLM), and optimize the features through a multi-task learning module. Finally, generate sentiment prediction results and calculate the loss value. The specific steps are as follows:

[0031] Step 3.1: Text data preprocessing and prompt design. First, perform task-specific prompt design on the preprocessed text data and input it into the BERT feature extractor. During this process, design specific prompts to guide the large language model BERT to extract the sentiment features in the text. The prompts can be designed as "Is the sentiment in this passage positive, neutral, or negative?" or as a guiding one like "Is the sentiment in this passage a positive emotion?" The prompt design formula can be expressed as:

[0032] T input = T 提示词 + T, (7)

[0033] where T input is the complete prompt and text combination input into BERT, and T is the transformed text data.

[0034] Step 3.2: Sentiment feature extraction. Input the designed prompts and text data into the BERT model, and extract the sentiment features in the text through the pre-trained model and self-attention mechanism. The BERT model establishes the connection between words through context information and uses multiple layers of Transformer encoders for feature extraction. Specifically, BERT captures the sentiment information in the text by mapping the input sequence to the hidden layer representation. The process of the feature extraction is as follows:

[0035] H = BERT(T input ), (8)

[0036] where H is the hidden layer feature extracted by the BERT model, and T input is the text input with prompts.

[0037] Step 3.3: Multi-task learning. Input the extracted sentiment features into the multi-task learning module for optimization and enhancement. The multi-task learning module enhances the model's learning ability of sentiment features by introducing tasks such as sentiment intensity prediction task and sentiment turn recognition task. The multi-task learning process can be expressed as:

[0038] L T = λ1L1 + λ2L2 + λ3L3, (9)

[0039] where L T is the sum of the losses of each task, L i is the loss value of each task, and λ i is the weight of the corresponding task.

[0040] Step 3.4: Sentiment Classification and Prediction. The optimized features output by the multi-task learning module are input into the sentiment classification layer to generate the sentiment prediction result of the text. The sentiment classification layer usually uses the Softmax function to predict the sentiment category and outputs the probability value of each sentiment category. The formula for the sentiment classification process is as follows:

[0041]

[0042] where represents the sentiment category prediction result, W h is the optimized feature matrix, b is the bias term, and Softmax represents the normalization calculation for each sentiment category.

[0043] Step 3.5: Backpropagation and Model Optimization. The loss value is backpropagated into the model through the backpropagation algorithm, and the model parameters are updated using the gradient descent method. The specific update process is as follows:

[0044]

[0045] where θ is the model parameter, η is the learning rate, is the gradient of the loss function with respect to the model parameter.

[0046] The spectrogram sequence F mfcc = [F1, F2, …, F k is processed by the speech emotion recognition model. The model first generates the preliminary features of the speech signal through the discriminative feature generation module, then extracts the emotion features through the backbone network, and then uses the emotion relationship mining module to capture the subtle emotion changes in the speech signal. Finally, the model generates the emotion prediction result through the decision Token selection and calculates the loss value for optimization. The specific steps are as follows:

[0047] Step 4.1: Discriminative Feature Generation. The input spectrogram sequence F mfcc = [F1, F2, …, F k is processed by the discriminative feature generation module to generate the preliminary features of the speech signal. This module uses a convolutional neural network to extract the time-frequency features and obtains the feature map where C represents the number of channels, and H′ and W′ are the height and width of the feature map. To enhance the extraction of local features, the strip pooling operation is introduced to retain the time-frequency information in the spectrogram. The update process of the feature map can be expressed as:

[0048] P updated = CNN(F mfcc ) · W stripes , (12)

[0049] where, P updatedRepresents the feature map after strip pooling, W stripes is the convolutional kernel weight.

[0050] Step 4.2: Emotional feature extraction. Process P through a residual network updated to extract more accurate emotional features. First, use the Shift Window Partitioning (SWP) method to partition the spectrogram into local regions, extract local features in different frequency bands, and further optimize their representations. The specific operations are as follows:

[0051] P local = SWP(P updated ), (13)

[0052] where P local is the local feature obtained through the SWP operation

[0053] Step 4.3: Capturing subtle emotional changes. Map the emotional features to audio Tokens [v1, v2,..., v m through a linear transformation, and add the positional encoding p i to preserve the temporal information. Then, the audio Tokens interact through the multi-head self-attention mechanism to calculate their importance in terms of emotional categories. The calculation formula for the attention matrix is:

[0054]

[0055] where Q, M, and N are the query, key, and value vectors respectively.

[0056] Step 4.4: Decision Token selection. Remove the Tokens that contribute less to emotional classification through a redundancy removal Token selection mechanism to improve the model performance. Calculate the importance weight a ij for each Token, and the emotional prediction formula based on the selected Tokens is:

[0057]

[0058] where is the emotional prediction result, FC represents the fully connected layer, and LN is the layer normalization operation.

[0059] Step 4.5: Loss function calculation and optimization. Calculate the cross-entropy loss function L ER by comparing with the true emotional label y. The calculation formula for the loss function is as follows:

[0060]

[0061] where N is the number of samples, y true,iis the true label of the i-th sample, and is the predicted sentiment category by the model. The loss value L ER will be backpropagated into the model through the backpropagation algorithm, and the model parameters θ will be updated using the gradient descent method:

[0062]

[0063] where η is the learning rate, and is the gradient of the loss function with respect to the model parameters.

[0064] By comparing the text sentiment and speech sentiment prediction results generated in Step 3 and Step 4, the annotation accuracy of the data is ensured and a high-quality sentiment annotation dataset is gradually constructed. The specific steps are as follows:

[0065] Step 5.1: Comparison and screening of text and speech sentiment prediction results. Denote the text sentiment prediction result generated in Step 3 as P T (x), and the speech sentiment prediction result generated in Step 4 as P S (x). Compare the text prediction result P T (x i ) and the speech prediction result P S (x i ). When the prediction results are the same, that is, P T (x i ) = P S (x i ), it is considered that the sentiment prediction of this sample is credible, and the sentiment is directly annotated as y i = P T (x i ). If the prediction results are inconsistent, that is, P T (x i ) ≠ P S (x i ), this sample is screened out to form a conflict dataset:

[0066] D confict = {x i | P T (xi) ≠ P S (x i}}, (18)

[0067] where P updated represents the feature map after strip pooling, and W stripes is the convolutional kernel weight.

[0068] Step 5.2: Iterative optimization of conflict data. Input the conflict dataset D confict into the text understanding module and the speech sentiment recognition module, and give the text understanding module guiding words, and respectively obtain the text sentiment prediction results and the speech emotion prediction result where k is the current iteration number. In each iteration, optimize the following consistency loss function:

[0069]

[0070] where, L T and L S represent the loss functions of the text module and the speech module respectively, and are the current predicted target values, and α and β are weight hyperparameters. Through multiple iterations of optimization, make the prediction result meet the consistency condition:

[0071]

[0072] When the consistency condition is met, add the labeled result y i = P T (x i ) to the high-quality emotion labeled dataset.

[0073] Step 5.3: Construction of the emotion labeled dataset. In each iteration, gradually add the samples that meet the consistency condition to the high-quality emotion labeled dataset. The final dataset can be expressed as:

[0074] D final = D final ∪{(x i , y i ) | P T (x i ) = P S (x i )}. (21) By continuously repeating the processes of annotation, screening, and optimization, ensure the accuracy and consistency of data annotation. The finally obtained high-quality emotion labeled dataset D final .

[0075] The present invention also provides a dialogue-based speech emotion automatic annotation system based on a graph language model, including:

[0076] A speech data acquisition module, which is used to collect the dialogue speech data between the operator and the citizen through the government service hotline system to form a speech data set D_voice;

[0077] A speech preprocessing module, which is used to perform noise reduction, speech segmentation, and spectrum processing on the speech data D_voice, and convert it into text data T to generate bimodal data D_bimodal;

[0078] A text information understanding module, which is used to receive text data T, extract text sentiment features based on a large-scale pre-trained language model (LLM), and generate a text sentiment prediction result R_text;

[0079] A speech emotion recognition module, which is used to receive a spectrogram sequence I_spec, extract speech emotion features, and generate a speech emotion prediction result R_audio;

[0080] An emotion annotation optimization module, which is used to compare the text sentiment prediction result R_text and the speech emotion prediction result R_audio. If they are consistent, direct emotion annotation is performed. Otherwise, the text information understanding module and the speech emotion recognition module are re-optimized to improve the accuracy of emotion prediction.

[0081] Furthermore, the speech preprocessing module specifically includes:

[0082] A noise suppression unit, which is used to remove static and dynamic background noises and improve the quality of the speech signal;

[0083] A speech segmentation unit, which is used to perform frame-level segmentation on the speech signal and perform time-frequency analysis based on the short-time Fourier transform (STFT);

[0084] A speech-to-text unit, which is used to convert the speech signal into text data T based on automatic speech recognition (ASR) technology;

[0085] A spectrogram conversion unit, which is used to convert the speech data into a spectrogram sequence I_spec for processing by the speech emotion recognition module.

[0086] Furthermore, the text information understanding module includes:

[0087] A text feature extraction unit, which is used to extract text sentiment features F_text using a large-scale pre-trained language model (LLM);

[0088] A multi-task learning unit, which is used to optimize and enhance the text sentiment features and improve the generalization ability of text sentiment recognition;

[0089] An emotion classification unit, which is used to generate a text sentiment prediction result R_text based on the optimized emotion features and calculate a loss value L_text to optimize the model parameters.

[0090] Furthermore, the speech emotion recognition module includes:

[0091] A feature generation unit, which is used to extract preliminary speech emotion features F_audio;

[0092] An emotion feature extraction unit, which is used to further extract emotion information such as the pitch, volume, and frequency distribution of the speech;

[0093] An emotional relationship mining unit for capturing subtle emotional changes in speech signals and optimizing the emotional feature F_audio^opt;

[0094] A classification decision-making unit for generating the final speech emotion prediction result R_audio based on the optimized emotional feature F_audio^opt and calculating the loss value L_audio to optimize the model parameters.

[0095] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0096] (1) The present invention utilizes the powerful semantic understanding ability of the large language model and deeply integrates it with the speech emotion recognition technology, breaking through the limitations of traditional manual annotation. Through the adaptive annotation technology, not only the high automation of emotion annotation is realized, but also the need for manual intervention is significantly reduced, thus saving a large amount of labor costs and improving the efficiency of large-scale data processing;

[0097] (2) Based on the dual-modal collaborative analysis framework of text and speech, the present invention can extract multi-dimensional emotional features from speech signals and semantic information, effectively reducing the subjective deviation in the annotation process. This method ensures the accuracy and consistency of the annotation results in complex scenarios, providing high-quality annotation data for the subsequent training of emotion recognition models;

[0098] (3) Aiming at the problem of the decline in the accuracy of emotion recognition caused by accent differences between the two parties in the government affairs call scenario, the present invention innovatively designs an accent adaptation module based on the language model. This module can intelligently capture different accent features and dynamically adjust the speech feature extraction strategy during the annotation process, effectively improving the recognition accuracy of speech emotions under different accents. Through this improvement, the present invention shows stronger robustness and adaptability when dealing with government service hotline call data in cross-regional and multilingual environments;

[0099] (4) The present invention can accurately capture the subtle emotional changes in speech signals during government affairs calls, especially showing excellent recognition ability in scenarios with complex contexts and frequent changes in emotional intensity. This feature makes the present invention not only applicable to conventional call scenarios, but also capable of maintaining high-efficiency emotion annotation performance in more complex multi-party conversations or dynamic situations. Description of the Drawings

[0100] Figure 1 is the flowchart of the dialogue-based speech emotion automatic annotation method based on the atlas language model in the embodiment of the present invention

[0101] Figure 2 is the schematic diagram of speech emotion annotation in the government service hotline scenario;

[0102] Figure 3 It is a schematic diagram of the model for text information understanding in the embodiments of the present invention;

[0103] Figure 4 It is a schematic diagram of the model for language emotion recognition in the embodiments of the present invention. Specific Embodiments

[0104] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0105] Embodiment 1: Dialogue Emotion Analysis Based on Government Service Hotline

[0106] In the government service hotline of a certain city, the operation center hopes to use voice emotion recognition technology to conduct emotion analysis on the calls between operators and citizens in order to optimize service quality and improve user satisfaction. The system first collects call voice data through a voice data collection module, and performs noise reduction, segmentation, and speech-to-text processing through a voice preprocessing module to generate text data T and corresponding spectrogram data I_spec. Subsequently, the text information understanding module extracts text emotion features based on the LLM model, and the voice emotion recognition module generates a voice emotion prediction result R_audio through feature extraction and emotion relationship modeling. Finally, the emotion annotation optimization module compares the text emotion prediction result R_text with the voice emotion prediction result R_audio to automatically annotate the emotion category. When the system detects that a citizen is emotionally excited (such as angry or anxious), it will send an alarm to the management personnel to facilitate manual intervention and improve the response ability and service quality of the government hotline.

[0107] Embodiment 2: Optimization of User Experience in Intelligent Customer Service System

[0108] An online banking intelligent customer service platform hopes to optimize the user experience and improve the accuracy of customer interaction by using an emotion recognition system. In this application, the system first collects the voice interaction data between users and the intelligent customer service in real time, and completes noise reduction, speech segmentation, and automatic transcription through the speech preprocessing module. Then, the text information understanding module extracts the emotional features in the user's question based on a large-scale pre-trained language model (LLM), while the speech emotion recognition module extracts the subtle emotional changes in the speech signal, such as the user's dissatisfaction or anxiety state. By comparing the text emotion prediction results and the speech emotion prediction results through the emotion annotation optimization module, when it is found that the user may be dissatisfied with the response of the customer service system, the system will actively adjust the dialogue strategy, such as switching to a human customer service, providing additional assistance, or adjusting the tone, so as to effectively improve the user experience and satisfaction.

[0109] Example 3: Emotion Analysis System for Online Education Platform

[0110] An online education platform hopes to monitor the learning status of students, improve the interactivity of online courses, and enhance the teaching quality through emotion recognition technology. In a live classroom, the system records the dialogue audio between teachers and students through the voice data collection module, and uses the speech preprocessing module to perform noise reduction, segmentation, and text transcription on the data. The text information understanding module extracts the emotional features in the students' speeches, such as confusion, excitement, or boredom, while the speech emotion recognition module analyzes the changes in tone and pitch to determine whether the students are interested in the course content. When the system detects a decline in the students' emotions or a long period of silence, the emotion annotation optimization module will notify the teacher to adjust the teaching method in a timely manner, such as increasing interactive questions or changing the teaching style, so as to enhance classroom participation and learning effects.

[0111] Example 4: Emotion Interaction Optimization for Intelligent Voice Assistants

[0112] An intelligent voice assistant (such as a smart home voice assistant or a car voice assistant) hopes to achieve a more natural human-computer interaction experience. In this application, the system first collects the user's voice commands in real time through the voice data collection module, and performs audio noise reduction, feature extraction, and speech-to-text processing through the speech preprocessing module. Subsequently, the text information understanding module analyzes the content of the user's commands, and the speech emotion recognition module combines parameters such as tone and speech rate to judge the user's emotions, such as whether the user is in a tense, angry, or happy state. When the system detects that the user's tone is anxious or angry, the emotion annotation optimization module will adjust the tone of the voice assistant so that it responds in a more peaceful or friendly manner, thereby enhancing the user experience. For example, when the user is in a hurry to travel, the car voice assistant can provide navigation information in a more concise and intuitive way, while when the user is relaxing and resting, it can use a more gentle tone to provide music recommendations and schedule reminders.

[0113] Such as Figure 1As shown in the figure, a method for automatically annotating conversational speech emotions based on a graph language model according to an embodiment of the present invention is particularly applicable to the conversation scenario between citizens and operators regarding social security issues in a government service hotline system. Its principle is as Figure 1 shown. The method includes steps 1 to 5, which are specifically as follows:

[0114] Step 1: Data acquisition: In the Figure 2 government service hotline scenario shown in the figure, comprehensively collect the dialogue voice data between operators and citizens regarding social security issues. Use a high-sensitivity microphone to capture the voice signals of operators and citizens in real time and convert them into digital signals. Ensure the authenticity and diversity of the data sources to provide a reliable data basis for subsequent emotion analysis. The specific steps are as follows:

[0115] Step 1.1): Capture the voice signal of the operator. Use a high-sensitivity microphone with a sampling rate of at least 44.1 kHz and a bit depth of 16 bits to capture the voice signal of the operator in real time and directly convert it into a digital signal. The signal collected by the microphone is recorded by the acquisition system to form clear digital voice data. For example, when the operator answers the citizen's question about "the length of social security payment", his voice signal is recorded in detail. The signal conversion process is defined as S′ agent = f mic (S agent ), where S′ agent is the processed digital voice signal of the operator, and f mic is the microphone signal processing function.

[0116] Step 1.2): Collect the voice signal of the citizen. The voice signal of the citizen is transmitted through the telephone network and automatically converted into a digital signal by the telephone system. Considering that citizens may come from different regions and have different language accents, special attention is paid to the stability of the audio quality during the collection process. The digital signal is then recorded by the acquisition system to generate the voice data of the citizen. The signal conversion process is expressed as S′ agent = f mic (S agent ), where S′ itizen is the digital voice signal of the citizen after processing, and f tel is the signal conversion function of the telephone system.

[0117] Step 1.3): Textualize the conversation content. Convert the collected voice data into corresponding text data through automatic speech recognition (ASR) technology. For the citizen's inquiry about "social security payment period" and the operator's detailed answer, generate text data containing the complete conversation content. For example, the text data may include: "Citizen: May I ask how many years of social security need to be paid? Operator: When a laborer reaches the legal retirement age, he / she needs to have accumulated 15 years of endowment insurance payment to receive a pension. If the payment period is insufficient, one can choose to make a lump-sum payment or liquidate the personal account." Form dual-modal data containing voice signal information and text content.

[0118] Step 2: Data preprocessing: Preprocess the collected conversation voice data, including noise reduction, voice segmentation, spectrum processing, and cleaning and formatting of text data. The specific steps are as follows:

[0119] Step 2.1): Noise removal. Use an adaptive noise suppression algorithm to remove background noise, eliminating both static and dynamic noise parts to ensure the clarity of the voice signal. For example, for the noisy background in a government service hotline environment, use a Wiener filter for noise removal. This denoising process is defined as S clean (t) = f wienerFilter (S noisy (t)), where S clean (t) represents the clear voice signal after denoising, f wienerFilter is the used noise suppression algorithm, and S noisy (t) is the original noisy voice signal.

[0120] Step 2.2): Voice signal segmentation. Perform frame-level segmentation on the clear voice signal. Each frame of the signal is windowed through a window function to form short-time signal segments. Then, perform a short-time Fourier transform on each frame of the signal to generate corresponding time-frequency segments. The specific time-frequency analysis formula is expressed as where S stft (f,t) represents the voice signal after time-frequency analysis, f stft is the short-time Fourier transform operation.

[0121] Step 2.3): Spectrogram conversion. Extract Mel-frequency cepstral coefficients from the processed voice signal and use a Mel filter bank to perform feature conversion on the signal to obtain a spectrogram sequence. The specific spectrogram conversion process is expressed as F mfcc = f MFCC (W), where F mfcc is the extracted spectrogram feature sequence, and f MFCC is the Mel-frequency cepstral coefficient extraction function.

[0122] Step 2.4): Text data cleaning and formatting. Clean the text data after ASR conversion, remove noise such as irrelevant words, punctuation marks, etc., and format the data into a standard format suitable for sentiment analysis. For example, convert the dialogue text into JSON format:

[0123]

[0124] Step 3: Text sentiment analysis: Input the preprocessed text data into a large-scale pre-trained language model (LLM) to perform sentiment analysis on the text information. Especially for the text recognition errors that may be caused by different language accents of citizens, utilize the powerful language understanding ability of the LLM to accurately extract sentiment features. The specific steps are as follows:

[0125] Step 3.1): Text data preprocessing and prompt design. First, perform task-specific prompt design on the preprocessed text data to guide the LLM to accurately extract the sentiment features in the text. Considering the possible language accent problems of citizens, pay attention to guiding the model to focus on the core sentiment expression of the text rather than specific accent features when designing the prompt. For example, the prompt can be designed as: "Please analyze the sentiment tendency of citizens in the dialogue text, ignoring possible accent differences." And input it into the BERT feature extractor. The prompt design formula is expressed as T input = T 提示词 + T, where T input is the complete prompt and text combination input into BERT, and T is the converted text data.

[0126] Step 3.2): Sentiment feature extraction. Input the designed prompt and text data into the BERT model, and extract the sentiment features in the text through the pre-trained model and self-attention mechanism. The BERT model establishes the connection between words through context information and uses multiple layers of Transformer encoders for feature extraction. For different accents of citizens, the BERT model can utilize its powerful generalization ability to accurately capture the core sentiment information in the text. The process of the feature extraction process is expressed as H = BERT(T input ), where H is the hidden layer feature extracted by the BERT model, and T input is the text input containing the prompt.

[0127] Step 3.3): Multi-task learning optimization. Input the extracted sentiment features into the multi-task learning module for optimization and enhancement. Considering the sentiment recognition errors that may be caused by citizens' accents, the multi-task learning module introduces tasks such as sentiment intensity prediction task and sentiment turn recognition task to enhance the model's learning ability for sentiment features. The multi-task learning process is defined as L T = λ1L1 + λ2L2 + λ3L3, where L TThe sum of losses for each task, L i The loss value for each task, λ i Is the weight corresponding to the task.

[0128] Step 3.4): Sentiment classification and prediction. Input the optimized features output by the multi-task learning module into the sentiment classification layer to generate the sentiment prediction result of the text. The sentiment classification layer uses the Softmax function to predict the sentiment category and outputs the probability value of each sentiment category. For the different accents of citizens, the sentiment classification layer can accurately judge the sentiment tendency in the text. The sentiment classification process is expressed as Where, Represents the sentiment category prediction result, W h Is the optimized feature matrix.

[0129] Step 3.5): Backpropagation and model optimization. Backpropagate the loss value into the model through the backpropagation algorithm and update the model parameters using the gradient descent method. Considering the diversity of citizens' accents, the model continuously optimizes its generalization ability during training to accurately identify the emotional expressions under different accents. The specific update process is expressed as Where, θ is the model parameter, η is the learning rate, Is the gradient of the loss function with respect to the model parameter.

[0130] Step 4: Speech emotion recognition: Input the transformed spectrogram sequence into the speech emotion recognition model to perform emotional analysis on the speech information. The model first extracts the emotional features in the speech signal through the backbone network, and these features will then be input into the emotional relationship mining module to capture more subtle emotional changes. Specifically for the different accents of citizens, the model adopts an adaptive feature extraction strategy to ensure the accuracy of emotion recognition. The specific steps are as follows:

[0131] Step 4.1): Discriminative feature generation. Process the input spectrogram sequence through the discriminative feature generation module to generate the preliminary features of the speech signal. This module uses a convolutional neural network to extract time-frequency features to obtain the feature map, and introduces strip pooling operation to retain the time-frequency information in the spectrogram. The update process of the feature map is defined as P updated = CNN(F mfcc )·W stripes , where, P updated Represents the feature map after strip pooling, W stripes Is the convolutional kernel weight.

[0132] Step 4.2): Emotional feature extraction. Through the residual network for P updatedProcess it to extract more accurate emotional features. Use the Shift Window Partitioning (SWP) method to divide the spectrogram into local regions, extract local features in different frequency bands, and further optimize their representation. The specific operation is denoted as P local = SWP(P updated ), where P local is the local feature obtained through the SWP operation.

[0133] Step 4.3): Subtle emotional change capture. Map the emotional features to audio tokens through linear transformation and add positional encoding to preserve the temporal information. Then, the audio tokens interact through the multi-head self-attention mechanism to calculate their importance in terms of emotional categories. Considering the different accents of the citizens, the model pays more attention to the overall trend and rhythm of the audio signal when capturing subtle emotional changes, rather than specific accent features. The calculation process of the attention matrix is denoted as

[0134] Step 4.4): Decision token selection. Remove the tokens that contribute less to the emotional classification through the redundant token selection mechanism to improve the model performance. Calculate the importance weights of each token and select the most representative decision token for emotional prediction according to the weights. Considering the diversity of the citizens' accents, the model pays more attention to the emotional expressiveness of the audio signal when selecting decision tokens, rather than specific accent features.

[0135] Step 4.5): Loss function calculation and optimization. Calculate the cross-entropy loss function L S . The loss function is denoted as where N is the number of samples, y true,i is the true label of the i-th sample, is the emotional category predicted by the model. The loss value L S will be backpropagated into the model through the backpropagation algorithm, and the process of updating the model parameters θ using the gradient descent method is denoted as where η is the learning rate, is the gradient of the loss function with respect to the model parameters.

[0136] Step 5: Emotional annotation and dataset construction: Compare the text emotion and speech emotion prediction results generated in Step 3 and Step 4 to ensure the annotation accuracy of the data and gradually construct a high-quality emotional annotation dataset. The specific steps are as follows:

[0137] Step 5.1): Comparison and screening of text and speech emotion prediction results. Denote the text emotion prediction result generated in Step 3 as P T (x), and denote the speech emotion prediction result generated in Step 4 as P S(x). For the text prediction result P T (x i ) and the speech prediction result P S (x i ) are compared. When the prediction results are consistent, i.e., P T (x i ) = P S (x i ), the emotional prediction of this sample is considered credible, and the emotion is directly labeled as y i = P T (x i ). If the prediction results are inconsistent, i.e., P T (x i ) ≠ P S (x i ), this sample is screened out to form a conflict data set:

[0138] Step 5.2): Iterative optimization of conflict data. The conflict data set D confict is input into the text understanding module and the speech emotion recognition module, and guiding words are given to the text understanding module to obtain the text emotion prediction result and the speech emotion prediction result where k is the current iteration number. In each iteration, the optimization consistency loss function is expressed as where, L T and L S respectively represent the loss functions of the text module and the speech module, and are the current prediction target values, and α and β are weight hyperparameters. Through multiple iterations of optimization, the prediction results are made to satisfy the consistency condition When the consistency condition is satisfied, the annotation result y i = P T (x i ) is added to the high-quality emotion annotation data set.

[0139] Step 5.3): Construction of the emotion annotation data set. In each iteration, the samples that meet the consistency condition are gradually added to the high-quality emotion annotation data set. The final data set is expressed as D final = D final ∪ {(x i , y i ) | P T (x i ) = P S (x i )}, and by continuously repeating the processes of annotation, screening, and optimization, the accuracy and consistency of data annotation are ensured.

[0140] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A dialogue-based speech emotion automatic annotation system based on a graph language model, characterized in that Including: A voice data acquisition module, which is used to collect the dialogue voice data between the operator and the citizen through the government service hotline system to form a voice data set D_voice; A voice preprocessing module, which is used to denoise, segment the voice, and perform spectrum processing on the voice data D_voice, and convert it into text data T to generate bimodal data D_bimodal; A text information understanding module, which is used to receive the text data T, extract text emotion features based on a large-scale pre-trained language model (LLM), and generate a text emotion prediction result R_text; A voice emotion recognition module, which is used to receive the spectrogram sequence I_spec, extract voice emotion features, and generate a voice emotion prediction result R_audio; An emotion annotation optimization module, which is used to compare the text emotion prediction result R_text and the voice emotion prediction result R_audio. If the two are consistent, emotion annotation is directly performed; otherwise, the text information understanding module and the voice emotion recognition module are re-optimized to improve the accuracy of emotion prediction.

2. The system according to claim 1, wherein The voice preprocessing module specifically includes: A noise suppression unit, which is used to remove static and dynamic background noise to improve the quality of the voice signal; A voice segmentation unit, which is used to perform frame-level segmentation on the voice signal and perform time-frequency analysis based on the short-time Fourier transform (STFT); A voice-to-text unit, which is used to convert the voice signal into text data T based on the automatic speech recognition (ASR) technology; A spectrogram conversion unit, which is used to convert the voice data into a spectrogram sequence I_spec for processing by the voice emotion recognition module.

3. The system according to claim 1, wherein The text information understanding module includes: A text feature extraction unit, which is used to extract text emotion features F_text using a large-scale pre-trained language model (LLM); A multi-task learning unit, which is used to optimize and enhance the text emotion features to improve the generalization ability of text emotion recognition; An emotion classification unit, which is used to generate a text emotion prediction result R_text based on the optimized emotion features and calculate a loss value L_text to optimize the model parameters.

4. The system according to claim 1, wherein The voice emotion recognition module includes: A feature generation unit, which is used to extract preliminary voice emotion features F_audio; An emotion feature extraction unit, which is used to further extract emotion information such as the pitch, volume, and frequency distribution of the voice; An emotion relationship mining unit, which is used to capture subtle emotion changes in the voice signal and optimize the emotion features F_audio^opt; A classification decision unit, which is used to generate a final voice emotion prediction result R_audio based on the optimized emotion features F_audio^opt and calculate a loss value L_audio to optimize the model parameters.

5. A method for automatically annotating conversational speech emotions based on a graph language model, characterized in that, Including steps: Step 1: Through the government service hotline system, comprehensively collect the dialogue voice data between operators and citizens to ensure the authenticity and diversity of data sources, and write it into the voice data set D voice , providing a reliable data basis for subsequent sentiment analysis; Step 2: Preprocess the collected dialogue voice data D voice including operations such as noise reduction, speech segmentation, and spectrum processing. At the same time, convert the voice data into corresponding text data T to form bimodal data D containing voice signal information and text content bimodal ; Step 3: Input the preprocessed text data T into a large-scale pre-trained language model (LLM) to perform sentiment analysis on the text information. The model first extracts the sentiment features F from the text based on the prompt words text , then optimizes and enhances these features through a multi-task learning module, and finally generates the text sentiment prediction result R through the sentiment classification layer text , and calculates the loss value L with the true value Y text ; text ; Step 4: Input the transformed spectrogram sequence I spec into the speech emotion recognition model for emotion analysis of the speech information. The model first extracts the emotional features in the speech signal through the backbone network. These features will then be input into the emotional relationship mining module to capture more subtle emotional changes, and finally generate the speech emotion prediction result R audio , and calculate the loss value L audio with the true value Y audio ; Step 5: Compare the text sentiment prediction result R text generated in Steps 3 and 4 and the speech sentiment prediction result R audio If the prediction results of the two are the same, the sentiment prediction result is considered credible and directly subjected to sentiment annotation; otherwise, the data with different predictions is re-input into the text information understanding module and the speech sentiment recognition module to iteratively optimize the model prediction performance until the required confidence level is reached. By repeatedly performing the annotation, evaluation, and optimization processes, a high-quality sentiment annotation dataset is finally constructed.

6. The method for automatically annotating the emotion of dialog-based speech according to claim 5, characterized in that, Step 1, through the government service hotline platform, citizens call into the service system by phone, and the system automatically records all call contents. The audio data of each call session is recorded and saved in real time to form a data set D voice ={S1, S2, …, S n}, ensuring that the data covers conversation samples with different time periods and different voice emotions.

7. The method for automatically annotating dialogic speech emotions according to claim 5, wherein Step 2, after the collection of voice data is completed, the collected voice data D voice is subjected to multiple preprocessings to ensure the quality and standardization of the data, and provide accurate input data for the subsequent voice emotion recognition model and text information understanding model. The specific steps are as follows: Step 2.1: Use a noise suppression algorithm to effectively suppress the background noise in the call, remove the static noise and dynamic noise parts to improve the clarity of the voice signal. For complex background noise, an adaptive filter is used to further optimize the voice data to enhance the quality and recognition accuracy of the voice; Step 2.2: Segment the processed speech data and perform frame-level segmentation of the speech signal using a fine-grained alignment method. Perform time-frequency analysis using the short-time Fourier transform to divide the speech signal into multiple time-frequency segments {P1, P2, …, P n}, providing precise speech segments for subsequent emotional feature extraction; Step 2.3: Convert the speech signal into text data through automatic speech recognition technology. Using speech recognition technology, convert the speech content in the speech into text data T = {t1, t2, …, t n}, forming bimodal data D bimodal = {I spec , T}, providing a basis for subsequent sentiment analysis and understanding; Step 2.4: Process each speech segment through Mel-frequency cepstral coefficients and a Mel filter bank to convert it into a spectrogram sequence I spec ={P1, P2, …, P n}, which serves as the standardized input for the emotion recognition model.

8. The method for automatically annotating dialogic speech emotions according to claim 5, characterized in that, Step 3: Perform sentiment analysis on the transformed text data through the text information understanding module. The model uses a large-scale pre-trained language model (LLM) to extract text sentiment features, optimizes and enhances the features through a multi-task learning module, and finally generates a sentiment prediction result and calculates the loss value. The specific steps are as follows: Step 3.1: Design prompt words according to the task requirements for the preprocessed text data T = {t1, t2, …, t n}, and input them into the BERT feature extractor; Step 3.2: The BERT feature extractor uses the context information and extracts the sentiment feature F in the text through the self-attention mechanism text ; Step 3.3: Input the extracted sentiment features into the multi-task learning module to optimize and enhance the features and improve the generalization ability of the model; Step 3.4: Input the optimized sentiment features into the sentiment classification layer to generate a text sentiment prediction result; Step 3.5: Compare the prediction result R of the text information understanding module text with the true value Y text , calculate the loss value L text , and use it for optimizing the model parameters.

9. The method for automatically annotating the emotion of dialog voice according to claim 5, characterized in that Step 4, perform sentiment analysis on spectrogram sequence I spec ={P1, P2, …, P n} by the voice emotion recognition model. First, the model generates preliminary features F audio ={f1′, f2′, …, f n ′} of the voice signal through the discriminative feature generation module. Then, extract the emotion features through the backbone network, and then use the emotion relationship mining module to capture the subtle emotion changes in the voice signal. Finally, the model generates the emotion prediction result through the decision Token selection and calculates the loss value for optimization. The specific steps are as follows: Step 4.1: Process the input spectrogram sequence I through the discriminative feature generation module spec to extract a preliminary emotional feature sequence F audio as the basis for subsequent feature extraction and sentiment analysis; Step 4.2: Use the backbone network to further process the sentiment features output by the discriminative feature generation module, including important information such as pitch, volume, and frequency distribution; Step 4.3: Input the extracted sentiment features into the sentiment relationship mining module. The sentiment relationship mining module models the subtle sentiment changes based on the context relationship and generates features Step 4.4: Input the optimized sentiment features into a multi-layer perceptron to generate the final speech sentiment prediction result R audio , and output the sentiment classification of the model for the current speech signal Step 4.5: Compare the prediction result R of the language emotion recognition module audio with the true value Y audio to calculate the loss value L audio , and optimize the model parameters through backpropagation.

10. The method for automatically annotating the emotion of a dialogue-based voice according to claim 5, characterized in that Step 5, by comparing the text sentiment prediction results R text generated in Steps 3 and 4 with the speech sentiment prediction results R audio to ensure the annotation accuracy of the data and gradually construct a high-quality sentiment annotation dataset. The specific steps are as follows: Step 5.1: Compare the generated text and speech sentiment prediction results. If the two prediction results are the same, the prediction result is considered credible and sentiment annotation is performed; If the predicted values are different, then filter out this data; Step 5.2: Re-input the filtered data into the text understanding module and the speech sentiment recognition module for iterative optimization until the two prediction results are the same; Step 5.3: By continuously repeating the annotation, evaluation, and optimization processes, finally construct a high-quality sentiment annotation dataset for subsequent model training.

Citation Information

Cited By

  • Speech emotion recognition method and system based on multi-modal feature fusion

    CN121565211A