Voice emotion recognition method and device based on context information, equipment and medium

By combining intelligent segmentation, speaker separation, speech recognition and semantic analysis with multimodal fusion, the problem of the inability of existing technologies to accurately identify emotional changes in complex conversations has been solved, and the accuracy and stability of emotion recognition in the fields of financial technology and healthcare have been improved.

CN120636474APending Publication Date: 2025-09-12PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510793541.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing speech emotion recognition technology cannot accurately identify the emotional changes of various characters in complex dialogue scenarios, and ignores contextual information, resulting in insufficient accuracy and stability of emotion recognition, which particularly affects decision-making and service quality in the fields of financial technology and healthcare.

Method used

By receiving the original voice stream and performing intelligent sentence segmentation and speaker separation, independent voice segments are generated. Automatic speech recognition and semantic analysis are combined to determine the role type, extract acoustic feature indicators, and call historical dialogue text to generate context information, and finally generate emotional judgment results in the multimodal fusion module.

Benefits of technology

Accurately identifying and understanding the emotional changes of each character in complex dialogue scenarios improves the accuracy and stability of emotion recognition, avoids errors in single-sentence emotion judgment, and improves the fluency of dialogue and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636474A_ABST
    Figure CN120636474A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a context information-based voice emotion recognition method, device, equipment and medium, which comprises the following steps: receiving an original voice stream and generating an independent voice segment, recognizing a text and determining a speaker role type, and extracting an acoustic feature index; and generating a preliminary emotion label, generating context information in combination with the historical dialogue text, and inputting the context information, the preliminary emotion label, the speaker role type and the acoustic feature index into a multi-modal fusion module to generate an emotion judgment result. According to the method, multi-modal fusion is realized on the basis of context information by combining voice, text and role information, so that the emotion change of each role can be accurately recognized and understood in a complex dialogue scene, the problems of large single sentence emotion judgment error and neglect of the context information in a traditional method are avoided, and the accuracy and stability of emotion recognition are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a speech emotion recognition method, apparatus, device and storage medium based on context information. Background Art

[0002] In customer service and robotics applications, speech emotion recognition technology has become a key tool for improving user experience. However, existing technologies still have significant shortcomings, particularly in the accuracy of emotion recognition and the comprehensive understanding of contextual information.

[0003] In the fintech business sector, interactions between customers and customer service representatives or robots often involve high-value transactions and sensitive information, and even slight changes in customer sentiment can directly impact the decision-making process. Therefore, accurately judging customer emotions and tone is crucial for preventing potential dissatisfaction, controlling risks, and improving customer satisfaction. While existing sentiment analysis technologies can identify sentiment within individual sentences, they often overlook the contextual information underlying changes in customer emotions. For example, when a customer is speaking with a financial services robot and expresses hesitation or impatience when answering certain questions, existing technologies struggle to effectively capture these details, leading to misjudgments or delayed responses, which in turn impact service quality and customer experience. Furthermore, traditional sentiment analysis technologies lack consideration of the entire conversation history, making it impossible to comprehensively assess changing trends in customer sentiment. Furthermore, they fail to process emotional feedback from customers who are reluctant to answer questions, thus missing opportunities to drive business success through emotional adjustment.

[0004] In the healthcare business, speech emotion recognition technology also plays a vital role in conversations between patients and medical staff or smart health assistants. In many healthcare scenarios, the patient's emotional state (such as anxiety, depression or fear) may be closely related to their condition. Therefore, accurately identifying the patient's emotional state and providing psychological comfort or guidance in a timely manner are crucial to improving treatment outcomes and patient satisfaction. Existing technologies can usually only analyze the emotional tone of the patient's voice, but are unable to combine emotional changes with multi-dimensional contextual information such as the patient's medical history and treatment progress for comprehensive judgment, resulting in the inability to identify the patient's potential emotional fluctuations or ignoring emotional feedback related to the condition. In particular, when the patient is slow to respond or expresses dissatisfaction in the conversation, traditional sentiment analysis systems often fail to capture it in time, thus missing the opportunity for timely intervention.

[0005] Despite progress in speech emotion recognition technology, the accuracy of single-sentence sentiment analysis remains limited by the limitations of speech and semantic context. For example, existing systems struggle to accurately identify emotional shifts in clients or patients when the change is subtle or masked by preceding context. Furthermore, existing technologies often rely on static models, unable to dynamically adjust the weight of emotions within a conversation and lacking the ability to adapt to changing context. As conversations become more complex, traditional emotion recognition technology fails to fully leverage the contextual information from multiple rounds of dialogue, limiting both the accuracy of emotion recognition and the fluency of conversations. Summary of the Invention

[0006] The main purpose of the present invention is to provide a speech emotion recognition method, device, equipment and storage medium based on contextual information, aiming to solve the technical problem that the existing technology is unable to accurately identify and understand the emotional changes of each character in the conversation based on multimodal context, especially the difficulty in reliably generating accurate emotion recognition results in complex conversation scenarios.

[0007] To achieve the above object, the present invention provides a method for speech emotion recognition based on context information, comprising:

[0008] Receive the original speech stream and perform intelligent segmentation processing to generate speech segments, and perform speaker separation operation on the speech segments to generate independent speech segments;

[0009] Inputting the independent speech segment into an automatic speech recognition model to generate a current speech text, and performing role differentiation on the current speech text through a semantic analysis model to determine the speaker role type;

[0010] Extracting acoustic feature indicators from the independent speech segment;

[0011] Inputting the independent speech segment into a speech emotion recognition model and outputting a preliminary emotion label;

[0012] Retrieving historical conversation text and combining it with the current voice text to generate context information;

[0013] The context information, the preliminary emotion label, the speaker role type and the acoustic feature index are input into a multimodal fusion module to generate an emotion judgment result.

[0014] Furthermore, to achieve the above-mentioned object, the present invention provides a speech emotion recognition device based on context information, comprising:

[0015] A speech preprocessing module is used to receive the original speech stream and perform intelligent sentence segmentation processing to generate speech segments, and perform speaker separation operations on the speech segments to generate independent speech segments;

[0016] A speech recognition module is used to input the independent speech segment into an automatic speech recognition model to generate a current speech text, and perform role differentiation on the current speech text through a semantic analysis model to determine the speaker role type;

[0017] An acoustic feature extraction module, configured to extract acoustic feature indicators from the independent speech segments;

[0018] An emotion recognition module, configured to input the independent speech segment into a speech emotion recognition model and output a preliminary emotion label;

[0019] A context processing module is used to retrieve historical conversation texts and combine them with the current voice text to generate context information;

[0020] The multimodal fusion module is used to input the context information, the preliminary emotion label, the speaker role type and the acoustic feature index into the multimodal fusion module to generate an emotion judgment result.

[0021] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a determination machine device, which includes a memory, a processor, and a speech emotion recognition program based on context information stored in the memory and run on the processor. When the speech emotion recognition program based on context information is executed by the processor, the steps of the speech emotion recognition method based on context information as described above are implemented.

[0022] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a machine-readable storage medium, on which a speech emotion recognition program based on context information is stored. When the speech emotion recognition program based on context information is executed by a processor, the steps of the speech emotion recognition method based on context information as described above are implemented.

[0023] Beneficial effects: The present invention relates to the field of speech processing technology and can be applied to business scenarios such as financial technology and medical health. A speech emotion recognition method, device, equipment and medium based on contextual information are disclosed, including: receiving the original speech stream and performing intelligent segmentation processing to generate speech segments, and performing speaker separation operations to generate independent speech segments; inputting the independent speech segments into the automatic speech recognition model to generate the current speech text, performing role differentiation on the current speech text through the semantic analysis model to determine the speaker role type; extracting acoustic feature indicators from the independent speech segments; inputting the independent speech segments into the speech emotion recognition model to output preliminary emotion labels; retrieving historical dialogue texts and combining them with the current speech text to generate contextual information; inputting the contextual information, preliminary emotion labels, speaker role types and acoustic feature indicators into the multimodal fusion module to generate emotion judgment results. By combining speech, text and role information and realizing multimodal fusion based on contextual information, the present invention can accurately identify and understand the emotional changes of each role in complex dialogue scenarios, avoiding the problems of large errors in single-sentence emotion judgment and neglect of contextual information in traditional methods, and effectively improving the accuracy and stability of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0025] Figure 1 A schematic diagram of an application environment of a method for speech emotion recognition based on contextual information in an embodiment of the present invention;

[0026] Figure 2 This is a flow chart of an embodiment of a method for speech emotion recognition based on context information according to the present invention;

[0027] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the speech emotion recognition device based on context information of the present invention;

[0028] Figure 4 A schematic structural diagram of a determination device according to an embodiment of the present invention;

[0029] Figure 5 FIG. 2 is another structural diagram of a determination device in an embodiment of the present invention. DETAILED DESCRIPTION

[0030] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0031] The speech emotion recognition method based on context information provided by the embodiment of the present invention can be applied in the following fields: Figure 1In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can receive the original voice stream through the user terminal and perform intelligent sentence segmentation processing to generate voice segmentation, and perform speaker separation operation to generate independent voice segments; input the independent voice segments into the automatic speech recognition model to generate the current voice text, and perform role differentiation on the current voice text through the semantic analysis model to determine the speaker role type; extract acoustic feature indicators from the independent voice segments; input the independent voice segments into the voice emotion recognition model to output preliminary emotion labels; retrieve the historical dialogue text and combine it with the current voice text to generate context information; input the context information, preliminary emotion labels, speaker role types and acoustic feature indicators into the multimodal fusion module to generate emotion judgment results. The present invention combines voice, text and role information, and realizes multimodal fusion based on context information. It can accurately identify and understand the emotional changes of each role in complex dialogue scenarios, avoids the problems of large errors in single-sentence emotion judgment and neglect of context information in traditional methods, and effectively improves the accuracy and stability of emotion recognition. Among them, the user terminal can be but is not limited to various personal computers, laptops, smart phones, tablets and portable wearable devices. The server side can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0032] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of a method for speech emotion recognition based on contextual information provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0033] like Figure 2 As shown, the speech emotion recognition method based on context information proposed by the present invention includes the following steps:

[0034] S10, receiving the original speech stream and performing intelligent segmentation processing to generate speech segments, and performing speaker separation operation on the speech segments to generate independent speech segments;

[0035] In this embodiment, the goal of intelligent segmentation processing is to segment the continuous speech stream into semantically independent segments for subsequent speech recognition and sentiment analysis. In the process of receiving the original speech stream, the system can receive various audio inputs, including real-time recordings, telephone call audio, voice messages, voice interaction recordings, etc. The audio format can be PCM, WAV, MP3, etc., and the sampling rate can be set according to the application scenario, such as 8kHz, 16kHz, 44.1kHz. The received audio signal enters the preprocessing stage, first performing data caching and streaming buffering to ensure the continuity of the audio data stream.

[0036] Intelligent sentence segmentation first applies speech endpoint detection technology, segmenting sentences by detecting energy discontinuities or silences in the audio signal. For example, in energy-based endpoint detection, the system divides the audio data into frames and calculates the short-term energy of each frame. If the energy exceeds a preset threshold, it is considered a speech signal; if it is below the threshold, it is considered a silence region. This method can identify the start and end points of speech and generate initial speech segments. For different audio environments, such as quiet customer service environments or noisy medical emergency environments, the system can dynamically adjust the energy threshold to adapt to environmental changes. Speech endpoint detection can also be based on spectral entropy, determining whether a speech frame is a speech frame by calculating the spectral entropy of the speech frame. For noisy environments, a dual threshold method can be combined: a higher threshold for identifying the start of speech and a lower threshold for identifying the end of speech, to avoid false positives caused by speech jitter.

[0037] The initial speech segments generated by intelligent segmentation processing are fed into the speaker separation module. Speaker separation selects different processing strategies based on the recording type of the input voice stream. When it is detected that the input is a single-track recording (such as a phone recording or a voice assistant conversation), the system uses a spectral clustering-based method to cluster speaker features of the initial speech segments. Specifically, the system extracts acoustic features such as MFCC (Mel-frequency cepstral coefficients) and acoustic spectrograms from each speech segment, calculates the similarity matrix between each segment, and distinguishes the speech of different speakers through a spectral clustering algorithm.

[0038] If the audio stream is a two-track recording (such as a two-channel recording of a customer service representative and a client), the system directly separates the speaker's voice data based on the left and right channels. The left and right channels of a two-track recording correspond to different speakers, so the system simply extracts the audio data from each channel and generates separate segments for each speaker. To ensure the clarity and reliability of the separated voice segments, the system can further perform noise suppression filtering after channel separation, using spectral subtraction, Wiener filtering, or adaptive noise suppression techniques to reduce background noise.

[0039] After speaker separation, the speech segments are further verified for identity consistency. Using a pre-trained speaker verification model (such as x-vector or ECAPA-TDNN), the system extracts speaker embedding features from each speech segment and determines whether the segments belong to the same speaker based on cosine similarity or Euclidean distance. For speech segments from the same speaker, the system automatically merges them into independent segments.

[0040] In the fintech business sector, intelligent segmentation processing can be applied to telephone customer service scenarios, where call recordings between customers and customer service representatives are segmented into independent voice segments in real time. For single-track recorded phone calls, the system uses spectral clustering to distinguish between customer and customer service voices; for dual-track recorded two-channel conversations, the left and right channels are used to directly separate the customer and customer service voices. Identity consistency verification based on the speaker verification model ensures that consecutive responses from the same customer are correctly aggregated into a single independent voice segment. The voice endpoint detection threshold can be automatically adjusted based on the ambient noise in the call center. For example, in a quiet environment, the threshold can be lowered to identify soft speech, and in a noisy environment, the threshold can be raised to avoid false triggering due to noise.

[0041] In the healthcare business, speech emotion recognition can be applied to remote diagnosis and treatment systems. The voice conversation between the doctor and the patient is intelligently segmented into independent voice segments, and the doctor's diagnostic recommendations and the patient's symptom descriptions are identified and separated separately. For single-track recordings, the system automatically identifies the doctor's and patient's voices through spectral clustering; in dual-track recordings, the left and right channels store the doctor's and patient's voices respectively, and the system directly separates the channels to generate independent voice segments. The system can be combined with a speaker verification model to avoid misidentifying the voice of the patient's family or accompanying personnel as the patient's voice. Noise suppression filtering ensures that voice clarity is maintained even in high-noise environments in emergency first aid scenarios.

[0042] This embodiment uses intelligent segmentation processing and speaker separation to automatically and accurately segment and identify speech segments of different speakers from the original voice stream without manual intervention. For single-track recording scenarios, spectral clustering is used to automatically identify and distinguish speakers, avoiding the tedious operation of manually labeling speakers in traditional methods. For dual-track recordings, the speaker voice data is directly separated based on the left and right channels to ensure a clear and reliable separation effect. The speaker verification model further improves the identity consistency of the voice segments, avoiding the missegmentation of the same speaker's voice segments or the confusion of multiple speakers' segments. This design not only achieves efficient and accurate speech separation in the fields of financial customer service and medical diagnosis and treatment, but also lays a solid data foundation for subsequent emotion recognition and multimodal fusion.

[0043] S20, inputting the independent speech segment into an automatic speech recognition model to generate a current speech text, and performing role differentiation on the current speech text through a semantic analysis model to determine the speaker role type;

[0044] In this embodiment, the process of inputting independent speech segments into the automatic speech recognition model is intended to convert audio signals into text data. Independent speech segments are generated through intelligent segmentation and speaker separation, and each segment contains only the speech data of a single speaker. The speech recognition model uses a pre-trained end-to-end speech recognition model, which can be a model based on a deep neural network (DNN), convolutional neural network (CNN), recurrent neural network (RNN) or Transformer architecture (such as DeepSpeech, Jasper, Wav2Vec 2.0, Conformer). These models can be pre-trained using large-scale speech datasets and have high-precision speech-to-text capabilities.

[0045] The input independent speech segment first undergoes audio preprocessing. The system normalizes the audio signal to ensure that the volume and sampling rate of the input audio meet the model's expectations, such as a 16kHz sampling rate and mono format. The audio signal then passes through a feature extraction module to generate acoustic features such as MFCC (Mel-Frequency Cepstral Coefficients) and audio spectrograms, which serve as input to the model. The model performs frame-by-frame decoding, gradually mapping the audio features of each frame into text characters, generating the current speech text with a timestamp. The introduction of timestamps ensures that each text character is accurately aligned with the corresponding audio segment.

[0046] The generated current voice text is input into the semantic analysis model to determine the speaker role type. The semantic analysis model can be a multi-task learning model based on natural language processing (NLP), using Transformer, BERT, RoBERTa or GPT models. First, the system performs word segmentation and word vectorization on the current voice text to convert the text content into a vector representation. Subsequently, the model extracts keywords from the text, especially professional terms, question sentences and answer patterns related to specific roles (such as customer service or customers). For example, customer service may frequently use service terms such as "Excuse me" and "For you", while customers may use language such as "I want to know" and "Can you help me" to express questions or requests.

[0047] The semantic analysis model calculates the probability of the customer service role and the customer role by counting the distribution density of these role keywords. The system quantifies the role probability of each individual speech segment and compares it to a preset role probability threshold. If the customer service role probability exceeds the threshold, the segment is marked as a customer service role; otherwise, it is marked as a customer role. To further ensure the accuracy of role labeling, the system verifies the sequence of conversational turns. For example, if the same character should logically speak continuously in consecutive conversational turns, but the labeling results show a role switch, the system will automatically correct this conflicting labeling.

[0048] In the fintech business field, the system can be used in voice customer service scenarios in call centers. The call recordings between customers and customer service are divided into independent voice segments. The system automatically inputs each segment into the voice recognition model to generate the current voice text with a timestamp. The semantic analysis model automatically distinguishes between customer service and customer roles by combining text content and keywords (such as "financial products", "loan interest rates", and "repayments"). If it is detected that the customer service role uses inappropriate language or loses control of emotions (such as a tough tone), the system can automatically generate an early warning or mark the call as a high-risk conversation.

[0049] In the healthcare business, speech emotion recognition can be applied to remote diagnosis and treatment or health consultation services. Voice calls between doctors and patients are automatically recognized as text, and the roles of doctors and patients are distinguished through semantic analysis models. Doctor roles usually include diagnostic terms (such as "recommended examination" and "taking medication"), while patient roles usually include symptom descriptions (such as "I feel dizzy" and "I keep coughing"). The system can automatically generate structured electronic medical records or medical records based on the role differentiation results.

[0050] In this embodiment, the system realizes efficient and accurate speech-to-text conversion and role recognition by inputting independent voice segments into the automatic speech recognition model and combining it with the semantic analysis model to distinguish roles. A pre-trained end-to-end speech recognition model is used to ensure the accuracy of text generation, and a set of role keywords is extracted from the text through the semantic analysis model to automatically distinguish roles such as customer service and customers, doctors and patients. Compared with the traditional speaker recognition method based on voice features, the semantic analysis model can flexibly judge roles based on text content, effectively avoiding role recognition errors caused by similar voiceprints. In addition, the accuracy of role recognition is further enhanced by verifying the order of dialogue turns, avoiding logical conflicts caused by labeling errors during role switching. This design not only improves the accuracy of speech emotion recognition, but also lays the foundation for subsequent multimodal sentiment analysis.

[0051] S30, extracting acoustic feature indicators from the independent speech segment;

[0052] In this embodiment, the process of extracting acoustic feature indicators from independent speech segments aims to convert audio signals into quantifiable acoustic features, facilitating subsequent sentiment analysis and multimodal fusion processing. Independent speech segments are audio segments generated through intelligent segmentation and speaker separation, each containing speech data from a single speaker. Acoustic feature indicators reflect the multidimensional information in speech signals, including pitch, intensity, speech rate, and their dynamic changes.

[0053] First, the system performs audio preprocessing on individual speech segments, converting the audio signal into a digital format, typically using a 16kHz sampling rate and a monophonic format. The system then extracts the absolute values ​​of pitch, intensity, and speech rate from the audio signal. These are known as raw acoustic feature metrics.

[0054] The absolute value of pitch is frequency information extracted from an audio signal using a short-time Fourier transform (STFT) or autocorrelation algorithm. It represents the highness of a sound, such as high or low pitch. The absolute value of intensity represents the amplitude of the speech signal, or its volume or loudness. It can be obtained by calculating the energy or logarithmic amplitude of each audio frame. The absolute value of speech rate indicates the speed of speech. It is calculated by counting the number of syllables or words per unit time and is suitable for distinguishing between fast, normal, and slow speech rates.

[0055] After extracting the original acoustic feature indicators, the system calculates the pitch difference, intensity difference, and speech rate difference between two adjacent independent speech segments. These are called dynamic acoustic feature indicators. Dynamic acoustic features reflect the changing trend of the speech signal, such as a sudden increase in pitch (indicating emotional excitement) or a gradual decrease in volume (indicating emotional fading). The difference value calculation formula is:

[0056] Pitch difference = current segment pitch - previous segment pitch. Intensity difference = current segment intensity - previous segment intensity. Speed ​​difference = current segment speed - previous segment speed.

[0057] To further enhance the expressiveness of acoustic features, the system filters out extreme values ​​in the original acoustic feature metrics that are above the third quartile or below the first quartile. These are labeled as high-fluctuation features. This extreme value screening uses a statistical quantile approach. For example, a pitch above the third quartile indicates emotional excitement, while a pitch below the first quartile indicates depression or weakness. This screening method enables the system to capture the salient features of emotional fluctuations.

[0058] Finally, the system combines the original acoustic feature indicators, dynamic acoustic feature indicators, and high-fluctuation feature markers to generate comprehensive acoustic feature indicators. These comprehensive acoustic features include information from multiple dimensions, such as:

[0059] Average pitch, maximum pitch, minimum pitch, pitch difference;

[0060] Average sound intensity, maximum sound intensity, minimum sound intensity, sound intensity difference;

[0061] Average speaking speed, maximum speaking speed, minimum speaking speed, speaking speed difference;

[0062] High-fluctuation feature markers (e.g., strong fluctuations in high notes, fluctuations in low notes’ pitch).

[0063] These comprehensive acoustic features can fully reflect the emotional information in speech signals, especially in scenarios with large emotional fluctuations and speech speed changes (such as disputes or complaints), and have higher emotional expression capabilities.

[0064] In the FinTech sector, acoustic feature extraction can be applied to call log analysis in customer service centers. The system extracts comprehensive acoustic feature metrics from voice conversations between customers and customer service representatives, analyzing changes in the customer's mood by comparing pitch and speech rate. For example, if a customer's speech speed suddenly increases, their pitch rises, and their voice intensity increases, the system will flag this as "highly emotional," prompting the customer service representative to adopt a soothing strategy. On the other hand, if a customer's speech speed slows, their voice intensity decreases, and their pitch is low, the system might flag this as "depressed," making it appropriate to recommend financial products to soothe them.

[0065] In the healthcare sector, acoustic feature extraction can be applied to remote health consultations or psychological counseling. The system extracts comprehensive acoustic feature indicators from voice calls between patients and doctors. For example, doctors' voices typically have a steady pitch and a fast speaking rate, while patients may experience pitch variations or abnormal speaking rates due to emotional fluctuations. The system can use high-volume feature markers in patients' voices to determine their emotional state, such as anxiety, depression, or agitation, providing additional reference for doctors.

[0066] This embodiment extracts acoustic feature indicators from independent voice segments, allowing the system to fully capture the emotional information in voice signals. Compared to emotion recognition methods based solely on speech text or a single acoustic feature, comprehensive acoustic feature indicators can reflect the multi-dimensional information of the sound, including the absolute values ​​and dynamic changes of pitch, intensity, and speech rate. In particular, through high-volatility feature markers, the system can accurately identify emotional fluctuations and emotional anomalies, effectively improving the accuracy and robustness of emotion recognition. Through this design, the system can accurately distinguish emotional fluctuations in complex conversation scenarios and is suitable for a variety of application scenarios such as financial technology and healthcare.

[0067] S40, inputting the independent speech segment into a speech emotion recognition model and outputting a preliminary emotion label;

[0068] In this embodiment, the process of inputting independent speech segments into a speech emotion recognition model and outputting preliminary emotion labels involves multimodal input construction, emotion recognition model reasoning, and emotion label generation. Independent speech segments refer to audio data obtained through intelligent segmentation and speaker separation, with each segment containing speech data from a single speaker. The preliminary emotion label represents the model's initial judgment of the emotional state based on the input data, typically including an emotion category (e.g., anger, satisfaction, sadness) and its corresponding confidence score.

[0069] First, the system concatenates independent speech segments with the corresponding current speech text into multimodal input data. This multimodal input combines speech signals and text information to more comprehensively capture speech emotions. The audio signal provides acoustic features of emotion, such as pitch, volume, and speaking speed, while the text information reflects the speaker's language content and expression. The system converts speech and text into vector representations through an encoder, where the audio part uses Mel-frequency cepstral coefficients (MFCC), spectrograms, or waveform encoding, and the text part uses word vectors (Word2Vec), contextual embeddings (BERT), or other natural language processing models.

[0070] After the multimodal input is constructed, the system feeds the multimodal input data into a pre-trained emotion recognition model. This emotion recognition model can be a deep learning model such as a convolutional neural network (CNN), a recurrent neural network (RNN), a multi-layer perceptron (MLP), or a multimodal transformer. The model extracts features from the multimodal input data using a multi-layer neural network and outputs an emotion category and confidence score at the final classification layer. The emotion category represents the model's emotional judgment of the current speech segment, such as anger, satisfaction, sadness, or neutrality, while the confidence score indicates the model's confidence in this judgment, typically a value between 0 and 1.

[0071] The system can also constrain the output categories of the emotion recognition model by setting preset text prompts. Preset text prompts are conditions added to the model's inference process. They usually use fixed text descriptions (such as "Please select the most appropriate emotion from anger, satisfaction, sadness, and neutral") to ensure that the emotion category output by the model falls within the preset category range.

[0072] After the emotion recognition model outputs the raw emotion results, the system further filters them using a confidence threshold. The confidence threshold represents the lower limit of the confidence level of the emotion judgment. When the confidence score output by the model falls below the preset threshold, the system marks the emotion category as pending review. This filtering mechanism effectively eliminates low-confidence emotion judgments, preventing erroneous results from interfering with subsequent emotion analysis.

[0073] For emotion labels with confidence levels below a threshold, the system corrects them through historical context analysis. The system retrieves historical emotion labels and the corresponding conversation text, and performs temporal correlation analysis in conjunction with the current emotion label. When there is a significant conflict between the current emotion label and the historical emotion label (e.g., the previous conversation was labeled "satisfied," while the current label is "angry"), the system corrects the emotion label using contextual consistency rules, such as re-labeling based on the historical label majority rule or emotion continuity.

[0074] Finally, the system combines the original emotion labels with the revised emotion labels that meet the confidence threshold to generate the final preliminary emotion labels. The preliminary emotion labels serve as input for multimodal sentiment analysis and provide the basis for subsequent sentiment determination and strategy generation.

[0075] Example: In the fintech sector, a customer service center's call logs are analyzed using a speech emotion recognition model. If a customer's voice clip contains the emotion tag "angry" with a confidence level higher than 0.8, the system generates an alert prompting the customer service representative to prioritize soothing tactics and monitor customer emotional changes using historical emotion tags. If the customer displays "angry" emotion for three consecutive rounds of conversation, the system automatically escalates the service request and transfers the call to a senior customer service representative or complaint handling specialist.

[0076] In the healthcare sector, voice calls in remote psychological counseling services use an emotion recognition model to generate real-time emotion labels. If a patient's emotion label is "sad" with a confidence level above 0.85, the system provides the doctor with psychological comfort strategies. If the emotion label is "anxious" with a confidence level above 0.9, the system prompts the doctor to guide the patient through relaxation training. If a patient's emotion consistently displays "anxiety" over multiple conversations, the system automatically records the trend of emotion changes and prompts the doctor to further assess the patient's mental health.

[0077] This embodiment significantly improves the accuracy and robustness of emotion recognition by inputting independent speech segments into a speech emotion recognition model and performing emotion analysis based on multimodal data. Compared to emotion recognition methods based solely on single acoustic or text features, multimodal fusion can capture richer emotional information from speech and text. Confidence filtering and context correction mechanisms can further enhance the reliability of emotion recognition and prevent low-confidence labels from affecting the emotion determination results.

[0078] S50, retrieving historical conversation text and combining it with the current voice text to generate context information;

[0079] In this embodiment, the process of retrieving historical conversation text and combining it with the current voice text to generate contextual information involves the retrieval and integration of text data, and the construction of a contextual structure. Historical conversation text refers to the conversation content preceding the current voice segment, typically including multiple rounds of interaction between the customer and customer service representatives or other roles. The current voice text is the text data obtained from the current independent voice segment using an automatic speech recognition model. By combining the current voice text with historical conversation text, the contextual information enables the system to consider the continuity and semantic relevance of the conversation in sentiment analysis.

[0080] First, the system retrieves the historical conversation text related to the current voice segment from a historical conversation database or real-time conversation cache. Historical conversation text can be retrieved based on timestamp, session identifier, or speaker role type to ensure that the retrieved text content is consistent with the current conversation context. In customer service scenarios, historical conversation text typically includes questions raised by customers and customer service replies; in healthcare, historical text may include patient descriptions and doctor responses.

[0081] After retrieving the historical conversation text, the system combines the current speech text with the historical conversation text to form a complete context. This combination can be simple text concatenation or a time-based conversation nesting. For example, the system can sort the historical conversation text and the current speech text chronologically to ensure the logical continuity of the context information. For multi-turn conversations, the system can also mark the speaker role (such as customer or customer service) and time information of each turn in the context information.

[0082] Contextual information can be further structured. The system can divide the text into multiple paragraphs based on conversation turns, speaker roles, and semantic topics, with each paragraph representing a single conversation turn. The system can also convert contextual information into vector representations using text vectorization techniques (such as large language models like BERT and GPT), preserving the semantic characteristics of the text. For longer conversations, the system can use a sliding window mechanism to ensure that the length of contextual information is controlled, retaining only the most recent conversation turns to avoid lengthy text that reduces the efficiency of sentiment analysis.

[0083] The generated contextual information includes not only the textual content itself but also semantic and emotional features. For example, the system can perform sentiment analysis on each historical conversation, generating emotion labels and confidence levels for each conversation. These labels are then embedded in the contextual information, providing a reference for multimodal sentiment fusion. The textual portion of the contextual information retains the original language expression, while the emotion labels and confidence levels are attached to the contextual information as structured metadata.

[0084] Contextual information, as one of the core inputs for multimodal sentiment analysis, can significantly improve the accuracy of sentiment recognition. By incorporating historical conversation text, the system can fully consider the continuity and contextual relevance of conversations in sentiment analysis, avoiding misjudgments in single-sentence sentiment recognition. For example, if historical conversation text shows a customer expressing dissatisfaction three times in a row, the system can more accurately identify the current sentiment as "angry" rather than "neutral."

[0085] Example: In the fintech sector, call logs in a customer service center are integrated through a contextual information generation module. When a customer asks, "Why is the charge so high?" or "I didn't sign up for automatic renewal," and becomes increasingly agitated, the system can identify the customer's dissatisfaction based on the contextual information and promptly provide soothing advice to the customer service representative, preventing the conflict from escalating.

[0086] In the healthcare sector, the system uses contextual information to help doctors understand patients' emotional trends during remote psychological consultations. For example, if a patient repeatedly expresses "I feel helpless" and their tone of voice gradually deteriorates, the system generates sentiment analysis based on this contextual information, prompting doctors to pay attention to the patient's emotional fluctuations and provide targeted guidance during subsequent consultations.

[0087] In a robot-assisted Q&A scenario, a user repeatedly asks questions like "How do I return a product?", "What is the return process?", and "How do I contact customer service?" The system uses contextual information to consolidate these multiple questions into a unified context. Using semantic analysis, it determines the user's interest in return-related information. Based on this context, the robot automatically selects the most appropriate answer, improving the accuracy and coherence of the response.

[0088] This embodiment generates contextual information by retrieving historical conversation text and combining it with the current speech text. This allows the system to consider the continuity and semantic relevance of conversations in sentiment analysis, avoiding misjudgments caused by single-sentence sentiment recognition. This contextual information enables the system to accurately capture the evolution of emotional changes in a conversation, particularly effectively identifying emotional continuations or sudden changes in multiple rounds of interaction. This design can effectively improve the accuracy of emotion recognition, especially in complex conversation scenarios such as customer service centers or psychological counseling.

[0089] S60: Input the context information, the preliminary emotion label, the speaker role type, and the acoustic feature index into a multimodal fusion module to generate an emotion determination result.

[0090] In this embodiment, contextual information, preliminary emotion labels, speaker role types, and acoustic feature indicators are input into the multimodal fusion module. The process of generating emotion determination results involves the fusion and analysis of multiple data types. This multimodal fusion design aims to improve the accuracy and stability of emotion recognition by integrating multiple dimensions of information, including speech, text, and emotion. Each input data type provides a different dimension of information during the fusion process, ensuring that the system can fully understand the emotional state of the conversation context.

[0091] First, the system receives contextual information, which is generated by retrieving historical conversation text and combining it with the current speech. This contextual information not only contains the current speech content but also preserves the semantic connections across multiple rounds of conversation, providing the contextual basis for sentiment analysis. This contextual information is preprocessed into a text vector representation to ensure data structure consistency during multimodal fusion. For long texts, the contextual information can be segmented using a sliding window mechanism to avoid computational delays caused by excessively long text.

[0092] Next, the system receives preliminary emotion labels, which are the emotion categories and confidence levels identified from individual speech segments by the speech emotion recognition model. Each preliminary emotion label represents the system's initial assessment of the emotional state of the current speech segment and is accompanied by a confidence score, indicating the reliability of this assessment. These preliminary emotion labels serve as emotional feature inputs in the multimodal fusion process, providing the system with a preliminary reference for emotion analysis.

[0093] The system then receives the speaker role type, which is role information determined from the current speech text using a semantic analysis model. The speaker role type identifies the current speaker's identity, such as customer service representative, client, doctor, or patient. In multimodal fusion, the speaker role type is used to distinguish the source of emotions, ensuring that sentiment analysis can distinguish the emotional states of different roles. For example, a customer expressing dissatisfaction and a customer service representative expressing reassurance belong to different emotion types.

[0094] Acoustic feature metrics represent speech feature data extracted from individual speech segments, including pitch, intensity, speaking rate, and their dynamic variations (e.g., pitch differential and speaking rate differential). These acoustic features can provide emotional information hidden in the speaker's speech expression; for example, a sudden rise in pitch may indicate excitement, while a slow speaking rate may indicate hesitation. In multimodal fusion, acoustic features are encoded as numerical vectors and normalized to ensure numerical consistency with the text vectors and sentiment labels during the fusion process.

[0095] In the multimodal fusion module, the above-mentioned multiple input data are encoded into multimodal feature vectors. Text data (context information) generates text feature vectors through large language models (such as BERT, GPT), emotion labels and speaker role types are input in the form of one-hot encoding or embedding, and acoustic features are vectorized to maintain numerical features. The system calculates the weight of each modal feature through a hierarchical attention network model and generates weighted fusion features based on the weights. The weight of each modality can be adjusted dynamically. For example, when the emotional confidence is high, the weight of the emotional label can be increased, and when the voice pitch fluctuates greatly, the weight of the acoustic feature can be increased.

[0096] The fused feature vectors are input into the fully connected classification layer, where the system generates a sentiment assessment based on the classification results. This assessment not only includes the final emotion category (e.g., anger, satisfaction), but also includes the basis for the assessment, such as key sentences in the context, high-volatility markers in the acoustic features, or indicators of interruption or hesitation. This basis ensures interpretability of sentiment analysis, facilitating subsequent applications in customer service assistance or automated speech generation.

[0097] Example: In the fintech sector, a customer service center system receives a customer voice stream. The extracted context shows the customer repeatedly asking, "Why are the charges so high?" The initial emotion tag indicates "angry," and acoustic features indicate a rapid increase in pitch. Using a multimodal fusion module, the system integrates the context, emotion tag, and acoustic features to ultimately determine the customer's emotion as "angry," generating the judgment criteria of "increased speech rate and rising pitch."

[0098] In the healthcare sector, a patient described themselves as "very stressed lately" during a remote psychological consultation. The current audio text displayed "I really don't know what to do," while the initial emotion label indicated "anxiety." The system used multimodal fusion to combine contextual information and acoustic features (strong bass, slow speech rate) to ultimately identify the emotion as "anxiety" and generated the criteria for this determination: "strong bass, slow speech rate, and context."

[0099] In a robot-based automated question-and-answer scenario, a user repeatedly asked questions like "How do I get a refund?" and "I want to cancel my refund." The initial sentiment label was "dissatisfied," and contextual information indicated that the user was repeatedly asking questions about refunds. The system, using its multimodal fusion module, determined the user's sentiment to be "dissatisfied" and generated the judgment criteria: "Continuous refund questions and dissatisfaction."

[0100] In this embodiment, by inputting contextual information, preliminary emotion labels, speaker role types, and acoustic feature indicators into the multimodal fusion module, the system is able to integrate information of multiple data types in sentiment analysis to ensure the accuracy of sentiment judgment results. Text information provides semantic and contextual references, emotion labels provide preliminary emotion judgments, role types distinguish the sources of emotions, and acoustic features reveal emotional clues in speech. Through multimodal fusion and hierarchical attention mechanisms, the system can dynamically adjust the weights of each modal feature and adaptively adjust the sentiment analysis results according to the actual context. This design significantly improves the robustness of emotion recognition in complex dialogue scenarios, especially in multi-round interactive scenarios such as customer service and psychological counseling, and can more accurately capture character emotional changes and provide explanations.

[0101] The present invention relates to the field of speech processing technology and can be applied to business scenarios such as financial technology and medical health. A speech emotion recognition method, device, equipment and medium based on contextual information are disclosed, comprising: receiving an original speech stream and performing intelligent segmentation processing to generate speech segments, and performing a speaker separation operation to generate independent speech segments; inputting the independent speech segments into an automatic speech recognition model to generate a current speech text, performing role differentiation on the current speech text through a semantic analysis model to determine the speaker role type; extracting acoustic feature indicators from the independent speech segments; inputting the independent speech segments into a speech emotion recognition model to output a preliminary emotion label; retrieving historical conversation texts and combining them with the current conversation text to generate contextual information; inputting the contextual information, preliminary emotion label, speaker role type and acoustic feature indicators into a multimodal fusion module to generate an emotion judgment result. By combining speech, text and role information and realizing multimodal fusion based on contextual information, the present invention can accurately identify and understand the emotional changes of each role in complex conversation scenarios, avoiding the problems of large errors in single-sentence emotion judgment and neglect of contextual information in traditional methods, and effectively improving the accuracy and stability of emotion recognition.

[0102] In one embodiment, the above step S10 includes:

[0103] S101, inputting the original speech stream into a speech endpoint detection model, dividing sentence boundaries according to speech energy mutation points, and generating initial speech segmentation;

[0104] S102, when it is detected that the original speech stream is a single-track recording, performing speaker feature clustering on the initial speech segments using a spectral clustering algorithm to generate speech segments of different speakers;

[0105] S103, when it is detected that the original voice stream is a dual-track recording, separating the speaker voice data in the initial voice segment according to the left and right channels to generate a voice segment containing only a single speaker;

[0106] S104, performing noise suppression filtering on the speech segments to generate pure speech segments;

[0107] S105 , performing identity consistency verification on the clean speech segments using a speaker verification model, and merging continuous speech segments of the same speaker to generate independent speech segments.

[0108] In this embodiment, the original audio stream is received and intelligently segmented to generate speech segments. Speaker separation is then performed on these segments to generate independent speech segments. This entire process, based on speech signal processing technology, ensures that each speaker's speech segment is accurately extracted from the original audio stream, forming independent speech segments to support subsequent speech recognition and sentiment analysis.

[0109] First, the system receives a raw voice stream, which is continuous audio data containing one or more conversations. It might come from a customer service call recording, a remote medical consultation, or an intelligent question-and-answer system. This raw voice stream, as input data, is typically acquired in real time or offline and can be a single-track or dual-track recording. A single-track recording means the voice data of all speakers is mixed in a single track, while a dual-track recording means the voices of different speakers are recorded separately in the left and right channels.

[0110] The original speech stream is input into a speech endpoint detection model, which determines sentence boundaries by analyzing energy changes in the audio signal. Specifically, the speech endpoint detection model detects the start and end points of sentences based on energy transitions in the audio signal. These energy transitions typically manifest as a rapid increase or decrease in the audio signal amplitude, indicating the beginning or end of speech. For example, when a customer suddenly begins speaking, the audio signal energy rapidly increases; when there is a long period of silence during a call, the audio signal energy drops sharply. By detecting these energy transitions, the system segments the continuous speech stream into multiple initial speech segments.

[0111] Next, the system uses different processing methods depending on the type of original voice stream (single-track or dual-track). When the original voice stream is detected as a single-track recording, the voice data of all speakers are mixed in one channel, and the system cannot directly distinguish the speakers by the channel. In this case, the system uses a spectral clustering algorithm to cluster the speaker features of the initial voice segments. The spectral clustering algorithm extracts the spectral features of each voice segment (such as Mel-frequency cepstral coefficients MFCC) and clusters voice segments with similar spectral features as the voice of the same speaker. The algorithm takes the feature vector as input, calculates the similarity matrix between the voice segments, and performs spectral clustering, ultimately separating the voice segments of different speakers.

[0112] When the system detects a dual-track recording of the original voice stream, it separates the voice data into left and right channels. The left and right channels each record the voices of different speakers, and the system directly extracts and processes the left and right channel audio separately. This approach avoids complex speaker separation algorithms and is more efficient in dual-track recordings. Dual-track separation is suitable for scenarios such as phone recording and online conferencing, where one party's voice is recorded in the left channel and the other party's voice is recorded in the right channel.

[0113] After completing sentence segmentation and speaker separation, the system performs noise suppression filtering on all speech segments. This process aims to eliminate background noise (such as ambient noise and call noise) to ensure clear speech signals. Noise suppression filters can be based on frequency domain methods (such as spectral subtraction) or time domain methods (such as adaptive filtering). Spectral subtraction calculates the power spectral density of the speech signal to identify and suppress background noise within a certain frequency range. Adaptive filtering suppresses noise by dynamically adjusting filter coefficients.

[0114] After obtaining clean speech segments, the system uses a speaker verification model to verify identity consistency. The speaker verification model analyzes the characteristics of each speech segment, such as timbre, speech frequency distribution, and pronunciation patterns, to determine whether the speech comes from the same speaker. Common methods include Gaussian mixture models (GMMs), i-vector models, and deep learning-based speech embedding models. For speech segments from the same speaker, the system automatically merges them to generate continuous and independent segments. This merging prevents the speech of the same speaker from being split into multiple segments, ensuring the integrity of each speaker's speech data.

[0115] Finally, the system outputs independent speech segments. Each independent speech segment represents the continuous speech of a single speaker, which is convenient for subsequent speech recognition and sentiment analysis.

[0116] This embodiment receives the original voice stream and performs intelligent segmentation, speaker separation, and noise suppression. The system ensures accurate separation and extraction of each speaker's independent voice segments in different recording types (single-track or dual-track). Whether it's a dual-track recording of a customer and customer service representative, or a single-track recording of a patient and doctor, the system ensures high fidelity and accurate separation of voice data. This design significantly improves voice processing accuracy in multi-speaker scenarios, laying the foundation for subsequent speech recognition and sentiment analysis.

[0117] In one embodiment, the above step S20 includes:

[0118] S201, performing audio gain equalization processing on the independent voice segment to generate a standardized voice segment;

[0119] S202, inputting the standardized speech segment into a pre-trained end-to-end speech recognition model, decoding frame by frame to generate a current speech text with a timestamp;

[0120] S203, extracting occupation-related terms, question sentences, and answer patterns from the current voice text to generate a role keyword set;

[0121] S204, determining the customer service role probability through a semantic analysis model based on the distribution density of the role keyword set;

[0122] S205, when the customer service role probability exceeds a preset probability threshold, marking the corresponding speaker role type as a customer service role, otherwise marking it as a customer role;

[0123] S206: Verify the dialog turn sequence of the marking results, correct the conflicting markings and generate the final speaker role type.

[0124] In this embodiment, independent speech segments are input into an automatic speech recognition model to generate a current speech text. This text is then analyzed using a semantic analysis model to identify the speaker's role type. This process combines speech recognition and semantic analysis technologies, using multi-layered feature extraction and model calculation to identify each speaker and ensure accurate role identification.

[0125] First, the system performs audio gain equalization on the individual speech segments to generate standardized segments. Audio gain equalization aims to eliminate potential volume fluctuations in the speech signal, ensuring consistent volume levels across all segments. This equalization not only improves speech recognition accuracy but also prevents character recognition bias caused by volume differences. Audio gain equalization can be achieved through linear amplification or attenuation, or by using adaptive gain control (AGC) technology to automatically adjust the volume, keeping the volume amplitude of all speech segments within the target range.

[0126] The standardized speech segments are input into a pre-trained end-to-end speech recognition model. The end-to-end speech recognition model is an integrated speech recognition system that can directly convert speech signals into text data without the need for traditional acoustic models, language models, and vocabulary separation. Common end-to-end models include those based on recurrent neural networks (RNNs), convolutional neural networks (CNNs), and Transformer architectures. The model converts speech signals into characters or words by decoding frame by frame and time step by time, and automatically annotates timestamps in the text. Timestamp information not only helps with subsequent semantic analysis, but also provides accurate time series basis for subsequent role differentiation and dialogue turn detection.

[0127] After generating the current voice text with a timestamp, the system extracts occupation-related terms, question sentences, and response patterns from the text. This feature extraction is mainly based on text analysis technology. The system uses word segmentation, part-of-speech tagging, and syntactic parsing technology to identify keywords with role characteristics from the text. For example, in customer service scenarios, customer service staff often use service-related terms such as "please ask," "need," and "for you," while customers often use terms such as "I want," "how," and "can" to express their needs. In addition, the detection of question sentences and response patterns further enhances the accuracy of role differentiation. For example, "Can you help?" is usually a question asked by a customer, while "I can provide you with the following services" is a response from a customer service representative.

[0128] The extracted career-related terms, question patterns, and response patterns are integrated into a role keyword set. This set is a set of keywords or phrases that reflect the characteristics of the roles in the text. Through this set, the system can quantify the distribution of various keywords in the text. The role keyword set can be matched not only based on a static vocabulary (such as "please ask" and "help"), but also based on an adaptive learning mechanism that dynamically expands the keyword library from the training data.

[0129] Based on the distribution density of the role keyword set, the system uses a semantic analysis model to calculate the probability of a customer service role for each text segment. Distribution density refers to the frequency of appearance of role keywords in a text segment. For example, if words like "please ask" and "help" appear frequently in a text segment, the system automatically increases the probability that the text is a customer service statement. The semantic analysis model can use a Transformer model based on the attention mechanism or a semantic classification model based on a convolutional neural network (CNN). This model takes the role keyword set as input and calculates the probability of a customer service role using a deep learning network.

[0130] When the system-calculated probability of a customer service role exceeds a preset probability threshold, the speaker in the text is labeled as a customer service role. This threshold is typically determined based on validation results during model training and represents the confidence level for role differentiation. For example, if the customer service role probability is greater than 0.7, the system automatically labels the speaker as a customer service role; otherwise, the speaker is labeled as a customer. This probability threshold can be flexibly adjusted in different scenarios to meet diverse business needs.

[0131] However, role differentiation based solely on single-sentence analysis can lead to misjudgments, especially in scenarios with highly continuous conversations. Therefore, the system further performs turn-by-turn verification on the labeled results. Turn-by-turn verification is a context-based verification mechanism that ensures the correctness of role labeling by analyzing the order of speech in a conversation turn. For example, in a typical customer service scenario, the customer service representative and the customer should speak alternately. Consecutive appearances of the same role may indicate a role labeling error. This sequential verification corrects potential role conflicts, such as a customer being mistakenly labeled as a customer service representative.

[0132] Finally, after verification, the system outputs the final speaker role type. This multi-level role differentiation mechanism ensures that the system can accurately identify each speaker's role in a variety of scenarios, providing an accurate role foundation for subsequent sentiment analysis and guidance speech generation.

[0133] This embodiment inputs independent speech segments into an automatic speech recognition model to generate a timestamp for the current speech. Using a semantic analysis model, the system automatically distinguishes character types, enabling accurate identification of each character in multi-character conversations. Audio gain equalization ensures speech signal quality and improves speech recognition accuracy. A set of character keywords captures semantic features to ensure accurate character differentiation. A combination of probability thresholds and turn-order verification further enhances the reliability of character labeling. This design effectively avoids character confusion in multi-speaker scenarios, ensuring that sentiment analysis and guidance speech generation are based on accurate character identification.

[0134] In one embodiment, the above step S30 includes:

[0135] S301, extracting the absolute value of pitch, absolute value of sound intensity, and absolute value of speech rate of the independent speech segment to generate original acoustic feature indicators;

[0136] S302, determining a pitch difference value, a sound intensity difference value, and a speech rate difference value between two adjacent independent speech segments to generate a dynamic acoustic feature index;

[0137] S303, screening the extreme values ​​of the original acoustic feature indicators that are higher than the third quartile or lower than the first quartile to generate high-fluctuation feature markers;

[0138] S304: Merge the original acoustic feature index, the dynamic acoustic feature index, and the high-fluctuation feature marker to generate a comprehensive acoustic feature index.

[0139] In this embodiment, acoustic feature indices are extracted from independent speech segments to extract multi-dimensional acoustic information from the speech signal that can reflect the speaker's voice characteristics and emotional changes. This process achieves a comprehensive analysis of the sound characteristics of the speech segment by calculating and screening multiple acoustic features.

[0140] First, the system extracts the absolute values ​​of pitch, intensity, and speaking rate from independent speech segments to generate raw acoustic feature metrics. The absolute value of pitch refers to the fundamental frequency (F0) of the speech signal and is typically extracted from the speech waveform using the autocorrelation method or the Harmonic Product Spectrum (HPS) algorithm. Pitch directly reflects the vibration characteristics of the speaker's vocal cords. In sentiment analysis, pitch changes can indicate anger (increased pitch) or sadness (decreased pitch). The absolute value of intensity refers to the amplitude or sound pressure level (SPL) of the speech signal and is typically calculated using short-time energy or root mean square (RMS) energy. Intensity reflects the force of the speaker's vocalization and is a key parameter for determining the intensity of speech. The absolute value of speaking rate refers to the number of syllables or words pronounced per second and is calculated by measuring the length of silent and spoken segments in the speech segment. Speaking rate is an effective indicator of a speaker's emotional tension or relaxation. Fast speaking rates are often associated with anger or anxiety, while slow speaking rates often indicate calmness or sadness.

[0141] After obtaining the raw acoustic feature indicators, the system further determines the pitch difference, intensity difference, and speaking speed difference between two adjacent independent speech segments to generate dynamic acoustic feature indicators. Dynamic acoustic features are indicators that measure the changing trend of speech features and can reflect the speaker's emotional fluctuations over a short period of time. For example, when the pitch difference value is continuously positive and the value is large, it usually indicates that the speaker is gradually becoming excited or angry. The intensity difference value reflects the change in volume. For example, a rapid increase in volume may indicate emotional excitement. The speaking speed difference value measures the change in the speaker's speaking speed. An increase in speaking speed often indicates urgency or nervousness, while a slowdown may indicate thinking or negativity.

[0142] After obtaining the original acoustic features and dynamic acoustic features, the system filters out extreme values ​​in the original acoustic feature indicators that are higher than the third quartile or lower than the first quartile to generate high-volume feature markers. Quantile is a feature selection method based on statistical distribution. The third quartile (75%) and the first quartile (25%) represent the upper quartile and lower quartile of the data distribution, respectively. Values ​​above the third quartile indicate higher feature values, while values ​​below the first quartile indicate lower feature values. This screening method can effectively identify abnormal sound features. For example, a sudden and sharp increase in pitch indicates an emotional outburst, while a rapid drop in sound intensity may indicate sudden silence. The system marks these extreme values ​​as high-volume features to identify abnormal emotions in subsequent sentiment analysis.

[0143] Finally, the system combines the original acoustic feature indicators, dynamic acoustic feature indicators and high-fluctuation feature markers to generate comprehensive acoustic feature indicators. Comprehensive acoustic features are a multidimensional feature vector that combines the absolute characteristics of the sound (such as pitch, sound intensity, and speaking speed), change characteristics (such as differential values), and abnormal characteristics (such as high-fluctuation markers). This multidimensional feature vector can provide rich information in the emotion recognition process. For example, in customer service scenarios, comprehensive acoustic features can not only identify the customer's emotion category (such as anger, satisfaction), but also detect emotion intensity and change trends (such as emotional outbursts or gradual calming). The system can further perform multimodal sentiment analysis based on this comprehensive feature to improve the accuracy of emotion recognition.

[0144] In this embodiment, by extracting the absolute values ​​of pitch, intensity and speaking speed from independent voice segments, as well as the differential values ​​of adjacent segments, the system can not only accurately capture the basic acoustic features of the voice signal, but also identify sound change trends and abnormal fluctuations. This multi-dimensional acoustic feature extraction method ensures the accuracy of emotion recognition, especially in multi-role dialogues or scenes with obvious emotional fluctuations. High-fluctuation feature markers help the system quickly identify extreme emotional changes such as emotional outbursts or depression, while comprehensive acoustic feature indicators provide comprehensive input data for subsequent multimodal emotion analysis. By combining absolute features, dynamic features and abnormal markers, the system can accurately capture the details of voice emotions and achieve accurate judgment of emotions.

[0145] In one embodiment, the above step S40 includes:

[0146] S401, splicing the independent voice segment and the corresponding current voice text into multimodal input data;

[0147] S402: inputting a preset text instruction into the pre-trained multimodal emotion recognition model, wherein the preset text instruction is used to limit the emotion category output by the multimodal emotion recognition model to a preset emotion category;

[0148] S403, inputting the multimodal input data into the multimodal emotion recognition model to generate an original output result including an emotion category and a confidence score;

[0149] S404, filtering the original output result by a confidence threshold, and when the confidence score is lower than a preset confidence threshold, marking the corresponding emotion category as a label to be reviewed;

[0150] S405: Obtain historical emotion labels corresponding to historical voice segments in the current dialogue round

[0151] S406, performing a temporal correlation analysis on the emotion categories in the to-be-reviewed labels, and mapping the emotion categories to corresponding categories in the preset emotion categories in combination with the historical emotion labels to generate a revised emotion label;

[0152] S407 , combining the corrected emotion label with the original output result whose confidence score is not less than a preset confidence threshold to generate a final preliminary emotion label.

[0153] In this example, inputting independent speech segments into a speech emotion recognition model and outputting preliminary emotion labels is a key step in achieving multimodal emotion recognition. This fusion analysis of multimodal information enables high-precision and reliable emotion recognition. This process encompasses the construction of multimodal input data, emotion recognition model inference, correction of low-confidence labels, and final emotion label generation, ensuring the accuracy and stability of emotion recognition results.

[0154] First, the system concatenates the independent speech segment and the corresponding current speech text into multimodal input data. The independent speech segment is a pure speech segment that has been intelligently segmented and speaker separated in the previous step and has clear speaker attributes. The current speech text is the text content converted from the speech segment through the automatic speech recognition model (ASR). The construction of multimodal input data is the basis of multimodal emotion recognition. Its core lies in the synchronous combination of speech signals (audio data) and text signals (speech text) to form a complete contextual expression. This multimodal data can be encoded in various forms, such as concatenating speech features (such as MFCC, pitch, and intensity) with text features (such as word vectors and sentence vectors), or using temporal alignment technology to correspond speech frames to text words one by one.

[0155] After the multimodal input data is constructed, the system inputs preset text instructions to the pre-trained multimodal emotion recognition model. The preset text instructions are used to limit the emotion categories output by the multimodal emotion recognition model to preset emotion categories. The preset text instructions are a control strategy for constraining the output of the emotion recognition model. Its content is usually presented in natural language, such as "identify emotions as anger, satisfaction, boredom, tension or neutrality." This type of text instruction not only ensures that the model is strictly limited to the preset category when outputting emotion categories, but also can dynamically adjust the output emotion categories according to the actual application scenario. For example, "anxiety" and "excitement" categories can be added in customer service, and "worry" and "depression" categories can be added in medical consultation.

[0156] The system then feeds the multimodal input data into a multimodal emotion recognition model, generating a raw output consisting of emotion categories and confidence scores. Multimodal emotion recognition models typically employ convolutional neural networks (CNNs), recurrent neural networks (RNNs), or Transformer architectures to simultaneously process audio and text features. The model's output is a probability distribution over multiple emotion categories, with the probability corresponding to each category representing the confidence score for that category. The confidence score represents the model's confidence in its emotion judgment; higher values ​​indicate a greater degree of confidence in the emotion classification.

[0157] To ensure the accuracy of emotion recognition results, the system filters the raw output using a confidence threshold. When the confidence score falls below a preset threshold, the system marks the corresponding emotion category as pending review. The preset confidence threshold is typically determined based on the application scenario; for example, it might be set to 0.7 in customer service, while it might be increased to 0.8 in medical consultations. This filtering mechanism eliminates emotion labels with insufficient confidence, reducing the impact of misidentification.

[0158] After the tags to be reviewed are marked, the system further initiates a correction mechanism based on the time context. In order to provide a reasonable basis for the correction of emotion categories, the system first obtains the historical emotion tags corresponding to the historical voice segments in the current dialogue round, specifically including the emotion categories that have been identified and have met the confidence level in the speaker's continuous sentences in this round and the previous round. Historical emotion tags are usually stored in a cache structure after being bound to the speaker's identity through a timestamp, and support calls by round, speaker identity, and relative time window. In complex dialogue scenarios, the length of the historical lookback window can also be limited according to business rules (such as the first 3 rounds, within 5 seconds, continuous speaking of the same character, etc.) to ensure that the timing dependency is stable and explainable.

[0159] Based on the above historical emotion labels, the system performs a temporal association analysis on the emotion categories in the labels to be reviewed. This analysis process constructs an emotion trajectory model, identifies the evolution trend of historical labels, and combines the contextual logical relationship to perform trend inference and mapping operations on the labels with lower confidence levels. For example, if a user presents "anxiety" or "anger" type labels in three consecutive rounds, and the label in the current round is "neutral" but the confidence level is only 0.45, the system can infer that the "neutral" emotion is at risk of misidentification out of context, and chooses to map it to the "anxiety" category that is consistent with the historical trend. The mapping process is performed under the premise of ensuring that the labels come from the preset emotion category set to avoid generating label categories that exceed the model pre-training boundaries.

[0160] Finally, the system merges the corrected emotion labels with the original output results whose confidence scores are at least a preset confidence threshold to generate the final preliminary emotion labels. This merging strategy ensures that high-confidence emotion labels are retained while low-confidence labels are corrected, forming a stable and accurate set of emotion recognition results. The final preliminary emotion labels contain both high-confidence emotion judgment results and the contextual correction results of low-confidence labels, fully reflecting the emotional information conveyed in the speech segment.

[0161] This embodiment combines independent voice segments with text content to construct multimodal input data and outputs emotion categories and confidence scores through a pre-trained emotion recognition model, so that the system can achieve high-precision emotion recognition. Confidence threshold filtering ensures the reliability of the output emotion labels, and the mechanism of correcting low-confidence labels through temporal correlation analysis further improves the accuracy of emotion judgment. Ultimately, the preliminary emotion labels generated by the system can not only accurately capture the emotion of the voice, but also dynamically adapt to the emotional changes in the conversation context, realizing flexible emotion recognition for multiple roles and multiple situations. Through this process, the system can provide stable and accurate emotion analysis support in a variety of scenarios such as customer service, medical consultation, and intelligent assistants.

[0162] In one embodiment, the above step S60 includes:

[0163] S601, detecting the speech overlap duration and overlap ratio between independent speech segments of different speakers in the same time period, and generating a speech-interruption indicator when the speech overlap duration exceeds a preset speech-interruption threshold;

[0164] S602, detecting the length of the response interval between independent speech segments of the questioner and the responder in adjacent conversation turns, and generating a hesitation indicator when the length of the response interval exceeds a preset hesitation threshold;

[0165] S603: uniformly encode the historical conversation text in the context information, the emotion category and confidence score in the preliminary emotion label, the customer service or client identifier in the speaker role type, the acoustic feature index, the speech overlap duration and overlap ratio in the talk-interruption index, and the response interval duration in the hesitation index into a time-aligned multimodal feature vector;

[0166] S604: extracting emotion conflict markers in each time window of the multimodal feature vector, and marking as potential service conflict events when the speaker role type is customer service and the emotion category is negative in the same time window;

[0167] S605: adjusting the fusion weight of the acoustic feature index based on the marker density of the potential service conflict event and the values ​​of the talk-interception index and the hesitation index to generate a weighted multimodal feature;

[0168] S606: Input the weighted multimodal features into a pre-trained hierarchical attention network model to determine the text modality weight, speech modality weight, conversation-interruption index weight, hesitation index weight, and context-related weight respectively;

[0169] S607, performing weighted fusion on the multimodal feature vector based on the text modality weight, the speech modality weight, the talk-interruption index weight, the hesitation index weight, and the context association weight to generate a fused emotional feature;

[0170] S608: Perform full-connection classification processing on the fused emotional features, and output an emotional determination result including a final emotional category and determination basis.

[0171] In this embodiment, contextual information, preliminary emotion labels, speaker role type, and acoustic feature indicators are fed into a multimodal fusion module to generate an emotion determination result. This is a core step in implementing comprehensive multimodal information analysis in a speech emotion recognition system. This process integrates contextual information, emotion recognition results, role information, and acoustic features. Through a multimodal feature fusion model, accurate emotion determination and conflict identification are achieved, ensuring high accuracy and reliability of the emotion determination results.

[0172] First, the system detects the speech overlap duration and overlap ratio between independent speech segments of different speakers in the same time period in the multimodal fusion module. When the speech overlap duration exceeds the preset interruption threshold, the system generates an interruption indicator. The speech overlap duration indicates the length of time different speakers (such as customer service and customers) speak simultaneously in the same time period, while the overlap ratio indicates the proportion of the overlapping duration to the entire segment duration. The preset interruption threshold can be dynamically adjusted based on the specific application scenario. For example, it can be set to 0.2 seconds in the customer service scenario, and can be relaxed to 0.5 seconds in medical consultation. Through this dual constraint of time and ratio, the system ensures the accuracy of interruption recognition and avoids misjudgment due to short sound overlap (such as background noise).

[0173] The system then detects the duration of the response interval between independent speech segments of the questioner and responder in adjacent conversation turns. When the response interval exceeds a preset hesitation threshold, the system generates a hesitation indicator. The response interval refers to the time difference between the end of the questioner's speech and the beginning of the responder's speech. For example, in a customer service scenario, if a customer doesn't respond within two seconds after a customer service question is asked, this can be considered a normal response; however, if it exceeds three seconds, it may indicate hesitation or confusion. By setting a preset hesitation threshold, the system can flexibly distinguish between normal responses and abnormal hesitation, helping to identify potential communication barriers in the conversation.

[0174] After generating the interruption and hesitation indicators, the system uniformly encodes the historical conversation text from the contextual information, the emotion category and confidence score from the preliminary emotion label, the customer service or client identification from the speaker role type, acoustic feature indicators, the speech overlap duration and overlap ratio from the interruption indicator, and the response interval duration from the hesitation indicator into a time-aligned multimodal feature vector. A multimodal feature vector is a high-dimensional feature representation that seamlessly integrates multi-dimensional information by aligning the time series of multiple modal features, including text, speech, emotion, and conversation rhythm. Specifically, text features can include sentence or word vectors of the conversation content; speech features include pitch, intensity, speech rate, and their dynamic changes; emotion features consist of preliminary emotion labels and confidence scores; role type is used to distinguish between customer service and client; and the interruption and hesitation indicators reflect the characteristics of the conversation rhythm.

[0175] After constructing the multimodal feature vector, the system further extracts emotional conflict markers within each time window. If the speaker's role type is customer service and the emotion category is negative (such as anger or boredom) within the same time window, the system marks it as a potential service conflict event. Emotional conflict markers are a dynamic monitoring mechanism used to identify emotional anomalies in conversations that could lead to a decline in service quality. For example, if a customer service representative exhibits negative emotions while the customer behaves normally, the system will identify the customer service representative as emotionally out of control, potentially impacting service quality.

[0176] Next, the system dynamically adjusts the fusion weights of the acoustic feature indicators based on the marker density of potential service conflict events, combined with the values ​​of the interruption index and the hesitation index, to generate weighted multimodal features. The marker density indicates the frequency of conflict events within a unit time window, while the values ​​of the interruption and hesitation indicators represent the degree of abnormal conversation rhythm. Based on this information, the system can increase or decrease the weight of the acoustic features so that the conflict markers have a more significant impact on emotional judgment. For example, when the conflict marker density is high and the hesitation index value is large, the system can increase the weight of emotional acoustic features (such as pitch fluctuation and speaking speed).

[0177] The weighted multimodal features are input into a pre-trained hierarchical attention network model, and the system determines the text modality weight, speech modality weight, talk-interruption indicator weight, hesitation indicator weight, and context-related weight. The hierarchical attention network is a deep learning model based on the Transformer or multi-layer self-attention mechanism that can automatically learn the importance of different modalities in multimodal features. The text modality weight is used to measure the influence of the conversation text in emotional judgment, the speech modality weight represents the contribution of speech features (such as timbre and emotion), the talk-interruption and hesitation indicator weights reflect the importance of the conversation rhythm, and the context-related weight ensures that the emotional judgment can take into account the conversation history and emotional trends.

[0178] Based on these weights, the system performs a weighted fusion of the multimodal feature vectors to generate a fused sentiment feature. Weighted fusion is a feature integration strategy that seamlessly integrates multimodal information by summing the features of each modality according to their weights. The resulting fused sentiment feature retains the core characteristics of each modality and ensures the rationality of information contribution through weighted distribution.

[0179] Finally, the system performs fully connected classification on the fused emotional features, outputting an emotional assessment result that includes the final emotion category and the basis for the assessment. This assessment includes analysis of the chat interruption and hesitation indicators, which explain the source and basis of the emotional assessment. The fully connected classifier is a standard deep learning output layer structure that uses a softmax activation function to convert multidimensional emotional features into a probability distribution to determine the final emotion category.

[0180] This embodiment uses a multimodal fusion module to integrate contextual information, preliminary emotion labels, speaker role types, and acoustic feature indicators into a multimodal feature vector. This is then dynamically weighted using indicators such as interruption and hesitation. This allows the system to accurately identify emotional changes in complex conversational scenarios. These indicators enable the system to detect abnormalities in conversational rhythm, while the hierarchical attention network ensures that different modal information is weighted according to their importance, enabling highly accurate emotion assessment. This technology effectively improves the accuracy and robustness of emotion recognition, enabling stable emotion recognition services in a variety of scenarios, including finance, healthcare, and intelligent assistants.

[0181] In one embodiment, after the above step S60, the method further includes:

[0182] S701, matching a target scene sub-library corresponding to the emotion type from a preset scene database according to the emotion type in the emotion determination result;

[0183] S702, extracting conversation context features from the emotion determination result;

[0184] S703: Convert the conversation context features into a text vector representation, determine the semantic similarity between the text vector representation and each guiding speech in the target scenario sub-library, and generate a similarity matching list;

[0185] S704: Filtering a predetermined number of candidate guiding phrases whose similarities are higher than a predetermined similarity threshold according to the semantic similarity values ​​in the similarity matching list;

[0186] S705, based on the logical relevance and emotional continuity between the candidate guiding speech and the historical conversation text, scoring the contextual coherence of the first preset number of candidate guiding speech to generate a contextual coherence score;

[0187] S706 , prioritizing the candidate guiding speech phrases according to the context coherence scores, and generating a guiding speech phrase prompt list;

[0188] S707: Associate and encapsulate the guidance speech prompt list with the emotion type and determination basis in the emotion determination result to generate guidance speech information.

[0189] In this embodiment, after the multimodal fusion module generates the emotion judgment result, the system further matches the target scene sub-library corresponding to the emotion type from the preset scene database based on the emotion type in the emotion judgment result. The preset scene database is a multi-level structured database that stores guiding words according to different application scenarios (such as customer service, medical consultation, smart assistant, etc.) and emotion categories (such as anger, satisfaction, anxiety, indifference). Each target scene sub-library contains guiding words designed for specific emotion types and dialogue scenarios, such as soothing words for angry emotions or product recommendation words for satisfied emotions. Through this classified storage, the system can quickly match the corresponding guiding words after emotion recognition to ensure dialogue continuity and response flexibility.

[0190] After completing the target scenario sub-library matching, the system further extracts conversation context features from the emotion judgment results. Conversation context features include the current conversation text, historical conversation text, and the judgment basis in the emotion judgment results. The current conversation text refers to the speech-to-text result between the user and the system or customer service in the current round, and the historical conversation text refers to all the previous speech texts of the user in this conversation, forming context information. The judgment basis in the emotion judgment results includes emotion type (such as anger, satisfaction), interruption or hesitation indicators, and emotion conflict markers (such as customer service role but negative emotion). These conversation context features are the core input for subsequent speech generation, ensuring that the system not only selects speech based on the current emotion, but also can make personalized adjustments based on conversation history and context logic.

[0191] The system converts the extracted conversation context features into text vector representations and generates a similarity match list based on the semantic similarity between the text vector representations and each guiding speech in the target scenario sub-library. Text vector representation is a high-dimensional numerical representation that can capture text semantic information and is often generated by pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers) or RoBERTa. The calculation of semantic similarity uses the cosine similarity algorithm, and the cosine value of the angle between text vectors is used as the similarity score. The similarity match list is sorted from high to low by similarity score, and each list item includes a guiding speech and its similarity score. Through semantic similarity calculation, the system can quickly screen the guiding speech that best matches the current conversation context.

[0192] After generating a similarity match list, the system selects a preset number of candidate lead lines based on the similarity score, with similarities exceeding a preset similarity threshold. The preset similarity threshold is a dynamically configurable parameter. For example, it can be set to 0.7 in customer service scenarios and increased to 0.8 in medical consultations. The preset number of candidate lead lines ensures that the system can select the best line from multiple options while preventing irrelevant lines with lower similarity from interfering with the output. Candidate lead lines must not only meet the threshold requirements in terms of similarity, but also meet a preset limit in terms of quantity, such as the top five or top ten.

[0193] Next, the system scores the contextual coherence of a preset number of candidate guiding lines based on their logical relevance and emotional continuity with the historical conversation text. Logical relevance indicates the semantic consistency of the candidate guiding line with the conversation context. For example, when a user expresses concern, the guiding line should be reassuring rather than recommending. Emotional continuity indicates that the candidate guiding line is emotionally consistent with the conversation history. For example, after a customer expresses dissatisfaction, the system should avoid recommending a product immediately. The contextual coherence score is generated by weighting the logical relevance and emotional continuity scores. A higher score indicates that the guiding line is more consistent with the current conversation logic.

[0194] The system prioritizes candidate guiding lines based on their contextual coherence scores and generates a list of guiding line suggestions. The ranking is from high to low, with higher-scoring lines receiving higher priority. This list of guiding line suggestions not only ensures the system outputs the optimal line from multiple options but also uses contextual coherence scores to avoid responses that conflict with the conversational content.

[0195] Finally, the system associates and encapsulates the list of guidance prompts with the emotion type and judgment basis from the emotion determination results to generate guidance information. Guidance information is structured data that includes the current emotion type (such as anger), context (such as the current conversation text and historical text), judgment basis (such as interruption and hesitation indicators), and the final list of selected guidance prompts. This structured guidance prompt information ensures that the system can be flexibly used in multiple rounds of conversation and can provide clear response suggestions to customer service or automated voice assistants.

[0196] Example: In the financial sector, when a customer calls the platform's customer service hotline, the system first receives the customer's voice stream and performs intelligent sentence segmentation. It then uses a voice endpoint detection model to identify sentence boundaries and generate voice segments. For single-track recorded conversations, the system uses a spectral clustering algorithm to separate the different speakers in the voice segments, generating independent voice segments representing the customer and the customer service representative respectively. For dual-track recorded conversations, the system directly separates the voice data based on the left and right channels to ensure that each voice segment contains only a single speaker. The system also further performs noise suppression filtering on the voice segments to eliminate background noise and improve sound quality, ultimately generating clean voice segments. On this basis, the system uses a speaker verification model to verify the identity consistency of the voice in the clean voice segments, ensuring that each independent voice segment always belongs to the same speaker.

[0197] After each independent voice segment is generated, the system inputs it into the automatic speech recognition model, decoding it frame by frame to generate the current voice text with a timestamp. Combining the keywords, question sentences and response patterns in the text, the system distinguishes between customer and customer service roles through a semantic analysis model. For example, when a customer uses expressions such as "Excuse me," "Can you," or "Need," the system identifies them as customer roles; and when a customer service representative uses "Hello," "I'm sorry," or "I can help you," the system identifies them as customer service roles. For scenarios that cannot be clearly distinguished, the system further analyzes the role probability through the distribution density of the role keyword set, and ultimately labels the role type.

[0198] The system then extracts acoustic feature indicators from each independent speech segment, including the absolute value of pitch, intensity, and speech rate, to form raw acoustic feature indicators. It also generates dynamic acoustic feature indicators, such as pitch difference, intensity difference, and speech rate difference, by comparing feature changes between adjacent independent speech segments. To capture emotional fluctuations, the system screens out high-volume features from the raw acoustic feature indicators (such as outliers above the third quartile or below the first quartile) and integrates all acoustic features into a comprehensive acoustic feature indicator.

[0199] On this basis, the system splices the independent voice clips and the corresponding current voice text into multimodal input data, and inputs it into the pre-trained multimodal emotion recognition model. The model limits the output emotion categories to "satisfaction", "anger", "anxiety" and other categories through preset text instructions, and generates original output results containing emotion categories and confidence scores. When the confidence is lower than the preset threshold, the system marks the emotion category as a label to be reviewed. At this time, the system also compares and corrects the low-confidence emotion category with the historical emotion label through time series association analysis to ensure the accuracy of emotion recognition, and finally outputs the corrected emotion label.

[0200] To comprehensively analyze customer sentiment, the system further retrieves historical conversation text and combines it with the current speech text to generate contextual information. This contextual information includes not only the current conversation text but also historical text content from multiple rounds of conversation between the customer and the customer service representative, ensuring that sentiment analysis is contextualized across multiple rounds. The system inputs this contextual information, preliminary emotion labels, speaker role type, and acoustic feature indicators into the multimodal fusion module to generate a sentiment assessment result. In multimodal fusion, the system first detects the duration and overlap ratio of speech overlap between independent speech segments to generate a talk-over indicator; it also detects the duration of the response interval between the customer and the customer service representative in adjacent conversation rounds to generate a hesitation indicator. Through multimodal feature vector encoding, the system integrates text, acoustic, role, and conversation rhythm features into a time-aligned multimodal feature vector.

[0201] After generating the multimodal feature vector, the system further extracts emotional conflict markers within each time window. For example, when the customer service role's emotions are negative, it is marked as a potential service conflict event. In combination with the interruption and hesitation indicators, the system adjusts the fusion weights of the acoustic features to generate weighted multimodal features. Based on the hierarchical attention network model, the system determines the weights of text, voice, interruption, hesitation, and context association, and performs weighted fusion on the multimodal feature vector to generate the final emotional judgment result, including the emotion category and the judgment basis (such as the analysis results of the interruption and hesitation indicators).

[0202] Based on the emotion type in the emotion determination result, the system matches the target scenario sub-library corresponding to the emotion type from the pre-set scenario database. For example, when the emotion type is "anger," the system selects guiding dialogue from the soothing scenario sub-library. The system further extracts conversation context features, including the current conversation text, historical conversation text, and the judgment basis, and converts these context features into text vector representations. Based on the semantic similarity between the text vector representation and the guiding dialogue in the scenario sub-library, the system generates a similarity match list and selects candidate guiding dialogues whose similarity exceeds a preset threshold.

[0203] After screening candidate guidance lines, the system further optimizes their ranking based on contextual coherence scores. This scoring is based on the logical relevance and emotional continuity of the candidate guidance lines with the historical conversation text, ensuring that the generated guidance lines are not only semantically accurate but also emotionally consistent. Finally, the system encapsulates the screened and ranked guidance lines into guidance information, which customer service representatives or intelligent assistants can flexibly utilize in subsequent conversations.

[0204] In the field of healthcare, intelligent voice assistants deploy the same voice emotion recognition system. Patients consult about their condition through voice, and the system generates independent voice segments based on the voice stream and distinguishes between the roles of doctor and patient. By analyzing the acoustic features in the patient's voice (such as slowed speech speed and pitch fluctuations), the system can identify the patient's emotions, such as "anxiety" or "fear." In multimodal fusion, the system combines the patient's voice and text information to generate an emotional judgment result, such as "anxiety." Based on this emotion, the system matches the guiding words from the "comfort" sub-library of the preset scenario database, such as "Please don't worry, your symptoms can be confirmed through professional examination." The contextual coherence score ensures that the system selects the comforting words that best match the patient's emotions, thereby improving patient satisfaction.

[0205] In the field of intelligent customer service, when a customer becomes agitated over a billing issue, the system accurately identifies the customer's emotion as "angry" through multimodal emotion recognition and selects appropriate soothing language based on historical conversation information, such as "We apologize for the inconvenience. We will verify your billing issue as soon as possible." The system also uses contextual coherence scoring to ensure the generated language effectively soothes the customer.

[0206] This embodiment matches the emotion type in the emotion determination results with a pre-set scenario database and selects guidance words based on contextual features and semantic similarity. This allows the system to generate real-time responses that match the user's emotional state in different conversation scenarios. Contextual coherence scoring ensures that guidance words are not only semantically consistent with the user's speech, but also maintain a reasonable emotional continuity.

[0207] In one embodiment, a speech emotion recognition device based on context information is provided, and the speech emotion recognition device based on context information corresponds one-to-one to the speech emotion recognition method based on context information in the above embodiment. Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of a speech emotion recognition device based on contextual information according to the present invention. These modules include a speech preprocessing module 10, a speech recognition module 20, an acoustic feature extraction module 30, an emotion recognition module 40, a context processing module 50, and a multimodal fusion module 60. Each functional module is described in detail below:

[0208] The speech preprocessing module 10 is used to receive the original speech stream and perform intelligent sentence segmentation processing to generate speech segments, and perform speaker separation operation on the speech segments to generate independent speech segments;

[0209] The speech recognition module 20 is used to input the independent speech segment into an automatic speech recognition model to generate a current speech text, and perform role differentiation on the current speech text through a semantic analysis model to determine the speaker role type;

[0210] an acoustic feature extraction module 30, configured to extract acoustic feature indicators from the independent speech segments;

[0211] An emotion recognition module 40 is configured to input the independent speech segment into a speech emotion recognition model and output a preliminary emotion label;

[0212] A context processing module 50 is used to retrieve historical conversation texts and combine them with the current voice text to generate context information;

[0213] The multimodal fusion module 60 is configured to input the context information, the preliminary emotion label, the speaker role type, and the acoustic feature index into the multimodal fusion module to generate an emotion determination result.

[0214] In one embodiment, the speech preprocessing module 10 is specifically configured to:

[0215] Inputting the original speech stream into a speech endpoint detection model, dividing sentence boundaries according to speech energy mutation points, and generating initial speech segmentation;

[0216] When it is detected that the original speech stream is a single-track recording, a spectral clustering algorithm is used to perform speaker feature clustering on the initial speech segments to generate speech segments of different speakers;

[0217] When it is detected that the original voice stream is a dual-track recording, separating the speaker voice data in the initial voice segment according to the left and right channels to generate a voice segment containing only a single speaker;

[0218] Performing noise suppression filtering on the speech segments to generate pure speech segments;

[0219] The identity consistency of the clean speech segments is verified by a speaker verification model, and continuous speech segments of the same speaker are merged to generate independent speech segments.

[0220] In one embodiment, the speech recognition module 20 is specifically configured to:

[0221] Performing audio gain equalization processing on the independent voice segments to generate standardized voice segments;

[0222] Input the standardized speech segment into a pre-trained end-to-end speech recognition model, and decode it frame by frame to generate a current speech text with a timestamp;

[0223] Extracting occupation-related terms, question sentences, and answer patterns from the current voice text to generate a role keyword set;

[0224] Determine the customer service role probability through a semantic analysis model based on the distribution density of the role keyword set;

[0225] When the customer service role probability exceeds a preset probability threshold, the corresponding speaker role type is marked as a customer service role, otherwise it is marked as a customer role;

[0226] The marking results are verified in the order of dialogue turns, conflict markers are corrected, and the final speaker role type is generated.

[0227] In one embodiment, the acoustic feature extraction module 30 is specifically configured to:

[0228] Extracting the absolute value of pitch, absolute value of sound intensity, and absolute value of speech rate of the independent speech segment to generate original acoustic feature indicators;

[0229] Determine the pitch difference, intensity difference, and speech rate difference between two adjacent independent speech segments to generate dynamic acoustic feature indicators;

[0230] Screening extreme values ​​of the original acoustic feature indicators that are higher than the third quartile or lower than the first quartile to generate high-fluctuation feature markers;

[0231] The original acoustic feature index, the dynamic acoustic feature index and the high-fluctuation feature marker are combined to generate a comprehensive acoustic feature index.

[0232] In one embodiment, the emotion recognition module 40 is specifically configured to:

[0233] splicing the independent speech segment and the corresponding current speech text into multimodal input data;

[0234] Inputting a preset text instruction into the pre-trained multimodal emotion recognition model, wherein the preset text instruction is used to limit the emotion category output by the multimodal emotion recognition model to a preset emotion category;

[0235] Inputting the multimodal input data into the multimodal emotion recognition model to generate an original output result including an emotion category and a confidence score;

[0236] Performing confidence threshold filtering on the original output result, and when the confidence score is lower than a preset confidence threshold, marking the corresponding emotion category as a label to be reviewed;

[0237] Get the historical emotion labels corresponding to the historical voice clips in the current conversation round;

[0238] Performing a temporal correlation analysis on the emotion categories in the to-be-reviewed labels, and mapping the emotion categories to corresponding categories in the preset emotion categories in combination with the historical emotion labels to generate a revised emotion label;

[0239] The modified emotion label is combined with the original output result whose confidence score is not less than a preset confidence threshold to generate a final preliminary emotion label.

[0240] In one embodiment, the multimodal fusion module 60 is specifically configured to:

[0241] Detecting the speech overlap duration and overlap ratio between independent speech segments of different speakers within the same time period, and generating a speech-interruption indicator when the speech overlap duration exceeds a preset speech-interruption threshold;

[0242] Detecting the duration of the response interval between independent speech segments of the questioner and the responder in adjacent conversation turns, and generating a hesitation indicator when the duration of the response interval exceeds a preset hesitation threshold;

[0243] The historical conversation text in the context information, the emotion category and confidence score in the preliminary emotion label, the customer service or customer identifier in the speaker role type, the acoustic feature index, the speech overlap duration and overlap ratio in the talk-interruption index, and the answer interval duration in the hesitation index are uniformly encoded into a time-aligned multimodal feature vector;

[0244] Extracting emotion conflict markers within each time window in the multimodal feature vector, and marking as potential service conflict events when the speaker role type is customer service and the emotion category is negative within the same time window;

[0245] According to the marking density of the potential service conflict event, combined with the values ​​of the talk-interception index and the hesitation index, the fusion weight of the acoustic feature index is adjusted to generate a weighted multimodal feature;

[0246] Inputting the weighted multimodal features into a pre-trained hierarchical attention network model to determine the text modality weight, speech modality weight, conversation-interruption index weight, hesitation index weight, and context-related weight respectively;

[0247] Based on the text modality weight, the speech modality weight, the talk-interruption index weight, the hesitation index weight and the context association weight, the multimodal feature vector is weightedly fused to generate a fused emotional feature;

[0248] Performing full-connection classification processing on the fused emotional features, and outputting an emotional judgment result including the final emotional category and judgment basis.

[0249] In one embodiment, the multimodal fusion module 60 is specifically configured to:

[0250] According to the emotion type in the emotion determination result, matching a target scene sub-library corresponding to the emotion type from a preset scene database;

[0251] Extracting conversation context features from the emotion determination result;

[0252] Converting the conversation context features into a text vector representation, determining the semantic similarity between the text vector representation and each guiding speech in the target scenario sub-library, and generating a similarity matching list;

[0253] According to the semantic similarity values ​​in the similarity matching list, a preset number of candidate guiding phrases having similarities higher than a preset similarity threshold are screened;

[0254] Based on the logical relevance and emotional continuity between the candidate guiding speech and the historical conversation text, the context coherence of the first preset number of candidate guiding speech is scored to generate a context coherence score;

[0255] Prioritizing the candidate guiding speech phrases according to the context coherence scores and generating a guiding speech phrase prompt list;

[0256] The guidance speech prompt list is associated with the emotion type and judgment basis in the emotion judgment result and packaged to generate guidance speech information.

[0257] In one embodiment, a determination device is provided. The determination device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The determination machine device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the determination machine device is used to provide determination and control capabilities. The memory of the determination machine device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a determination machine program and a database. The internal memory provides an environment for the operation of the operating system and the determination machine program in the non-volatile storage medium. The network interface of the determination machine device is used to communicate with an external user terminal through a network connection. When the determination machine program is executed by the processor, it realizes the functions or steps on the service side of a speech emotion recognition method based on context information.

[0258] In one embodiment, a determination device is provided. The determination device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The determination machine device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the determination machine device is used to provide determination and control capabilities. The memory of the determination machine device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a determination machine program. The internal memory provides an environment for the operation of the operating system and the determination machine program in the non-volatile storage medium. The network interface of the determination machine device is used to communicate with an external server through a network connection. When the determination machine program is executed by the processor, it realizes the functions or steps on the user side of a speech emotion recognition method based on context information.

[0259] In one embodiment, a determination machine device is provided, including a memory, a processor, and a determination machine program stored in the memory and executable on the processor. When the processor executes the determination machine program, the following steps are implemented:

[0260] Receive the original speech stream and perform intelligent segmentation processing to generate speech segments, and perform speaker separation operation on the speech segments to generate independent speech segments;

[0261] Inputting the independent speech segment into an automatic speech recognition model to generate a current speech text, and performing role differentiation on the current speech text through a semantic analysis model to determine the speaker role type;

[0262] Extracting acoustic feature indicators from the independent speech segment;

[0263] Inputting the independent speech segment into a speech emotion recognition model and outputting a preliminary emotion label;

[0264] Retrieving historical conversation text and combining it with the current voice text to generate context information;

[0265] The context information, the preliminary emotion label, the speaker role type and the acoustic feature index are input into a multimodal fusion module to generate an emotion judgment result.

[0266] In one embodiment, a determination machine readable storage medium is provided, on which a determination machine program is stored. When the determination machine program is executed by a processor, the following steps are implemented:

[0267] Receive the original speech stream and perform intelligent segmentation processing to generate speech segments, and perform speaker separation operation on the speech segments to generate independent speech segments;

[0268] Inputting the independent speech segment into an automatic speech recognition model to generate a current speech text, and performing role differentiation on the current speech text through a semantic analysis model to determine the speaker role type;

[0269] Extracting acoustic feature indicators from the independent speech segment;

[0270] Inputting the independent speech segment into a speech emotion recognition model and outputting a preliminary emotion label;

[0271] Retrieving historical conversation text and combining it with the current voice text to generate context information;

[0272] The context information, the preliminary emotion label, the speaker role type and the acoustic feature index are input into a multimodal fusion module to generate an emotion judgment result.

[0273] It should be noted that the above functions or steps that can be implemented by the machine-readable storage medium or the machine device can be referred to the relevant descriptions on the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0274] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a determination machine program, and the determination machine program can be stored in a non-volatile determination machine-readable storage medium. When the determination machine program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0275] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0276] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A speech emotion recognition method based on context information, characterized in that: The following steps are involved: Receive the original speech stream and perform intelligent segmentation processing to generate speech segments, and perform speaker separation operation on the speech segments to generate independent speech segments; Inputting the independent speech segment into an automatic speech recognition model to generate a current speech text, and performing role differentiation on the current speech text through a semantic analysis model to determine the speaker role type; Extracting acoustic feature indicators from the independent speech segment; Inputting the independent speech segment into a speech emotion recognition model and outputting a preliminary emotion label; Retrieving historical conversation text and combining it with the current voice text to generate context information; The context information, the preliminary emotion label, the speaker role type and the acoustic feature index are input into a multimodal fusion module to generate an emotion judgment result.

2. The method for speech emotion recognition based on contextual information according to claim 1, wherein: Receive the original speech stream and perform intelligent segmentation processing to generate speech segments, and perform speaker separation operation on the speech segments to generate independent speech segments, including: Inputting the original speech stream into a speech endpoint detection model, dividing sentence boundaries according to speech energy mutation points, and generating initial speech segmentation; When it is detected that the original speech stream is a single-track recording, a spectral clustering algorithm is used to perform speaker feature clustering on the initial speech segments to generate speech segments of different speakers; When it is detected that the original voice stream is a dual-track recording, separating the speaker voice data in the initial voice segment according to the left and right channels to generate a voice segment containing only a single speaker; Performing noise suppression filtering on the speech segments to generate pure speech segments; The identity consistency of the clean speech segments is verified by a speaker verification model, and continuous speech segments of the same speaker are merged to generate independent speech segments.

3. The method for speech emotion recognition based on contextual information according to claim 1, wherein: Inputting the independent speech segment into an automatic speech recognition model to generate a current speech text, and performing role differentiation on the current speech text through a semantic analysis model to determine the speaker role type, including: Performing audio gain equalization processing on the independent voice segments to generate standardized voice segments; Input the standardized speech segment into a pre-trained end-to-end speech recognition model, and decode it frame by frame to generate a current speech text with a timestamp; Extracting occupation-related terms, question sentences, and answer patterns from the current voice text to generate a role keyword set; Determine the customer service role probability through a semantic analysis model based on the distribution density of the role keyword set; When the customer service role probability exceeds a preset probability threshold, the corresponding speaker role type is marked as a customer service role, otherwise it is marked as a customer role; The marking results are verified in the order of dialogue turns, conflict markers are corrected, and the final speaker role type is generated.

4. The method for speech emotion recognition based on contextual information according to claim 1, wherein: Extracting acoustic feature indicators from the independent speech segment includes: Extracting the absolute value of pitch, absolute value of sound intensity, and absolute value of speech rate of the independent speech segment to generate original acoustic feature indicators; Determine the pitch difference, intensity difference, and speech rate difference between two adjacent independent speech segments to generate dynamic acoustic feature indicators; Screening extreme values ​​of the original acoustic feature indicators that are higher than the third quartile or lower than the first quartile to generate high-fluctuation feature markers; The original acoustic feature index, the dynamic acoustic feature index and the high-fluctuation feature marker are combined to generate a comprehensive acoustic feature index.

5. The method for speech emotion recognition based on contextual information according to claim 1, wherein: The independent speech segment is input into the speech emotion recognition model to output a preliminary emotion label, including: splicing the independent speech segment and the corresponding current speech text into multimodal input data; Inputting a preset text instruction into the pre-trained multimodal emotion recognition model, wherein the preset text instruction is used to limit the emotion category output by the multimodal emotion recognition model to a preset emotion category; Inputting the multimodal input data into the multimodal emotion recognition model to generate an original output result including an emotion category and a confidence score; Performing confidence threshold filtering on the original output result, and when the confidence score is lower than a preset confidence threshold, marking the corresponding emotion category as a label to be reviewed; Get the historical emotion labels corresponding to the historical voice clips in the current conversation round; Performing a temporal correlation analysis on the emotion categories in the to-be-reviewed labels, and mapping the emotion categories to corresponding categories in the preset emotion categories in combination with the historical emotion labels to generate a revised emotion label; The modified emotion label is combined with the original output result whose confidence score is not less than a preset confidence threshold to generate a final preliminary emotion label.

6. The method for speech emotion recognition based on contextual information according to claim 1, wherein: Inputting the context information, the preliminary emotion label, the speaker role type, and the acoustic feature index into a multimodal fusion module to generate an emotion determination result, including: Detecting the speech overlap duration and overlap ratio between independent speech segments of different speakers within the same time period, and generating a speech-interruption indicator when the speech overlap duration exceeds a preset speech-interruption threshold; Detecting the duration of the response interval between independent speech segments of the questioner and the responder in adjacent conversation turns, and generating a hesitation indicator when the duration of the response interval exceeds a preset hesitation threshold; The historical conversation text in the context information, the emotion category and confidence score in the preliminary emotion label, the customer service or customer identifier in the speaker role type, the acoustic feature index, the speech overlap duration and overlap ratio in the talk-interruption index, and the answer interval duration in the hesitation index are uniformly encoded into a time-aligned multimodal feature vector; Extracting emotion conflict markers within each time window in the multimodal feature vector, and marking as potential service conflict events when the speaker role type is customer service and the emotion category is negative within the same time window; According to the marking density of the potential service conflict event, combined with the values ​​of the talk-interception index and the hesitation index, the fusion weight of the acoustic feature index is adjusted to generate a weighted multimodal feature; Inputting the weighted multimodal features into a pre-trained hierarchical attention network model to determine the text modality weight, speech modality weight, conversation-interruption index weight, hesitation index weight, and context-related weight respectively; Based on the text modality weight, the speech modality weight, the talk-interruption index weight, the hesitation index weight, and the context association weight, the multimodal feature vector is weightedly fused to generate a fused emotional feature; Performing full-connection classification processing on the fused emotional features, and outputting an emotional judgment result including the final emotional category and judgment basis.

7. The method for speech emotion recognition based on contextual information according to claim 1, wherein: After inputting the context information, the preliminary emotion label, the speaker role type, and the acoustic feature index into a multimodal fusion module to generate an emotion determination result, the method further includes: According to the emotion type in the emotion determination result, matching a target scene sub-library corresponding to the emotion type from a preset scene database; Extracting conversation context features from the emotion determination result; Converting the conversation context features into a text vector representation, determining the semantic similarity between the text vector representation and each guiding speech in the target scenario sub-library, and generating a similarity matching list; According to the semantic similarity values ​​in the similarity matching list, a preset number of candidate guiding phrases having similarities higher than a preset similarity threshold are screened; Based on the logical relevance and emotional continuity between the candidate guiding speech and the historical conversation text, the context coherence of the first preset number of candidate guiding speech is scored to generate a context coherence score; Prioritizing the candidate guiding speech phrases according to the context coherence scores and generating a guiding speech phrase prompt list; The guidance speech prompt list is associated with the emotion type and judgment basis in the emotion judgment result and packaged to generate guidance speech information.

8. A speech emotion recognition device based on context information, characterized in that: The speech emotion recognition device based on context information includes: A speech preprocessing module is used to receive the original speech stream and perform intelligent sentence segmentation processing to generate speech segments, and perform speaker separation operations on the speech segments to generate independent speech segments; A speech recognition module is used to input the independent speech segment into an automatic speech recognition model to generate a current speech text, and perform role differentiation on the current speech text through a semantic analysis model to determine the speaker role type; An acoustic feature extraction module, configured to extract acoustic feature indicators from the independent speech segments; An emotion recognition module, configured to input the independent speech segment into a speech emotion recognition model and output a preliminary emotion label; A context processing module is used to retrieve historical conversation texts and combine them with the current voice text to generate context information; The multimodal fusion module is used to input the context information, the preliminary emotion label, the speaker role type and the acoustic feature index into the multimodal fusion module to generate an emotion judgment result.

9. A determination device, characterized in that: The determination device includes a memory, a processor, and a context-based speech emotion recognition program stored in the memory and capable of running on the processor. When the context-based speech emotion recognition program is executed by the processor, the steps of the context-based speech emotion recognition method as described in any one of claims 1-7 are implemented.

10. A machine-readable storage medium, characterized in that: The storage medium stores a speech emotion recognition program based on context information, and when the speech emotion recognition program based on context information is executed by the processor, the steps of the speech emotion recognition method based on context information as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Data classification method and platform based on AI and electronic equipment

    CN120832411A

  • Multi-modal data collaborative interaction system and method applied to customer service platform

    CN121327164A

  • Intelligent voice interaction method and device

    CN121459790A

  • Interaction method based on emotional recognition and electronic equipment

    CN121561761A

  • Multi-modal big data analysis processing method and device based on artificial intelligence

    CN121580302A