Large model-based interpretable emotion recognition method

By extracting and screening basic acoustic features in speech data and training with large models, the problem of lack of transparency in existing speech emotion recognition technology is solved, and high accuracy and interpretive emotion recognition effect is achieved, which is suitable for psychological counseling and other applications.

CN119943095AActive Publication Date: 2025-05-06SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 18 Cites 0 Cited by

Patent Information

Application Number
CN202411872204.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The existing speech emotion recognition technology lacks transparency and is difficult to explain how the model draws emotional judgments, which limits its effectiveness in applications such as psychological counseling.

Method used

By extracting a variety of basic acoustic features from the original speech data, performing feature screening, building a data set, and using a large model for model training, an interpretable emotion recognition model is obtained. This model can not only accurately identify emotions, but also generate detailed reasoning processes to improve the explanatory nature of the model.

Benefits of technology

It realizes the high accuracy and interpretability of speech emotion recognition, can provide transparent and detailed emotion analysis results for users such as psychological counselors, and enhances the feasibility of emotion recognition application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943095A_ABST
    Figure CN119943095A_ABST
Patent Text Reader

Abstract

The invention provides an interpretable emotion recognition method based on a large model. The interpretable emotion recognition method comprises the steps of extracting various basic acoustic features for emotion classification from original voice data; performing feature screening processing on the basic acoustic features to obtain explanatory acoustic features closely related to specific emotions; constructing a data set based on the specific emotion, the corresponding interpretive acoustic features and emotion category labels; performing model training by using data in the data set to obtain an interpretable emotion recognition model; and performing audio emotion recognition through the interpretable emotion recognition model to obtain a complete emotion analysis result. The method aims at solving the problem that the decision making process in current emotion recognition is not transparent, the understanding ability of a large model for audio emotion is improved, and then the accuracy of emotion recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to an explainable emotion recognition method based on a large model. Background Art

[0002] Speech is the fastest and most natural form of human communication, and therefore, it is often used as a convenient bridge for human-computer interaction. This signal contains a lot of valuable emotional information. Currently, speech emotion recognition (SER) has become an indispensable component of human-computer interaction and other high-end speech processing systems. Usually, SER systems extract and classify key features from pre-processed speech signals to identify whether the speaker shows a specific emotion. However, there are significant quantitative and qualitative differences between humans and machines in the way they recognize and associate emotions in speech signals, which poses a huge challenge when integrating interdisciplinary knowledge, especially in the fields of speech emotion recognition, applied psychology, and human-computer interaction.

[0003] With the development of deep learning, many studies have adopted neural networks for emotion recognition. However, the current deep learning methods have poor interpretability and are usually unable to explain the specific reasons and basis for the model's classification decisions. This "black box" nature has brought challenges in applications that require transparent decision-making, such as emotion recognition. In recent years, large models have shown certain capabilities in emotion understanding. For example, natural language processing and emotional reasoning have improved the interpretability of decisions. By generating emotional descriptions and analysis, these models not only enhance the accuracy of emotion recognition, but also provide transparency for the model's decision-making process, which is conducive to more natural emotional interaction applications. Although the emotion recognition method based on deep learning has achieved good results, its decision-making process lacks transparency, making it difficult to understand how the model makes a certain emotional judgment. This "black box" nature has limitations in scenarios such as psychological dialogues that require accurate emotion tracking and interpretation, affecting the psychological counselor's judgment of the client's emotional state.

[0004] Existing emotion recognition technology has deficiencies in many aspects, which limits its effectiveness in practical applications. Most current emotion recognition models output emotion categories in the form of black boxes, lack a transparent decision-making process, and are difficult to provide detailed reasoning explanations, making it difficult for users such as psychological counselors to understand the basis of emotion recognition results; and existing technologies often cannot effectively guarantee the logical coherence of emotion analysis, and the generated emotion results lack a clear reasoning chain, which limits the interpretability and application feasibility of emotion recognition. Summary of the invention

[0005] In view of this, the present invention provides an interpretable emotion recognition method based on a large model to solve the above problems.

[0006] The present invention provides an interpretable emotion recognition method based on a large model, comprising: extracting a plurality of basic acoustic features for emotion classification from original speech data; performing feature screening processing on the basic acoustic features to obtain interpretative acoustic features closely related to specific emotions; constructing a data set based on the specific emotions, corresponding interpretative acoustic features and emotion category labels; performing model training using data in the data set to obtain an interpretable emotion recognition model; performing audio emotion recognition using the interpretable emotion recognition model to obtain a complete emotion analysis result.

[0007] In another implementation of the present invention, the basic acoustic features include prosody features, sound quality features and spectrum features.

[0008] In another implementation of the present invention, the prosodic features include duration-related features, fundamental frequency-related features and energy-related features; the duration-related features include speech rate and short-time average zero-crossing rate; the fundamental frequency-related features include fundamental frequency and its mean, variation range, variation rate, and mean square error; the energy-related features include short-time average energy and short-time average amplitude.

[0009] In another implementation of the present invention, the sound quality characteristics include time base error, amplitude perturbation, and harmonic-to-noise ratio.

[0010] In another implementation of the present invention, the frequency spectrum feature is a reflection of the correlation between the vocal tract shape change and the vocalization movement, and includes linear prediction cepstral coefficients and frequency cepstral coefficients.

[0011] In another implementation of the present invention, the data in the data set is used to perform model training to obtain an interpretable emotion recognition model, including: inputting the explanatory acoustic features in the data set into a pre-trained model to obtain a continuous high-dimensional acoustic feature encoding; decomposing the continuous high-dimensional acoustic feature encoding into a series of discrete speech units; using the speech units as input and the emotion labels and interpretable emotion clues as outputs to perform model training to obtain an interpretable emotion recognition model.

[0012] In another implementation of the present invention, the loss function of the interpretable emotion recognition model includes the cross entropy loss of emotion classification and the similarity loss of interpretability clues.

[0013] In another aspect of the present invention, an interpretable emotion recognition device based on a large model is provided, comprising: a feature extraction module: extracting a variety of basic acoustic features for emotion classification from original speech data; a feature processing module: performing feature screening processing on the basic acoustic features to obtain explanatory acoustic features closely related to specific emotions; a data preparation module: constructing a data set based on the specific emotion, the corresponding explanatory acoustic features and the emotion category label; a model training module: performing model training using the data in the data set to obtain an interpretable emotion recognition model; an emotion recognition module: performing audio emotion recognition through the interpretable emotion recognition model to obtain a complete emotion analysis result.

[0014] In another aspect of the present invention, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of an interpretable emotion recognition method based on a large model as described in any one of the above items are implemented.

[0015] In another aspect of the present invention, a computer storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in a large model-based interpretable emotion recognition method as described in any one of the above are implemented.

[0016] In the interpretable emotion recognition method based on a large model of the present invention, a multi-level feature extraction mechanism is introduced, which combines with basic acoustic features to enable the model to more deeply understand complex emotion information, enhance the overall emotion recognition effect, and improve the recognition accuracy; by guiding the large model to generate emotion understanding descriptions, emotion recognition is not limited to simple classification results, but also includes a detailed reasoning process, which facilitates psychological counselors to understand the emotional state of visitors; in addition, the thinking chain method is adopted to enhance the logical coherence of emotion analysis, and the output is not only an emotion label, but also includes a reasoning process, thereby significantly improving the interpretability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. By reading the detailed description of the following implementations, the advantages and benefits of the solutions become clear to those skilled in the art. The drawings are only used to illustrate the preferred implementations and are not to be considered as limiting the present invention.

[0018] In the attached picture:

[0019] Figure 1 The figure is a flowchart of an explainable emotion recognition method based on a large model according to an embodiment of the present invention.

[0020] Figure 2 This is a schematic diagram of the framework of an explainable emotion recognition system based on a large model according to an embodiment of the present invention.

[0021] Figure 3 The figure is a flow chart of a feature processing module according to an embodiment of the present invention.

[0022] Figure 4 The figure is a schematic diagram of the process of obtaining explainable emotional clues according to an embodiment of the present invention.

[0023] Figure 5 The figure is a flow chart of a model training module according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be described clearly and in detail below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in the field based on the embodiments in the embodiments of the present invention should fall within the scope of protection of the embodiments of the present invention.

[0025] Figure 1 A flow chart of an interpretable emotion recognition method based on a large model provided by an embodiment of the present invention is as follows: Figure 1 As shown, this embodiment mainly includes the following steps:

[0026] S101. Extract multiple basic acoustic features for emotion classification from original speech data.

[0027] For example, audio signals carry rich emotional information, and by extracting specific acoustic features, it is possible to analyze the potential changes in the speaker's emotions. In order to improve the accuracy of LLM's understanding of speech emotions, a variety of basic acoustic features that can be used for emotion classification are extracted from the original speech data. These basic acoustic features include but are not limited to pitch, pause duration, etc. They show obvious changes in different emotional states, which helps to identify and track emotions.

[0028] S102: Perform feature screening processing on the basic acoustic features to obtain explanatory acoustic features that are closely related to specific emotions.

[0029] For example, Figure 3 As shown in the figure, the extracted features are subjected to feature selection to obtain emotional clues. By screening out the explanatory acoustic features that are most closely related to specific emotions, LLM is ensured to focus on features with strong emotional discrimination and used in the data preparation module to generate explanatory sentiment analysis results.

[0030] S103: construct a data set based on the specific emotion, the corresponding explanatory acoustic features and the emotion category label.

[0031] Exemplarily, the data preparation module is used to obtain interpretable emotional clues, and by introducing emotion-related explanatory features into the data set, the large model can output reasonable explanations in emotional judgment. Each data sample pair includes audio, interpretable output (emotional clues), and emotional category labels. This data set design not only provides rich training data for the emotion recognition task, but also provides a framework for the large model to output explanations, so that it can give the reasoning process while outputting the category label, thereby helping application scenarios such as psychological counseling to more clearly understand the specific state and changing trend of user emotions.

[0032] S104: Use the data in the data set to perform model training to obtain an explainable emotion recognition model.

[0033] For example, Figure 4 As shown in the figure, the obtained acoustic feature set (such as pitch, loudness, etc.) is added to Prompt to ensure that the large model can better understand the emotional information behind the acoustic signal when analyzing audio emotions, rather than simply giving the emotional classification results; and the designed Prompt adds the chain of thought method (CoT) to obtain emotional clues, so that the model can explain the emotional judgment process in detail, output the emotional category and its reasoning process, and further manually correct the obtained emotional clues, and finally obtain the interpretable emotional clues corresponding to the speech. Instruct LLM to include not only the emotional category but also the corresponding emotional analysis description when outputting, that is, the interpretable emotional clues. These explanatory acoustic features can effectively reflect the expression of emotions and make the decision-making process of the model more transparent.

[0034] S105. Perform audio emotion recognition through the explainable emotion recognition model to obtain a complete emotion analysis result.

[0035] Exemplarily, for emotion recognition tasks, a fine-tuned large model can be used to replace the traditional emotion recognition model to achieve high-precision and explainable output of audio emotion recognition. The fine-tuned large model combines speech features and explainable feature descriptions, so it can not only accurately judge the emotion category in emotion recognition tasks, but also generate explanatory instructions at the time of output, making the emotion recognition process more transparent. It is no longer a simple category judgment, but outputs complete emotion analysis results in an explainable reasoning manner. This can provide more reliable and transparent emotion recognition services in application scenarios that require a deep understanding of emotions, such as psychological counseling, emotion tracking, and human-computer interaction.

[0036] In the interpretable emotion recognition method based on a large model of the present invention, a multi-level feature extraction mechanism is introduced, which combines with basic acoustic features to enable the model to more deeply understand complex emotion information, enhance the overall emotion recognition effect, and improve the recognition accuracy; by guiding the large model to generate emotion understanding descriptions, emotion recognition is not limited to simple classification results, but also includes a detailed reasoning process, which facilitates psychological counselors to understand the emotional state of visitors; in addition, the thinking chain method is adopted to enhance the logical coherence of emotion analysis, and the output is not only an emotion label, but also includes a reasoning process, thereby significantly improving the interpretability of the model.

[0037] In another implementation of the present invention, the basic acoustic features include prosody features, sound quality features and spectrum features.

[0038] For example, in the task of speech emotion recognition, feature extraction and selection are crucial. The standard features in speech emotion recognition can be divided into three categories: prosodic features, sound quality features, and spectral features. Acoustic features related to emotion include multiple dimensions such as speech rate, pitch, and pause duration. Existing large models often find it difficult to fully extract and understand these complex acoustic emotion features. To solve this problem, the present invention proposes a multi-level feature extraction method that combines a large model with interpretable acoustic emotion factors, so that the model can understand the emotional information extracted from the audio, improve the large model's ability to understand audio emotions, and thereby improve the accuracy of emotion recognition.

[0039] By combining basic acoustic features with thinking chains, the model can have a deeper understanding of complex acoustic emotional information, and use interpretable methods to enhance the large model's ability to understand acoustic emotional features; by fine-tuning the open source large model, it can understand and process audio emotional information, improve the accuracy and interpretability of emotion recognition, and is suitable for application scenarios such as psychological counseling and emotion recognition.

[0040] In another implementation of the present invention, the prosodic features include duration-related features, fundamental frequency-related features and energy-related features, etc.; the duration-related features include speech rate, short-time average zero-crossing rate, etc.; the fundamental frequency-related features include fundamental frequency and its mean, variation range, variation rate, mean square error, etc.; the energy-related features include short-time average energy, short-time average amplitude, etc.

[0041] For example, prosodic features are a common feature type in speech recognition, which involve changes in rhythm, intonation, amplitude, and frequency in a sentence.

[0042] In another implementation of the present invention, the sound quality features include jitter, shimmer, harmonic-to-noise ratio (HNR), etc.

[0043] For example, voice quality features are closely related to emotional states and have also been proven to be very effective in speech emotion recognition.

[0044] In another implementation of the present invention, the spectral features are a reflection of the correlation between vocal tract shape changes and vocalization movements, including linear prediction cepstral coefficients (LPCC), frequency cepstral coefficients (MFCC), and the like.

[0045] For example, spectral features are often used to supplement acoustic feature analysis, and different emotions are expressed differently in spectral morphology. Important spectral features commonly used in emotion recognition include: Mel frequency cepstral coefficients (MFCC), linear prediction coefficients (LPC), linear prediction cepstral coefficients (LPCC), gammatone frequency cepstral coefficients (GFCC), perceptual linear prediction (PLP) and formants.

[0046] In another implementation of the present invention, the data in the data set is used to perform model training to obtain an interpretable emotion recognition model, including: inputting the explanatory acoustic features in the data set into a pre-trained model to obtain a continuous high-dimensional acoustic feature encoding; decomposing the continuous high-dimensional acoustic feature encoding into a series of discrete speech units; using the speech units as input and the emotion labels and interpretable emotion clues as outputs to perform model training to obtain an interpretable emotion recognition model.

[0047] For example, open source large models, such as GPT4 and Claude, are used to optimize audio emotion recognition tasks. In order to improve the performance of the model in emotion recognition and ensure that its output is highly interpretable, interpretable emotion clues and voice tokens are introduced during the model training process for fine-tuning, so that the large model can effectively understand and process audio data.

[0048] Specifically, Figure 5 As shown in the figure, in the data preprocessing stage, the acoustic features are first extracted and tokenized. The acoustic features generate a set of high-dimensional embeddings through the aforementioned pre-training model, and then further converted into a set of speech tokens. These tokens not only reflect the acoustic information of the audio, but also have the temporal structure required for emotion recognition, providing rich feature input for the subsequent training process.

[0049] During the model fine-tuning process, the LLM is trained with interpretable emotion cues and speech tokens as inputs and emotion labels and interpretable emotion cues as outputs. This ensures that the LLM not only includes emotion categories during reasoning, but also generates reasonable emotion analysis to enhance the transparency of its decision-making. Finally, the fine-tuned large model can achieve high accuracy and interpretable output for audio emotion recognition.

[0050] It should be understood that although the traditional emotion recognition model can output emotion categories, it cannot explain its decision-making process. To this end, the present invention uses the powerful natural language processing capabilities of large models (such as GPT-4 or LLaMA3) to design appropriate prompts to guide the large model to perform emotion analysis in a human-understandable way and generate corresponding emotion understanding descriptions; and combines the Chain of Thought (CoT) method to explain the emotion judgment process in detail, thereby improving the logical coherence of the model in emotion analysis. The model not only outputs emotion categories, but also generates emotion categories and detailed reasoning processes, achieving transparency in the emotion recognition decision-making process, solving the "black box" problem of emotion recognition, and improving the interpretability of emotion recognition, helping psychological counselors understand and track the emotional changes of visitors, and providing well-founded emotional descriptions to assist personalized diagnosis.

[0051] In another implementation of the present invention, the loss function of the interpretable emotion recognition model includes the cross entropy loss of emotion classification and the similarity loss of interpretability clues.

[0052] Exemplarily, during the training process, the model's ability to understand and recognize emotions is enhanced by jointly optimizing the loss function, so that the model can effectively balance emotion classification and explanatory output. The loss function includes the cross-entropy loss of emotion classification and the similarity loss of interpretable cues. The latter is used to evaluate the similarity between the generated description and the expected explanation, so as to improve the interpretability of the sentiment analysis results. Through such a dual loss optimization strategy, the model gradually learns how to generate explanatory emotional output based on speech tokens. The model generated by the audio large model training module can accept audio token input and output category labels and corresponding explanatory descriptions during emotion classification, which provides higher accuracy and transparency for emotion recognition tasks and meets the high requirements for sentiment analysis in application scenarios such as psychological counseling.

[0053] The present invention aims to achieve transparency and accuracy in the emotion recognition process and enhance its practical value in applications such as psychological counseling.

[0054] In another aspect of the present invention, a large model-based interpretable emotion recognition device is provided, such as Figure 2 As shown, including:

[0055] Feature extraction module: extracts a variety of basic acoustic features for emotion classification from the original speech data.

[0056] Feature processing module: performs feature screening processing on the basic acoustic features to obtain explanatory acoustic features that are closely related to specific emotions.

[0057] Data preparation module: construct a data set based on the specific emotion, the corresponding explanatory acoustic features and the emotion category label.

[0058] Model training module: using the data in the dataset to perform model training to obtain an interpretable emotion recognition model.

[0059] Emotion recognition module: audio emotion recognition is performed through the explainable emotion recognition model to obtain complete emotion analysis results.

[0060] In the interpretable emotion recognition device based on a large model of the present invention, a multi-level feature extraction mechanism is introduced, which combines with basic acoustic features to enable the model to more deeply understand complex emotion information, enhance the overall emotion recognition effect, and improve the recognition accuracy; by guiding the large model to generate emotion understanding descriptions, emotion recognition is not limited to simple classification results, but also includes a detailed reasoning process, which facilitates psychological counselors to understand the emotional state of visitors; in addition, the thinking chain method is adopted to enhance the logical coherence of emotion analysis, and the output is not only an emotion label, but also includes a reasoning process, thereby significantly improving the interpretability of the model.

[0061] The present invention further provides an electronic device, which may include: a processor, a memory, a communication bus, and a communication interface.

[0062] in:

[0063] The processor, memory and communication interface communicate with each other through a communication bus.

[0064] Communication interface, used to communicate with other electronic devices or servers.

[0065] The processor is used to execute the program, and specifically can execute the steps of any one of the large model-based interpretable emotion recognition methods in the above-mentioned embodiments.

[0066] Specifically, the program may include program codes including computer operation instructions.

[0067] The processor may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0068] The memory is used to store programs. The memory may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0069] The program can be specifically used to enable the processor to execute to implement the steps of any one of the interpretable emotion recognition methods based on a large model described in the embodiment. The specific implementation of each step in the program can refer to the corresponding descriptions in the steps and units executed by any one of the interpretable emotion recognition methods based on a large model in the above steps, which will not be repeated here. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the aforementioned method embodiment.

[0070] The present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the methods of the various embodiments of the present application.

[0071] The above-described method according to an embodiment of the present invention may be implemented in hardware, firmware, or as software or computer code that may be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded over a network and will be stored in a local recording medium, so that the method described herein may be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, processor, or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0072] Thus far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired results. Additionally, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing may be advantageous.

[0073] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0074] In the description of the present invention, the terms "first" and "second" are only used to facilitate the description of different components or names, and cannot be understood as indicating or implying a sequential relationship, relative importance, or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features.

[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0076] It should be noted that although the specific embodiments of the present invention are described in detail in conjunction with the accompanying drawings, it should not be understood as limiting the scope of protection of the present invention. Within the scope described in the claims, various modifications and variations that can be made by those skilled in the art without creative work still belong to the scope of protection of the present invention.

[0077] The examples of the embodiments of the present invention are intended to concisely illustrate the technical features of the embodiments of the present invention so that those skilled in the art can intuitively understand the technical features of the embodiments of the present invention, and are not intended to be improper limitations of the embodiments of the present invention.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large model-based interpretable emotion recognition method, characterized in that: include: Extract a variety of basic acoustic features for emotion classification from raw speech data; Performing feature screening processing on the basic acoustic features to obtain explanatory acoustic features closely related to specific emotions; Constructing a data set based on the specific emotion, the corresponding explanatory acoustic features and the emotion category label; Using the data in the data set to perform model training, to obtain an interpretable emotion recognition model; Audio emotion recognition is performed through the interpretable emotion recognition model to obtain a complete emotion analysis result.

2. The method according to claim 1, characterized in that: The basic acoustic features include rhythmic features, sound quality features and spectrum features.

3. The method according to claim 2, characterized in that The prosody features include duration-related features, fundamental frequency-related features and energy-related features; The duration-related features include speech rate and short-term average zero-crossing rate; The fundamental frequency related features include fundamental frequency and its mean, variation range, variation rate, and mean square error; The energy-related characteristics include short-time average energy and short-time average amplitude.

4. The method according to claim 2, characterized in that: The sound quality characteristics include time base error, amplitude perturbation, and harmonic-to-noise ratio.

5. The method according to claim 2, characterized in that: The frequency spectrum features reflect the correlation between vocal tract shape changes and vocalization movements, and include linear prediction cepstral coefficients and frequency cepstral coefficients.

6. The method according to claim 1, characterized in that The method of using the data in the data set to perform model training to obtain an interpretable emotion recognition model includes: Inputting the explanatory acoustic features in the data set into a pre-trained model to obtain continuous high-dimensional acoustic feature encoding; Decomposing the continuous high-dimensional acoustic feature code into a series of discrete speech units; The speech unit is taken as input, and the emotion label and interpretable emotion clues are taken as output to perform model training, thereby obtaining an interpretable emotion recognition model.

7. The method according to claim 6, characterized in that The loss function of the explainable emotion recognition model includes the cross entropy loss of emotion classification and the similarity loss of interpretability cues.

8. An interpretable emotion recognition device based on a large model, characterized in that: include: Feature extraction module: extracts a variety of basic acoustic features for emotion classification from the original speech data; Feature processing module: performing feature screening processing on the basic acoustic features to obtain explanatory acoustic features closely related to specific emotions; Data preparation module: constructing a data set based on the specific emotion, the corresponding explanatory acoustic features and the emotion category label; Model training module: using the data in the data set to perform model training to obtain an interpretable emotion recognition model; Emotion recognition module: audio emotion recognition is performed through the explainable emotion recognition model to obtain complete emotion analysis results.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of an interpretable emotion recognition method based on a large model as described in any one of claims 1 to 7 are implemented.

10. A computer storage medium, characterized in that: The computer storage medium stores a computer program, which, when executed by a processor, implements the steps in a large model-based interpretable emotion recognition method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Remote Chinese language teaching system based on voice affection identification

    CN101201980A

  • Multi-modal depression detection method and system based on context awareness

    CN110728997A

  • Speech recognition system for cognitive impairment

    CN112908317A

  • Voice emotion recognition method and device based on artificial intelligence, equipment and medium

    CN115312033A

  • Speech emotion recognition method based on emotion embedding and feature fusion

    CN115881162A