Deep learning-based visual emotion expression system
By integrating computer vision and emotion recognition algorithms for data collection, labeling, feature extraction, and visualization, the method addresses the limitations of existing emotion analysis technologies, achieving improved accuracy and diversity in emotion expression.
Patent Information
- Application Number
- PCT/KR2024/012524
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-08-22
- Publication Date
- 2025-05-30
AI Technical Summary
Existing emotion analysis and visual emotion expression technologies face limitations in accurately analyzing and expressing emotions, particularly in automated systems, and struggle to capture and visually represent a diverse range of emotions.
The method combines advanced computer vision technology and emotion recognition algorithms to accurately identify and visually express various emotions through a process involving data collection, labeling, feature extraction, emotion classification model creation, emotion visualization, and text-based emotion expression.
This approach enhances the accuracy of emotion analysis and increases the diversity of visual emotion expressions, improving user experience and supporting emotion-based decision-making across various application fields.
Smart Images

Figure KR2024012524_30052025_PF_FP_ABST
Abstract
Description
Deep Learning-Based Visual Emotion Expression System
[0001] The present invention relates to a method for visual emotional expression through emotional analysis. This technology improves emotional recognition and expression, and is intended for use in various fields. For example, it can be used in various fields, such as emotional analysis software, human-machine interaction, marketing, and education.
[0002] In modern society, emotion analysis and emotional expression play a crucial role in various fields. These include emotion-based marketing, user experience improvement, human-machine interaction, education, and entertainment. Emotion analysis and visual emotional expression are crucial tools for understanding and interacting with human emotions in these fields.
[0003] However, existing emotion analysis and visual emotion expression technologies have several limitations. Existing technologies struggle to accurately analyze and express emotions, or they use limited methods, making it difficult to recognize and express emotions in automated systems. Furthermore, existing technologies have limitations in capturing and visually expressing diverse emotions, limiting their use in various applications.
[0004] The present invention was developed to overcome these problems and limitations. This technology innovatively combines advanced computer vision technology with emotion recognition algorithms to provide a method for accurately identifying and visually expressing a variety of emotions.
[0005] This invention can be used in a variety of applications, including sentiment analysis software, human-machine interaction systems, marketing campaigns, educational content, and entertainment, to enhance user experience and support emotion-based decision-making.
[0006] These technological innovations will offer superior performance and versatility compared to existing technologies, opening up new possibilities and opportunities in the fields of emotion recognition and expression.
[0007] The purpose of the present invention is to solve the problem of accuracy and lack of diversity in visual emotional expression that arise in existing emotional analysis technology.
[0008] The emotional analysis visual emotion expression method of the present invention for solving the above-described problem is characterized by comprising a data collection step, a labeling step which is a process of assigning an emotional category to audio data, a feature extraction step which is a process of extracting features from audio data, an emotional classification model creation step which is a process of creating a model that recognizes and classifies emotions based on the extracted audio features, an emotion visualization step which is a process of visually expressing emotions by considering the model prediction result and accuracy, and a step of visualizing and expressing emotions in text based on the emotional intensity.
[0009] The present invention improves the accuracy of sentiment analysis, enabling more accurate sentiment analysis than existing methods, thereby providing users with more accurate information. Furthermore, it increases the diversity of visual emotional expressions, enabling more effective visual expression of emotions, making information easier and more intuitive for users to understand. Furthermore, it can be used in a variety of applications, including sentiment analysis software, human-machine interaction systems, marketing campaigns, educational content, and entertainment, thereby enhancing the user experience and supporting emotion-based decision-making. Furthermore, compared to existing sentiment analysis and visual emotion expression technologies, it offers superior performance and diversity, opening up new possibilities and opportunities in the fields of emotion recognition and expression.
[0010] Figure 1 illustrates an overall method for visual emotional expression through emotional analysis using data of the present invention.
[0011] Figure 2 illustrates a model performance table of the present invention.
[0012] The present invention will be described in detail with reference to certain drawings. Reference numerals are added to components in each drawing, providing a detailed description of the components.
[0013] From now on, with reference to the attached drawings, the emotional analysis visual emotional expression method of the present invention will be described in detail.
[0014] Figure 1 illustrates an overall visual emotion expression method for emotion analysis using data of the present invention, and Figure 2 illustrates a model performance table of the present invention.
[0015] Figure 1 is a method for solving the problem of lack of diversity in visual emotional expression. The method for analyzing emotions through data and expressing emotions visually includes a data collection step (v100), a labeling step (v110), a feature extraction step (v120), an emotional classification model creation step (v130), an emotional visualization step (v140), and an emotional visualization and expression step (v150).
[0016] The data collection phase (v100) represents the step of collecting voice and emotion-related information from various data sources.
[0017] The labeling stage (v110) represents the process of assigning emotional categories to the collected data. For labeling, people listen to voice samples and label them with the corresponding emotions.
[0018] The feature extraction step (v120) represents the process of extracting acoustic features from voice data, and extracts information such as frequency, volume, and voice features and expresses them as numbers.
[0019] The emotion classification model creation step (v130) describes the process of developing and training an emotion classification model. Using extracted acoustic features, the model recognizes and classifies emotions.
[0020] The emotion visualization step (v140) shows the process of visually expressing emotions by considering the model's prediction results and accuracy.
[0021] The Emotion Visualization and Expression step (v150) describes the process of visualizing emotions in text based on their intensity. This is a crucial step in clearly communicating the results to users and facilitating their understanding.
[0022]
[0023] Data Preprocessing
[0024] Data preprocessing refers to the process of processing data into a format easily usable by deep learning models. This process transforms raw data into a format that deep learning models can appropriately process. It is a crucial step for maintaining data quality and transforming it into an analyzable form. Data preprocessing ensures that data is input into a format more suitable for the model, enabling the model to learn more effectively and make more accurate predictions on new data.
[0025]
[0026] Sentimentrogram model
[0027] [Correction under Rule 91 11.09.2024] Figure 4 is the structure of the Sentimentrogram model.
[0028] [Correction pursuant to Rule 91, September 11, 2024]
[0029] Voice emotion recognition, which recognizes emotions from audio data, is largely composed of five parts: 'Raw Audio Data Labeling', 'Audio Feature Extraction', 'Audio Model Creation', 'Emotion Visualization and Emotion Intensity Mapping', and 'Emotion visualization and size representation in text based on intensity mapping'.
[0030]
[0031] The "Raw Audio Data Labeling" step defines labels for the given audio data. First, initial setup for the experiment is performed, defining the parameters necessary for data labeling and audio feature extraction. For example, this setup includes the file name for storing the labeled data, the file name for storing the audio features, and the name of the data source. The data for the experiment is then loaded and used in the experiment. In the present invention, data containing labels is loaded from a CSV file. Otherwise, the task of identifying and labeling audio files from the data source is performed.
[0032] The "Audio Feature Extraction" step defines settings for extracting audio features using the Librosa library. These settings include a list of audio features to be extracted and parameter settings for each feature. MFCC (Mel-frequency cepstral coefficients) are primarily used, generating features containing frequency and time information.
[0033] For the "MFCC" feature, the 'n_mfcc (Mel-frequency cepstral coefficients)' and 'sr (sample rate)' values are set. Therefore, the number of n_mfcc is 12, the sampling frequency is 48000, and the dictionary is specified, and the number of cores to be extracted is 4 for parallel processing. The existence of the extracted audio feature data is checked, and if the data already exists, the data is loaded and used for the experiment.
[0034] Otherwise, audio features are extracted based on the data labeled in the previous step.
[0035] The "Audio Model Creation" step creates a model that recognizes and classifies emotions based on extracted audio features. The emotion classification model is based on a 1D Convolusional Neural Network (1DCNN). The model uses MFCC (Melfrequency cepstral coefficients) as input, which captures frequency and temporal information. The model consists of the following layers.
[0036] [Correction under Rule 91 11.09.2024] Figure 5 is a diagram showing the Sentimentrogram 1DCNN structure.
[0037] [Correction pursuant to Rule 91, September 11, 2024]
[0038] Conv1D Layer: This layer applies convolution to the input data to extract temporal features. Multiple Conv1D layers are stacked, each followed by a ReLU activation function and a Dropout layer. The Conv1D layer uses 256 filters and 5 kernel sizes. The 'Padding' option is set to 'same' to keep the input and output sizes identical.
[0039] Flatten Layer: Flattens the extracted features into a one-dimensional vector.
[0040] Fully Connected Layer: Performs multi-class classification and outputs predicted probabilities for each class using the softmax activation function.
[0041] The model uses the Adam optimizer to optimize the categorical cross-entropy loss function, and the model's performance is evaluated by accuracy, which represents the proportion of correctly classified predictions.
[0042] In the "Emotion Visualization and Emotion Intensity Mapping" step, emotion labels were mapped to RGB values based on the colors of the PAD theory described in Section 3.3.2 to facilitate intuitive understanding of the emotion analysis results. Emotional intensity was estimated based on the predicted results and model accuracy, and font size was expressed differently depending on the emotional intensity. Equations (1, 2) represent functions that estimate emotional intensity and adjust font size accordingly.
[0043] In addition, the Zero Crossing Rate (ZCR) can be used as a function expressing the above emotions, which is the number of times the signal value changes from positive to negative or negative to positive divided by the frame length (WL), and is expressed in the following equation (3).
[0044] (Formula 1)
[0045]
[0046] (Formula 2)
[0047]
[0048] (Formula 3)
[0049]
[0050] [Revised 11.09.2024 under Rule 91] This paper describes a method for extracting emotional elements from a melspectrogram, the subject of the present invention, converting it into STT, and expressing the emotional elements in text. The proposed model operates according to the flow shown in Figure 6. First, preprocessing is performed to enable the voice data to be utilized by a deep learning model. Then, various emotion recognition models are utilized to perform emotion analysis, and based on this, emotional information is assigned to the text and visually expressed in color and size.
[0051]
[0052] [Correction under Rule 91 11.09.2024] Figure 6 is an example of a model operation according to the present invention.
[0053] [Correction pursuant to Rule 91, September 11, 2024]
[0054] Data preprocessing in the present invention involves transforming data to fit the input of the emotion expression model. The sentimentrogram model can identify audio files (".wav") and perform sentiment analysis using labeled CSV files, regardless of the data source. Therefore, the performance of the present invention was demonstrated using three data sets: Emo_DB, RAVDESS, and LDC Emotional Speech and Transcripts.
[0055] The initial 'LDC Emotional Speech and Transcripts' speech data is provided in SPH file format, a dataset in which eight voice actors pronounce several words each expressing 17 emotions.
[0056] To convert this data into a Melspectrogram, the SPH file format must be converted to a wav file format.
[0057] This wav file cannot be converted to a Melspectrogram because it contains a single voice actor pronouncing hundreds of words continuously for several tens of minutes with various emotions.
[0058] Therefore, in order to distinguish each utterance, all wav files were cut according to time zone and all file names were changed.
[0059] And when a wav file and a corresponding text format transcript file are input, the (start) time and (end) time in the first and second columns of the transcript file are automatically extracted and the file is divided by utterance.
[0060] At this time, the format of the transcripts is, for example, as follows: 43.26 43.28 A: neutral, Two thousand one. Here, the first two numbers represent the interval (in seconds), the "neutral" that follows indicates the emotional state, and the part after that contains the content of the speech.
[0061] In this process, all sentences with errors and parts that did not correspond to the utterance were deleted.
[0062] Next, the names of the data were encoded to train the Semtimentrogram model. The name format of each file is as follows: [Voice actor number: 1 digit][Emotion: 2-digit uppercase alphabet initial][Conversation situation: 1 digit][Speech content: s + 3 digits][Duplicate numbering: 1 digit].wav. Here, the voice actors were numbered 1 to 8, consisting of 3 men and 5 women, and only 7 emotions were left: AX (anxiety), SN (sadness), NT (neutral), HP (happy), HA (hot anger), DG (disgust), and BD (boredom). In this process, in order to increase the accuracy of classification, emotions that are not commonly used, such as dominant and contempt, were deleted, and only 7 emotions were used out of 17 emotions.
[0063] The conversation situations were numbered into three categories: "conversation," "distant," and "tete a tete."
[0064] Here, conversation is a conversational tone, distant is a loud voice speaking to a distant object, and tete a tete is a whispering voice. The utterances are s001, s002,... approximately 400. If there are duplicate file names, a 0 is added to the end of the file name.
[0065] Therefore, the file name is expressed as follows: '1AX1s0010.wav'. Then, a folder is collected with each file preceded by a voice actor number and the csv file is extracted.
[0066] This CSV file is created by modifying the information in columns such as "emotion," "statements," "gender," and "situation" in a dictionary-type JSON file. [Table 1] shows the format of the CSV file.
[0067] [Table 1] Format of csv file
[0068]
[0069] Figure 2 illustrates the model performance of the present invention and visualizes that this model can be used to evaluate the effectiveness of sentiment analysis visual representations.
[0070] This figure includes model accuracy, model precision, model F-score, and model recall, and the equations for each metric are presented below.
[0071] Figure 3 is an explanation of the contents presented as an SMI file utilizing an embodiment of implementing the above emotional information of the present invention, as follows.
[0072] For example, between the times 00:00:11,951 and 00:00:15,573 in the video, let's look at the 4th line. The 'Ah yeah!' part is the emotion of happiness (Happy), and if this is expressed in yellow RGB, it becomes #D2FF0A, and the intensity of the emotion can be 40.
[0073] In the case of the next line, 'See', the emotion is disgust and the intensity is 30. Here is a detailed example of the text expression of that emotion.
[0074] <font color="#D2FF0A" font size="40">Ah yeah! < / font> <font color="#0070c0" font size="30"> See < / font> <font color="#D2FF0A" font size="40"> games < / font> <font color="#0070c0" font size="25">are my <font color="#0070c0" font size="30"> thing < / font> <font color="#D2FF0A" font size="40"> you know. < / font>
[0075] For the following dialogue, for the video between the times 00:00:28,099 and 00:00:30,000, dialogue 9 is 'What's wrong with you' and this is the anger emotion (Angry) in red RGB #FF0000 and the size of the emotion is 40, and the following dialogue 'with you You' is the happiness emotion (Happy) in yellow RGB #D2FF0A and the size of the emotion is 30, and for the following dialogue freaking pervert?, if the anger emotion's RGB #FF0000 and the size of the emotion etc. is 60, the above content can be expressed as follows in a file format such as SMI.
[0076] <font size="40" font color="#FF0000"> What's wrong < / font> <font size="30" font color="#D2FF0A">with you You <font size="40" font color="#FF0000"> freaking < / font> <font size="60" font color="#FF0000"> pervert?< / font>
[0077] Accordingly, the contents of the present invention are considered to be for the purpose of explaining rather than limiting the technical concept of the present invention, and there is no limitation on the scope of the technical concept of the present invention.
[0078] The scope of the present invention should be interpreted according to the following claims, and similar technical concepts will be considered to be included within the scope of the present invention.
[0079] < / font> < / font>
Claims
1. In an emotion expression system that expresses the emotion in the text using emotion data obtained by using STT (Speech to Text) that converts voice into text and a voice-based emotion recognition (SER) deep learning model, An emotion module that receives emotion elements such as the type and size of emotion as input in advance and has the function of managing the input emotion; and An emotion expression system comprising a text module that expresses the intensity of emotion provided by the above emotion module by changes in the visual size of the text.
2. In paragraph 1, An emotion expression system further comprising an emotion text style module that expresses the type of emotion by mapping the type of emotion provided by the emotion module to the style of the text (e.g., Gothic, round, italic, etc.).
3. In paragraph 1 or 2, An emotion expression system further comprising a text color module that expresses the type of emotion by mapping the type of emotion provided by the emotion module to the color of the text.
4. In paragraph 3, An emotion expression system further characterized by including a method of displaying a color similar to the background color of the text and the surrounding area in the above text color module.
5. In paragraph 1, An emotion expression system further characterized by including a method for expressing emotions by positioning text at an arbitrary location on the screen based on the above text data.
6. In paragraph 4, An emotion expression system further characterized by including a method of expressing emotions visually, including a method of expressing emotions by expressing the emotion of text as a label based on the above text data.
7. In paragraph 4, An emotion expression system further characterized by including a method of visually expressing emotions, including a method of expressing emotions by expressing the emotion of text as a legend based on the above text data.
8. In paragraph 4, An emotion expression system further characterized by including a method for expressing emotions using speech bubbles in text based on the above text data.
9. In paragraph 4, An emotion expression system further comprising a method for reproducing the voice including the emotion by inputting at least one of the size, color, and style of the text to the emotion recognition deep learning model, etc.
Citation Information
Patent Citations
Method and system for analysing feeling
KR1020130010134A
Language interpreter, speech synthesis server, speech recognition server, alarm device, lecture local server, and voice call support application for deaf auxiliaries based on the local area wireless communication network
KR1020160142079A
A method and apparatus for determining evaluation technology for cardiac respiratory physiotherapy of patient
KR1020210132393A
Carbon dioxide capture composite particles and manufacturing method thereof
KR1020240065887A
Method and system for generating caption
KR102067446B1