Ai stress feedback device having a voice pattern recognition capability

KR103013350B1Active Publication Date: 2026-09-02IND UNIV COOP FOUND SUNMOON UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020240010246
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2026-09-02
Estimated Expiration
2044-01-23

Smart Images

  • Figure R1020240010246_ABST
    Figure R1020240010246_ABST
Patent Text Reader

Abstract

The present invention discloses a device for analyzing stress and providing feedback by determining the user's emotions based on voice pattern recognition. The device according to the present invention is characterized by comprising: an acoustic analysis module that extracts pronunciation information corresponding to phonemes or phonological information from an input voice signal; a pronunciation analysis module that performs voice recognition by considering user-specific intonation, accent, and pronunciation habits in the phonological information; a language modeling module that performs appropriate voice recognition for the voice signal from the acoustic analysis module; a sentiment analysis module that performs deep learning to predict sentiment and determines positive, negative, and neutrality regarding contextual sentiment; and a text conversion module that converts the context of the voice determined by the sentiment analysis module into text and derives the sentiment as an emoticon. Therefore, the present invention has the effect of deriving a result corresponding to the user's emotions by determining the emotion of the voice based on AI voice recognition. In addition, the present invention provides the effect of precisely analyzing an individual's emotions by recognizing voice while considering the intonation, accent, and pronunciation habits of the user's pronunciation.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to an artificial intelligence feedback device, and more specifically, to an artificial intelligence stress feedback device having a voice pattern recognition function capable of recognizing a user's stress state by precisely determining the user's voice pattern. Background Technology

[0003] Generally, speech recognition using AI is an algorithmic technology that classifies and learns the characteristics of input data on its own, and the elemental technology is a technology that mimics the functions of the human brain, such as cognition and judgment, by utilizing machine learning algorithms such as deep learning, and consists of technology fields such as linguistic understanding, visual understanding, reasoning / prediction, knowledge representation, and motion control.

[0004] Conventionally, speech processing is performed into natural sentences by implementing an AI-based speech processing method that can determine speech intent for incomplete utterances, implementing an AI-based speech processing method that can determine speech intent by merging two or more utterances with different pronunciation sequences, or determining speech intent by merging two or more utterances with various patterns.

[0005] The attached patent document implements such voice processing functions, and as can be seen in FIG. 1, various components are required to process voice events in an end-to-end voice UI environment. The sequence for processing voice events performs signal acquisition and playback, speech pre-processing, voice activation, speech recognition, natural language processing, and finally, speech synthesis in which the device responds to the user.

[0006] The client device (50) includes an input module. The input module can receive user input from a user. For example, the input module receives user input from a connected external device (e.g., a keyboard, a headset). Also, for example, the input module includes a touch screen, and the input module includes a hardware key located on the user terminal.

[0007] The input module includes at least one microphone capable of receiving a user's speech as a voice signal, and the input module includes a speech input system and receives the user's speech as a voice signal through the speech input system. The at least one microphone determines a digital input signal for the user's speech by generating an input signal for audio input. That is, a plurality of microphones may be implemented as an array. The array may be arranged in a geometric pattern, for example, a linear geometric shape, a circular geometric shape, or any other configuration. For example, at a given point, an array of four sensors is arranged in a circular pattern divided by 90 degrees to receive sound from four directions.

[0008] In some implementations, the microphone may include spatially different arrays of sensors within the data communication, and may include a networked array of sensors. The microphone may include omnidirectional, directional (shotgun microphone), etc.

[0009] The client device (50) includes a pre-processing module (51) capable of pre-processing user input (voice signal) received through the input module (e.g., microphone). The pre-processing module (51) can remove echoes included in the user voice signal input through the microphone by including an adaptive echo canceller (AEC) function. The pre-processing module (51) removes background noise included in the user input by including a noise suppression (NS) function.

[0010] A client device (50) can transmit user voice input to a cloud server. Automatic speech recognition (ASR) and natural language understanding (NLU) operations, which are core components for processing user voice, are typically executed in the cloud due to computing, storage, and power constraints. The cloud includes a cloud device (60) that processes user input transmitted from the client.

[0011] The cloud device (60) includes an Auto Speech Recognition (ASR) module (61), an Artificial Intelligent Agent (62), a Natural Language Understanding (NLU) module (63), a Text-to-Speech (TTS) module (64), and a service manager (65). The ASR module (61) converts user voice input received from the client device (50) into text data.

[0012] When the ASR module (61) generates a recognition result containing a text string (e.g., words, or a sequence of words, or a sequence of tokens), the recognition result is transmitted to the natural language processing module (732) for intent inference. The ASR module (730) generates multiple candidate text representations of the speech input. Each candidate text representation is a sequence of words or tokens corresponding to the speech input. In this way, conventional speech pattern recognition determines speech intent by merging two or more utterances with different pronunciation sequences, and by merging two or more different utterances of various patterns to determine speech intent, speech processing into natural sentences is possible.

[0013] However, conventional speech processing methods aim for speech recognition of utterances, making it difficult to determine the user's sentiment contained in the speech. In other words, speech analysis-based speech processing only analyzes spoken words, so there is a problem in that it cannot determine the sentiment of the speech contained in the words. Prior art literature

[0015] Republic of Korea Published Patent No. 10-2021-0062838, Title of Invention 'Artificial Intelligence-based Voice Management Method' The problem to be solved

[0016] The present invention was created to solve such problems, and the objective of the present invention is to provide an artificial intelligence stress feedback device having a voice pattern recognition function capable of deriving a result corresponding to the user's emotion by determining the emotion of the voice based on AI voice recognition.

[0017] Another objective of the present invention is to provide an artificial intelligence stress feedback device having a voice pattern recognition function capable of precisely analyzing an individual's emotions by recognizing speech while considering intonation, accent, and pronunciation habits of the user's pronunciation. means of solving the problem

[0019] An artificial intelligence stress feedback device having a voice pattern recognition function according to an aspect of the present invention for achieving the above objective is an artificial intelligence device capable of determining a user's stress state based on voice pattern recognition and providing corresponding feedback, comprising: an acoustic analysis module that extracts pronunciation information corresponding to phonemes or phonological information from an input voice signal; a pronunciation analysis module that determines pronunciation by comparing the voice signal and pronunciation information from the acoustic analysis module, and performs voice recognition by considering the user's intonation, accent, and pronunciation habits in the phonological information; a language modeling module that performs context-appropriate voice recognition by performing an N-gram model or a statistical language model on the voice signal from the acoustic analysis module; and a sentiment analysis module that determines the emotion embedded in a sentence by calculating similarity in a vector space while preserving the meaning and context of words through embedding that converts a sentence into a numerical vector and presents it as a high-dimensional vector based on the pronunciation analysis module and the language modeling module, and then determines the positive, negative, and neutrality of the contextual sentiment by receiving the embedded sentence vector as input and performing deep learning to predict sentiment. It is characterized by being composed of a text conversion module that converts the context of the voice determined by the sentiment analysis module into text and derives the sentiment as an emoticon.

[0020] In addition, the acoustic analysis module according to the present invention extracts feature information from sequence data of a speech signal by using a deep learning algorithm or a Hidden Markov Model (HMM) method, and the feature information is characterized by using a feature extraction algorithm of Mel Frequency Cepstral Coefficients (MFCC) that can extract features of frequency, energy, and frequency change from the speech signal.

[0021] In addition, the pronunciation analysis module according to the present invention is characterized by using a Deep Neural Network (DNN) to predict a phoneme sequence by receiving an MFCC feature vector extracted from the acoustic analysis module.

[0022] In addition, the language modeling module according to the present invention models the structure and probabilistic characteristics of language, and is characterized by the application of an N-gram model or a statistical language model to predict the probability of the next word or sentence in a given context.

[0023] In addition, the sentiment analysis module according to the present invention is characterized by being able to adjust the weights of the model in a direction that minimizes the difference between the predicted value and the actual sentiment label while training the model using training data, and by applying Gradient Descent, which performs optimization during the training process by defining a loss function.

[0024] In addition, the sentiment analysis procedure of the sentiment analysis module according to the present invention comprises: a data collection and preprocessing procedure for data processed by the model, which performs tasks such as tokenizing sentences from sentiment label data provided from the pronunciation analysis module and converting words into numbers; and a process of converting the sentences into numeric vectors. Embedding is characterized by comprising: an embedding procedure that represents words or sentences as high-dimensional vectors; a deep learning procedure that receives the embedded sentence vector as input and predicts sentiment; a training optimization procedure that trains and optimizes the deep learning model, which minimizes the difference between the predicted value and the actual sentiment label while training the model using training data; and a prediction procedure that, after embedding the input sentence, gives it as input to the deep learning model and classifies the sentiment of the sentence into the class with the highest probability among positive, negative, and neutral classes based on the predicted value. Effects of the invention

[0026] The artificial intelligence stress feedback device having a voice pattern recognition function presented in the present invention has the effect of deriving a result corresponding to the user's emotions by determining the emotion of the voice based on AI voice recognition. In addition, the present invention provides the effect of precisely analyzing an individual's emotions by recognizing voice while considering the intonation, accent, and pronunciation habits of the user's pronunciation. Brief explanation of the drawing

[0028] Figure 1 is a configuration diagram illustrating conventional speech recognition technology. FIG. 2 is a configuration diagram for explaining an artificial intelligence stress feedback device having a voice pattern recognition function according to the present invention. FIG. 3 is a diagram illustrating the acoustic analysis procedure of the acoustic analysis module according to the present invention. FIG. 4 is a diagram illustrating the operation procedure of a pronunciation analysis module according to the present invention. FIG. 5 is a diagram illustrating the operation procedure of a language modeling module according to the present invention. FIG. 6 is a diagram illustrating the operation procedure of the sentiment analysis module according to the present invention. Specific details for implementing the invention

[0029] This project (result) is the result of the Local Government-University Cooperation-based Regional Innovation Project, conducted in 2023 with funding from the Ministry of Education and support from the National Research Foundation of Korea (2021RIS-004).

[0030] Hereinafter, a preferred embodiment of the present invention will be described in detail with reference to the attached illustrative drawings as follows.

[0031] FIG. 2 is a configuration diagram for explaining an artificial intelligence stress feedback device having a voice pattern recognition function according to the present invention. As described above, the system comprises: an acoustic analysis module (201) that extracts pronunciation information corresponding to phonemes or phonological information from an input voice signal; a pronunciation analysis module (203) that determines pronunciation by comparing the voice signal and pronunciation information from the acoustic analysis module (201) and performs voice recognition by considering user-specific intonation, accent, and pronunciation habits in the phonological information; a language modeling module (205) that performs context-appropriate voice recognition by performing an N-gram model or a statistical language model on the voice signal from the acoustic analysis module (201); a sentiment analysis module (207) that determines the emotion embedded in a sentence by calculating similarity in a vector space while preserving the meaning and context of the word through embedding that converts the sentence into a numerical vector and presents it as a high-dimensional vector based on the pronunciation analysis module (203) and the language modeling module (205), and then determines the positive, negative, and neutrality of the contextual sentiment by receiving the embedded sentence vector as input and performing deep learning to predict the sentiment; and the sentiment analysis It consists of a text conversion module (209) that converts the context of the determined voice of the module (207) into text and derives the emotion as an emoticon.

[0032] The above acoustic analysis module (201) may use a deep learning algorithm, but the present invention applies a Hidden Markov Model (HMM) method that is efficient for acoustic recognition. The HMM algorithm is a statistical model that extracts meaningful feature information from various sequence data such as acoustic signals, speech, text, and gestures. The feature information is extracted from speech signals by extracting features such as frequency, energy, and frequency changes, and the Mel Frequency Cepstral Coefficients (MFCC) feature extraction algorithm is applied.

[0033] The above pronunciation analysis module (203) receives MFCC feature vectors extracted from the above acoustic analysis module (201) and uses a Deep Neural Network (DNN) to predict phoneme sequences. This is mainly used to improve the performance of speech recognition and is trained with a large dataset to analyze acoustic characteristics between phonemes.

[0034] The above language modeling module (205) models the structure and probabilistic characteristics of language and is a model that predicts the probability of the next word or sentence in a given context. This may use an N-gram model that predicts the probability of the next word when N-1 words are given in a commonly used sentence, or a statistical language model that uses statistical methods to predict the probability of the next word in a given context. As such, the N-gram model or the statistical language model is utilized in various natural language processing tasks, such as maintaining the naturalness of a sentence or predicting the next word.

[0035] Meanwhile, the sentiment analysis module (207) can adjust the weights of the model in a direction that minimizes the difference between the predicted value and the actual sentiment label while training the model using training data, and can apply Gradient Descent, which performs optimization during the training process by defining a loss function.

[0036] The text conversion module (209) uses a speech alignment algorithm that performs mapping between speech and text, and converts speech into text by comparing the pronunciation characteristics of the speech signal with the pronunciation characteristics of the text to find the temporal correspondence between the speech and the text, and then determining the text corresponding to each part of the speech signal.

[0037] The operation of the present invention will be described in detail below with reference to the attached exemplary drawings.

[0038] First, the voice signal input through the microphone is used to extract phoneme or phonological unit features from the voice signal through an acoustic analysis module (201) and to recognize the voice based on these features. An algorithm for this purpose is a Hidden Markov Model (HMM), which is a statistical model used to model and analyze sequence data and extracts meaningful information from various sequence data such as acoustic signals, voice, text, and gestures. As with the acoustic analysis procedure recognized in FIG. 3, state analysis is performed in step S310.

[0039] The above state analysis step represents various situations or conditions that the system may assume, such as a state of emitting sound and a state without sound. Subsequently, the system enters the S320 step to perform the observation step. The above observation step extracts observation symbols representing values ​​observed in each state. The above observation symbols are assumed to include the frequency spectrum of the acoustic signal for acoustic recognition, Mel frequency cepstrum (MFCC), filter bank energy, etc.

[0040] Entering step S330, the Transition Probability step is performed. This refers to the transition probability between states, representing the probability of transitioning from one state to another. In other words, it signifies the probability of transitioning from the current state to the next; the transition probability models the dependencies between states and reflects the characteristics of sequence data through state transitions.

[0041] Step S340 is the Emission Probability step, where the HMM has a probability that an observation symbol will be generated in each state, representing the probability that a specific observation symbol will be observed in a specific state. The emission probability models the relationship between state characteristics and observation symbols, and infers the state through the observation symbols. Then, entering Step S350, the Initial Probability step, the HMM analyzes the probability that each state will be selected first as the starting state based on the initial state probabilities. These initial probabilities represent the starting point of the HMM model and determine the first state of the sequence data.

[0042] Accordingly, the above HMM models sequence data based on state transition probabilities, observation probabilities, and initial probabilities, and to train the model, mapping between observation data and states is performed. In the present invention, the Baum-Welch algorithm or the Viterbi algorithm may be used for state mapping.

[0043] That is, the acoustic analysis module (201) extracts features such as frequency, energy, and frequency change from the speech signal, and for this purpose, a feature extraction algorithm such as Mel Frequency Cepstral Coefficients (MFCC) is used. Acoustic unit segmentation is performed by the MFCC algorithm, which divides the speech signal into acoustic units such as phonemes or words. Thus, phoneme or word boundaries are determined by utilizing the features of the speech signal and a language model.

[0044] The above pronunciation analysis module (203) receives voice input, separates it into phonemes, and analyzes the pronunciation characteristics of each phoneme. The pronunciation characteristics of the phonemes include pitch, length, intensity of pronunciation, etc. The above pronunciation analysis module (203) utilizes voice signal processing technology to convert the voice input into a digital signal and performs the task of analyzing it to divide it into phonemes.

[0045] In the present invention, since the acoustic analysis module (201) is a core element of speech recognition technology, it is important to train an accurate acoustic model, and for this purpose, a large amount of speech data and ground truth text for the data are required. Therefore, the performance of speech recognition technology can be improved by training the model using a machine learning algorithm such as deep learning.

[0046] Subsequently, the above-mentioned acoustic analysis module (201) extracts pronunciation information corresponding to phonemes or phonological information from the input speech signal based on state mapping, and the pronunciation information is provided to the above-mentioned pronunciation analysis module (203) and language modeling module (205).

[0047] The above pronunciation analysis module (203) accurately models the pronunciation of a specific language and is used to accurately understand and convert the user's speech. The above pronunciation analysis module (203) includes a process of training a pronunciation model using a pronunciation transcript.

[0048] That is, phoneme separation is performed in step S410, as shown in the operation procedure of the pronunciation analysis module illustrated in Fig. 4. A phoneme is the smallest phonological unit of language and is the basic unit of pronunciation, and since the pronunciation analysis module (203) is based on phonemes, the collected speech data is separated into phoneme units.

[0049] Next, the process enters step S420 to perform phoneme model training. This involves training a phoneme model using separated phoneme data, and the phoneme model is used to learn the relationship between phonemes and their corresponding pronunciations. In other words, it trains various phoneme combinations and their corresponding pronunciation variations. Subsequently, in step S430, a pronunciation model is constructed based on the phoneme model, and the pronunciation model performs the role of mapping phoneme sequences to their corresponding pronunciations.

[0050] The aforementioned phoneme sequences are aligned for pronunciation modeling; alignment is a process of mapping text and speech data to determine the accurate phoneme sequence for each word or sentence. Alignment is performed to maximize the correspondence between the phoneme dictionary and the speech data.

[0051] Accordingly, the pronunciation analysis module (203) is used to predict and convert pronunciation in a speech recognition system. Such pronunciation analysis is learned by taking into account diversity and variation, and the model is trained by taking into account the pronunciation rules of a specific language and contextual changes in pronunciation.

[0052] Meanwhile, the above language modeling module (205) is intended to process speech signals provided from the above acoustic analysis module (201) into natural language, predict the probability of the next word or sentence in a given context, and perform natural language understanding and generation tasks. FIG. 5 is a diagram illustrating the language modeling procedure.

[0053] As described, sentence-unit segmentation is performed in step S510. This involves dividing the collected text data into sentence units, using a sentence segmentation algorithm or based on sentence delimiters. Then, language model training is performed in step S520. This involves training a language model using the segmented sentence data to build a model that predicts the probability of the next word in a given context; for this purpose, a statistical-based approach or a deep learning-based approach may be used.

[0054] Entering step S530, language model evaluation takes place. Specifically, this involves evaluating the performance of the trained language model by predicting the actual next word within a given context and assessing the accuracy of that prediction. Perplexity may be used as an evaluation method.

[0055] Accordingly, the language modeling module (205) makes the translation result of a given sentence natural, or in speech recognition, uses a language model to select the sentence that matches most probabilistically when converting a speech signal into text.

[0056] The sentiment analysis module (207) analyzes the sentiment of sentences with sentiment labels such as positive, negative, and neutral generated from the pronunciation analysis module (203). This process follows the sentiment analysis procedure illustrated in the attached FIG. 6, and data collection and preprocessing are performed in step S610.

[0057] A large amount of training data is required to train the sentiment analysis, and this data includes sentences with sentiment labels such as positive, negative, and neutral, and is provided by the pronunciation analysis module (203). In the preprocessing step after data collection, the sentences are tokenized and words are converted into numbers, and the data is processed into a form that the model can process.

[0058] Then, at step S620, embedding is performed, which is a process of converting sentences into numeric vectors. Embedding is a technique that represents words or sentences as high-dimensional vectors, calculating similarity in a vector space while preserving the meaning and context of the words; Word2Vec, GloVe, FastText, etc., can be used. Entering step S630, deep learning is performed, and the deep learning receives the embedded sentence vectors as input to predict sentiment.

[0059] The deep learning described above utilizes Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTMs), and modified Transformers, and predicts sentiment by considering the context of the sentence and the interaction between words. In step S640, a learning optimization process is performed. That is, the deep learning model is trained and optimized; the model is trained using training data, and the model's weights are adjusted in a direction that minimizes the difference between the predicted value and the actual sentiment label. As previously mentioned, the Gradient Descent optimization algorithm is applied to perform optimization. The optimization algorithm defines a loss function to perform optimization during the training process.

[0060] In step S650, the sentiment of the sentence is predicted. The input sentence is embedded and then input to a deep learning model to output a predicted value. The predicted value is usually represented as a probability value, and the sentiment of the sentence is classified into the class with the highest probability among classes such as positive, negative, and neutral.

[0061] Accordingly, natural language processing data generated from the language modeling module (205) and sentence sentiment prediction data predicted by the sentiment analysis module (207) are generated, and the context of the voice is converted into text through the text conversion module (209), and the sentiment is derived as an emoticon to express the sentiment contained in the user's voice. Ultimately, voice pattern recognition is possible during the user's voice recognition process, and based on this, it is possible to determine the user's stress and derive analysis results that can respond to it.

[0062] As such, the present invention recognizes speech by considering various intonations, accents, and pronunciation habits, thereby recognizing differences in pronunciation between people or dialects, and accurately deriving the user's emotional judgment using data including pronunciation information at the phoneme or phonological unit level, which enables natural speech recognition and thus enables AI stress feedback based on speech pattern recognition. Explanation of the symbols

[0064] 201: Acoustic Analysis Module 203: Pronunciation Analysis Module 205: Language Modeling Module 207: Sentiment Analysis Module 209 : Text conversion module

Claims

Claim 1 An artificial intelligence device capable of determining a user's stress state based on voice pattern recognition and providing corresponding feedback, comprising: an acoustic analysis module (201) that extracts pronunciation information corresponding to phonemes or phonological information from an input voice signal; a pronunciation analysis module (203) that determines pronunciation by comparing the voice signal and pronunciation information from the acoustic analysis module (201), and performs voice recognition by considering the user's intonation, accent, and pronunciation habits in the phonological information; a language modeling module (205) that performs context-appropriate voice recognition by performing an N-gram model or a statistical language model on the voice signal from the acoustic analysis module (201); and a sentiment analysis module that determines the emotion embedded in a sentence by calculating similarity in a vector space while preserving the meaning and context of words through embedding that converts a sentence into a numerical vector and presents it as a high-dimensional vector based on the pronunciation analysis module (203) and the language modeling module (205), and then performs deep learning to predict sentiment by receiving the embedded sentence vector as input to determine the positive, negative, and neutrality of the contextual sentiment. An artificial intelligence stress feedback device having a voice pattern recognition function, characterized by comprising a module (207); and a text conversion module (209) that converts the context of the voice determined by the sentiment analysis module (207) into text and derives the sentiment as an emoticon. Claim 2 An artificial intelligence stress feedback device having a voice pattern recognition function, wherein, in claim 1, the acoustic analysis module (201) extracts feature information from sequence data of a voice signal by using a deep learning algorithm or a Hidden Markov Model (HMM) method, and the feature information is characterized by using a Mel Frequency Cepstral Coefficients (MFCC) feature extraction algorithm capable of extracting features of frequency, energy, and frequency change from the voice signal. Claim 3 An artificial intelligence stress feedback device having a speech pattern recognition function, wherein, in claim 1, the pronunciation analysis module (203) receives an MFCC feature vector extracted from the acoustic analysis module (201) and uses a Deep Neural Network (DNN) to predict a phoneme sequence. Claim 4 An artificial intelligence stress feedback device having a speech pattern recognition function, wherein the language modeling module (205) models the structure and probabilistic characteristics of language, and an N-gram model or a statistical language model is applied to predict the probability of the next word or sentence in a given context. Claim 5 An artificial intelligence stress feedback device having a speech pattern recognition function, characterized in that, in claim 1, the sentiment analysis module (207) can adjust the weights of the model in a direction that minimizes the difference between the predicted value and the actual sentiment label while training the model using training data, and applies Gradient Descent, which performs optimization during the training process by defining a loss function. Claim 6 An artificial intelligence stress feedback device having a voice pattern recognition function, characterized in that, in any one of claims 1 to 5, the sentiment analysis procedure of the sentiment analysis module (207) comprises: a data collection and preprocessing procedure for data processed by the model, which performs operations such as tokenizing sentences from sentiment label data provided from the pronunciation analysis module (203) and converting words into numbers; and a process of converting the sentences into numeric vectors. Embedding comprises: an embedding procedure that represents words or sentences as high-dimensional vectors; a deep learning procedure that receives the embedded sentence vector as input and predicts sentiment; a learning optimization procedure that trains and optimizes the deep learning model, which minimizes the difference between the predicted value and the actual sentiment label while training the model using training data; and a prediction procedure that, after embedding the input sentence, gives it as input to the deep learning model and classifies the sentiment of the sentence into the class with the highest probability among positive, negative, and neutral classes based on the predicted value.

Citation Information

Patent Citations

  • Regional features based speech recognition method and system

    KR1020200007983A

  • Method and Apparatus for Determining Stress in Speech Signal Using Weight

    KR102300599B1

  • Apparatus and method for recognizing emotion in speech

    KR102607373B1