Emotion Detection in Audio Interactions
A machine learning model aligns phoneme boundaries with acoustic features to enhance emotion detection accuracy in contact centers, addressing subjective annotation issues and improving response strategies for customer agitation.
Patent Information
- Application Number
- JP2022529849
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2020-12-21
- Publication Date
- 2025-11-10
- Estimated Expiration
- 2040-12-21
AI Technical Summary
Existing emotion detection systems in audio interactions, particularly in contact centers, face challenges due to the subjective nature of human annotation and the lack of alignment between phoneme sequences and acoustic features, leading to inaccurate emotion classification.
A machine learning model is trained using acoustic features extracted from audio segments, aligned with phoneme boundaries, and employs a BiLSTM-CRF architecture to classify emotions by considering the sequential nature of conversations, with features normalized at frame, speaker, and phoneme levels.
Improves emotion detection accuracy by aligning phoneme boundaries with acoustic features, providing consistent classification and enabling effective response strategies to customer agitation in contact centers.
Smart Images

Figure 0007766594000002 
Figure 0007766594000003 
Figure 0007766594000004
Abstract
Description
[Technical Field]
[0001] (Priority Claim) This application claims priority to U.S. Patent Application No. 16 / 723,154, filed December 20, 2019, and entitled "EMOTION DETECTION IN AUDIO INTERACTIONS."
[0002] FIELD OF THE INVENTION The present invention relates to the field of automatic, computerized emotion detection. [Background technology]
[0003] Many commercial enterprises conduct and record multiple audio interactions with customers, users, or other persons every day. Often, these organizations may want to extract as much information as possible from the interactions, for example, to improve customer satisfaction and prevent customer churn.
[0004] The measurement of negative emotions conveyed in customer voice serves as a key performance indicator of customer satisfaction. Furthermore, managing customers' emotional responses to the services provided by organizational representatives can improve customer satisfaction and reduce customer churn.
[0005] The foregoing examples of the related art and limitations associated therewith are intended to be illustrative and not exhaustive. Other limitations of the related art will become apparent to those skilled in the art upon reading this specification and studying the drawings. Summary of the Invention
[0006] The following embodiments and aspects thereof are described and illustrated in conjunction with systems, tools, and methods that are meant to be exemplary and illustrative, not limiting in scope.
[0007] In one embodiment, a method is provided that includes receiving a plurality of audio segments comprising a speech signal, the audio segments representing a plurality of verbal interactions; receiving labels associated with an emotional state expressed in each of the audio segments; dividing each of the audio segments into a plurality of frames based on a specified frame duration; extracting a plurality of acoustic features from each of the frames; calculating statistics across the acoustic features for a sequence of frames representing phoneme boundaries within the audio segment; during a training phase, training a machine learning model on a training set including (i) the statistics associated with the audio segments and (ii) the labels; and during an inference phase, applying the trained machine learning model to one or more target audio segments comprising the speech signal to detect the emotional states expressed in the target audio segments.
[0008] Also, in one embodiment, a system is provided, comprising: at least one hardware processor; and a non-transitory computer-readable storage medium storing program instructions, the program instructions being executable by the at least one hardware processor to: receive a plurality of audio segments comprising a speech signal, the audio segments representing a plurality of verbal interactions; receive a label associated with an emotional state expressed in each of the audio segments; divide each of the audio segments into a plurality of frames based on a specified frame duration; extract a plurality of acoustic features from each of the frames; calculate statistics across the acoustic features with respect to a sequence of frames representing phoneme boundaries within the audio segment; during a training phase, train a machine learning model on a training set including (i) the statistics associated with the audio segments and (ii) the labels; and during an inference phase, apply the trained machine learning model to one or more target audio segments comprising the speech signal to detect the emotional states expressed in the target audio segments.
[0009] Further, in one embodiment, there is provided a computer program product comprising a non-transitory computer-readable storage medium having program instructions embodied thereon, the program instructions being executable by at least one hardware processor to: receive a plurality of audio segments comprising a speech signal, the audio segments representing a plurality of verbal interactions; receive a label associated with an emotional state expressed in each of the audio segments; divide each of the audio segments into a plurality of frames based on a specified frame duration; extract a plurality of acoustic features from each of the frames; calculate statistics across the acoustic features with respect to a sequence of frames representing phoneme boundaries within the audio segment; during a training phase, train a machine learning model on a training set including (i) the statistics associated with the audio segments and (ii) the labels; and during an inference phase, apply the trained machine learning model to one or more target audio segments comprising the speech signal to detect the emotional states expressed in the target audio segments.
[0010] In some embodiments, the audio segments are arranged in a temporal sequence based on their association with a specified interaction of verbal interactions.
[0011] In some embodiments, the boundaries of the temporal sequence are determined based at least in part on the continuity of the audio signal within the audio segment.
[0012] In some embodiments, statistics are computed for time-sequenced audio segments, and labels are associated with emotional states expressed in the audio segments.
[0013] In some embodiments, the training set further includes vector representations of phonemes defined by phoneme boundaries.
[0014] In some embodiments, the emotional state is one of neutral and negative.
[0015] In some embodiments, the acoustic features include Mel-frequency cepstral coefficients (MFCCs), Probability-of-Voicing (POV) features, pitch features, cutoff frequencies, signal-to-noise-ratio (SNR) characteristics, speech descriptors, vocal tract characteristics, loudness, signal energy, spectral distribution, slope, clarity, spectral flux, Chroma The feature is selected from the group consisting of: a zero-crossing rate (ZCR);
[0016] In some embodiments, the statistic is selected from the group consisting of the mean and the standard deviation.
[0017] In some embodiments, the phoneme boundaries are obtained based on applying a speech-to-text machine learning model to the audio segment.
[0018] In some embodiments, the extracting further includes a feature normalization step, wherein normalization is performed on at least one of features associated with all of the frames, features associated with frames representing speech by a particular speaker within the verbal interaction interaction, and features associated with a sequence of frames representing phoneme boundaries associated with speech by a particular speaker within the verbal interaction interaction.
[0019] In some embodiments, the verbal interaction represents a conversation between a customer and a call center agent.
[0020] In some embodiments, audio segments containing speech signals representing speech by an agent are removed from the training set.
[0021] In some embodiments, the target audio segment is a temporal sequence of audio segments from an individual verbal interaction.
[0022] In addition to the exemplary aspects and embodiments described above, further aspects and embodiments will become apparent by reference to the figures and by study of the following detailed descriptions. [Brief explanation of the drawings]
[0023] Exemplary embodiments are illustrated in the referenced drawings, in which dimensions of components and features shown are generally chosen for convenience and clarity of presentation and are not necessarily shown to scale. The drawings are listed below.
[0024] [Figure 1] 1 illustrates an exemplary frame-level feature extraction scheme according to some embodiments of the present invention.
[0025] [Figure 2] 10 illustrates mid-term features computed from frame-level features according to some embodiments of the present invention.
[0026] [Figure 3] 1 illustrates an exemplary neural network according to some embodiments of the present disclosure.
[0027] [Figure 4] 1 is a flowchart illustrating functional steps in a process for training a machine learning model to classify audio segments containing speech, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0028] Disclosed herein are methods, systems, and computer program products for automated and accurate emotion recognition and / or detection in audio signals.
[0029] In some embodiments, the present disclosure provides a machine learning model trained to classify one or more audio segments containing speech based on detected emotions expressed by speakers in the audio segments.
[0030] In some embodiments, the machine learning model is trained on a training set that includes acoustic features extracted from a plurality of audio segments containing speech utterances, and the audio segments are manually annotated with associated emotions expressed by the speaker, for example, negative (upset) or neutral emotions.
[0031] In some embodiments, acoustic features may be extracted from a speech signal, e.g., an input audio segment comprising a speech utterance, at the phoneme level detected in the signal. Thus, in some embodiments, acoustic features may be associated with identified phoneme boundaries of the speech signal.
[0032] In some embodiments, the present disclosure provides for detecting phoneme sequences in an audio signal. In some embodiments, the present disclosure provides for detecting phoneme boundaries and / or expressions within an audio signal.
[0033] In some embodiments, the present disclosure generates an alignment between phoneme boundaries detected in an audio signal and a sequence of frames, the frames having a specified frame duration, for example, 25 ms.
[0034] In some embodiments, phoneme speech signal alignment provides for extracting acoustic and / or related features from the speech signal and associating these acoustic features with particular phoneme boundaries within the speech signal.
[0035] In some embodiments, input to a trained machine learning model includes a sequence of audio segments containing vocal utterances by a speaker. In some embodiments, emotions, and specifically negative emotions, expressed by a speaker in a verbal interaction may follow a structure that includes several peaks spaced over a period of time. Thus, the present disclosure attempts to capture multiple audio segments that represent the "vibration plot" of a conversation or interaction, and learn and potentially exploit the interdependencies between emotions expressed in neighboring segments to improve prediction accuracy and effectiveness.
[0036] In some embodiments, the present disclosure provides emotion detection methods that can be deployed to perform emotion detection analysis on verbal interactions, such as telephone conversation recordings. The techniques disclosed herein are particularly useful for emotion detection in telephone conversation recordings in call or contact center contexts. Contact center interactions typically have at least two sides, for example, an agent and a customer. These interactions can reflect conversations of various lengths (e.g., from a few minutes to an hour or more), can shift in tone and emotion over the course of the interaction, and typically have a defined emotion flow or "agitation plot."
[0037] Thus, the present disclosure may provide contact center operators with the ability to detect audio contacts with high emotional content expressed by either customers or agents. Agitated customers are likely to call repeatedly, initiate escalations to a supervisor, complain to a customer service representative, or transfer to another company. Therefore, it is important to quickly identify contacts in which agitation is expressed and respond promptly to prevent further escalation. The response may be, for example, to send an experienced agent to resolve the argument with the customer through careful negotiation.
[0038] Emotion recognition features can also be used to evaluate agent performance, for example, based on their ability to resolve and de-escalate tense situations, to identify successful approaches to customer dispute resolution or, on the other hand, to identify agents who may need more advanced training in de-escalation techniques. Additionally, emotion recognition features can also help identify agents who tend to become agitated, e.g., express anger, which can further help improve agent training.
[0039] The task of speech emotion recognition has been studied for decades, and numerous methods have been developed using various algorithmic approaches. The accepted approach is to train a classifier using audio segments of different emotions and teach the classifier to distinguish emotions by various audio characteristics, such as pitch and energy. Several algorithms, such as support vector machines (SVMs), hidden Markov models (HMMs), and various neural network architectures, such as recurrent neural networks (RNNs) or attention neural networks, have been used. Various approaches have attempted to train machine learning models using, for example, audio spectrograms representing speech signals or combinations of audio spectrograms and associated phoneme sequences. However, these approaches use raw associated phoneme sequences but do not align them with the audio signal, so these sequences do not influence how acoustic features are ultimately extracted. Emotion detection task - experimental results
[0040] We conducted an experiment seeking to demonstrate the subjective nature of the task of classifying and detecting emotions in speech. To this end, we obtained 84 recorded speech segments from multiple speakers expressing negative emotions (e.g., emotional upset). For each "negative" segment, two "neutral" speech samples from the same speaker, sampled during the same conversation (e.g., using a text recognition algorithm and / or similar method), were selected. The test set included a total of 238 segments, with a negative-to-neutral ratio of approximately 1:2. All of these segments were then manually annotated by three human annotators.
[0041] The results show that 113 of the 238 segments were tagged as "negative" by at least one of the annotators. 24 were consensus annotations by all three annotators. 43 were annotated as negative by only two annotators. 46 were annotated as negative by only one annotator.
[0042] These results are consistent with the scientific literature and mitigate the highly subjective nature of speech emotion recognition tasks, which is even more pronounced in the contact center domain, where discussions involving financial issues and service quality complaints are frequent, and human annotators may tend to identify with the customer, which in turn may affect their interpretation of the gist of the conversation.
[0043] Based on these experimental results, we set an objective for our emotion detection model: classification of an audio signal as "negative" should only be consistent with similar annotations by at least two human annotators. Collecting training data
[0044] In some embodiments, a training dataset may be constructed to train the audio emotion detection machine learning model of the present disclosure.
[0045] In some embodiments, the training data set may include a plurality (eg, thousands) of relatively short audio segments containing speech utterances, having an average duration of about 5 seconds.
[0046] In some embodiments, multiple audio segments are manually annotated as expressing either negative or neutral emotions.
[0047] In some embodiments, the manual annotation is based solely on the acoustic aspects of the audio segment, while ignoring the textual content.
[0048] In some embodiments, the training dataset includes audio segments that represent entire and / or significant portions of individual verbal interactions. In some embodiments, the training dataset identifies and / or groups together segments that represent individual interactions. In some embodiments, the training dataset may also include, for example, only a single segment from one or more interactions, multiple randomly selected segments from one or more interactions, and / or disaggregated segments from multiple interactions.
[0049] In some embodiments, using segments that represent the entirety and / or significant portions of individual verbal interactions may provide more accurate prediction results. Because emotional expression is highly individual and speaker-dependent, in some implementations, it may be advantageous to include samples covering a range of tones and / or tenors expressed by the same speaker, e.g., associated with negative and neutral speech samples from the same speaker. Furthermore, multiple segments covering the entirety and / or significant portions of individual verbal interactions may provide interrelated context. For example, our experimental results showed that negative segments tend to be adjacent to each other within a conversation arc. Therefore, the temporal placement of segments may possess indicative power for their emotional content and thus help improve prediction results.
[0050] In some embodiments, this approach may be advantageous for the task of normalizing acoustic features. In some embodiments, this approach may be advantageous in the context of manual annotation, as it may provide the annotator with a wider range of different expressive voices from the same speaker to help distinguish between different emotions. For example, a high pitch may generally indicate agitation, but if a speaker's voice has a natural high pitch, then that speaker's agitation segments are likely to reflect a particularly high pitch.
[0051] In some embodiments, the training data set includes a negative-to-neutral segment ratio that reflects and / or approximates an actual ratio from call center interactions, hi some embodiments, the training data set has a negative-to-neutral ratio of about 4:1. Audio Feature Extraction
[0052] In some embodiments, audio segments containing speech signals included in a training dataset may be subjected to an acoustic feature extraction stage. In some embodiments, in the preprocessing stage, the audio segments may be divided into frames of 25 ms in length, for example, in steps of 10 ms. Thus, for example, a 5-second segment may be divided into 500 frames. In other examples, different frame lengths and step sizes, for example, longer or shorter frame lengths and overlapping sections, may be used.
[0053] In some embodiments, an 8000 Hz sampling rate may be used to sample the acoustic features, and a 25 ms frame may be sampled 200 times. In some embodiments, other sampling rates, e.g., higher or lower sampling rates, may be used.
[0054] In some embodiments, Mel-Frequency Cepstral Coefficients (MFCCs), Probability of Vocalization (POV) features, pitch features, cutoff frequencies, signal-to-noise ratio (SNR) characteristics, speech descriptors, vocal tract characteristics, loudness, signal energy, spectral distribution, slope, clarity, spectral flux, Chroma Short-term acoustic features may be extracted or calculated, which may include one or more of the following:
[0055] In some embodiments, certain features may be extracted using, for example, the Kaldi ASR toolkit (see, "The Kaldi speech recognition toolkit," IEEE 2011 workshop on automatic speech recognition and understanding, Daniel Povey et al., no. CONF. IEEE Signal Processing Society, 2011). In other implementations, additional and / or other feature extraction tools and / or techniques may be used.
[0056] In some embodiments, between 5 and 30, eg, 18, short-term acoustic features may be extracted for each frame, for a total of 1800 values per second of audio.
[0057] In some embodiments, medium-term features may be extracted and / or calculated for windows of the audio signal that are longer than a single frame (e.g., windows having a duration of tens to hundreds of milliseconds). In some embodiments, the medium-term features include statistical calculations of short-term features over the duration of such windows, e.g., by calculating means, standard deviations, and / or other statistics regarding the features over two or more consecutive frames or portions thereof.
[0058] It has been shown that emotional states have different effects on different phonemes (see, for example, Chul Min Lee et al., "Emotion Recognition based on Phoneme Classes," Eighth International Conference on Spoken Language Processing, 2004). Based on this observation, an obvious problem with arbitrary division into medium-term time windows is that the features of different phonemes are mixed, which makes no sense for, for example, the average pitch values of vowels and consonants. Therefore, in some embodiments, the present disclosure seeks to provide an alignment between frames and phonemes, which better enables setting the boundaries of medium-term windows based on the phonemes they represent.
[0059] Phoneme alignment is the task of properly positioning a sequence of phonemes relative to a corresponding continuous speech signal. This problem is also called phoneme segmentation. Accurate and fast alignment procedures are necessary tools for developing speech recognition and text-to-speech systems. In some embodiments, phoneme alignment is provided by a trained speech-to-text machine learning model.
[0060] In some embodiments, the medium-term windows used to calculate the feature statistics provide windows of various lengths and / or number of frames, with the window boundaries being determined according to aligned phoneme boundaries in the speech signal. Feature Normalization
[0061] In some embodiments, before calculating statistics of the short-term features, the short-term features may be normalized, for example, with respect to the following categories: (a) Every frame in the data, (b) all frames associated with a particular speaker in the interaction; and / or (c) All frames that represent similar phonemes and belong to the same speaker.
[0062] FIG. 1 illustrates an exemplary frame-level feature extraction scheme according to some embodiments of the present invention.
[0063] Figure 2 shows the mid-term features calculated from the frame-level features in Figure 1, where M denotes the mean function and S denotes the standard deviation function.
number
[0064] In some embodiments, an additional phoneme-level feature, local phoneme rate, may be calculated, for example, measuring the phoneme rate in a 2-second window around each phoneme. This feature is calculated as (#num-phonemes / #frames) taking into account speaking rate and normalized once per speaker and once for the entire data. Classification Algorithm - Architecture
[0065] In some embodiments, the present disclosure provides an emotion detection machine learning model based on a neural network, in which the neural network architecture is configured to take into account the sequential characteristics of individual audio segments whose basic units are phonemes, as well as the sequential nature of segments in a conversation.
[0066] In some embodiments, the present disclosure uses a BiLSTM-CRF (bidirectional LSTM with a conditional random field layer on top) architecture that receives, at each timestamp, a set of segment-level features computed by another BiLSTM.
[0067] In some embodiments, the inner neural network, which runs for each audio segment, receives as input a sequence of phoneme-level features (i.e., their timestamps correspond to phoneme changes in the segment). These features are 108 medium-term statistics plus two phoneme rate values, as described above. These are supplemented with additional vector representations of the phonemes (e.g., one-hot representations) embedded in layers of size 60, which are also learned during training. The inner network (or segment-level network) is connected to the outer network (or document-level network) via the last hidden state of the former, which is forwarded as input to the latter. In some embodiments, the document-level network is a BiLSTM-CRF, as it takes into account both the output features in the "tag" layer (containing two neurons, "neutral" and "negative") and the transition probabilities between tags of consecutive segments.
[0068] Figure 3 illustrates an exemplary neural network of the present disclosure. In some embodiments, this architecture resembles similar networks used in the context of named entity recognition, where the timestamps of the outer network correspond to words in a sentence, and the inner network generates word features by scanning the characters of each word. Classification Algorithm - Training
[0069] In some embodiments, a training dataset of the present disclosure may include a series of subsequent audio segments representing speech signals from multiple interactions, for example, calls or conversation recordings in a contact center context.
[0070] In some embodiments, a training dataset of the present disclosure may include a sequence of sequential audio segments from an interaction. In some embodiments, each such sequential sequence includes two or more audio segments comprising an audio signal from a speaker. In some embodiments, one or more particular interactions may each be represented by multiple sequential sequences of varying lengths.
[0071] In some embodiments, audio segments within each sequence are positioned based, for example, on the start time value of each segment. In some embodiments, audio segment sequence boundaries within an interaction are determined based, for example, on audio signal gaps. For example, natural conversations include pauses or gaps, such as when a call is placed on hold or when a speaker remains silent for a period of time. In some embodiments, a gap in the audio signal of, for example, 3 to 7 seconds, e.g., 5 seconds, may define the start of a new sequence in an interaction.
[0072] In some embodiments, the training data set of the present disclosure may be constructed from audio segment sequences. In some embodiments, the training data set may include portions (i.e., subsequences) of sequences, each subsequence consisting of, for example, between 5 and 15 adjacent segments. In some embodiments, entire sequences having fewer than 5 consecutive segments may be used in their entirety.
[0073] In some embodiments, the training dataset construction process of the present disclosure further includes removing audio segments from sequences and subsequences that include voice signals associated with agents in call center interactions. In some embodiments, audio segments that include customer voice signals can be combined in sequences even if they are interspersed with agent voice signals. This approach reflects the idea that the customer's agitated state is maintained throughout the sequence, even when interrupted by agent voice signals. Classification Algorithms - Estimation
[0074] In some embodiments, a trained machine learning model of the present disclosure may be applied to an audio signal, e.g., a target audio section containing dialogue, to generate a prediction regarding the emotional state of the speaker.
[0075] In some embodiments, the target audio section may be divided into segments, each having a duration of several seconds, based on detected gaps (e.g., silence points) in the speech signal. In some embodiments, a trained machine learning model may be applied to the sequence of segments to generate an output including, for example, a classification (e.g., negative or neutral) and a confidence score.
[0076] In some embodiments, the BiLSTM-CRF layer of the neural network of the present machine learning model includes a trained transition score matrix that reflects the likelihood of a transition between any two classes in a subsequent segment. In some embodiments, the most likely path, i.e., between classes o1,...,o2, is selected. n Segments s1,...,s for outputting n The assignments of can be decoded using the Viterbi algorithm, which takes into account both the class scores and the transition matrix for each segment.
[0077] In some embodiments, the confidence score may represent a predicted emotion level for an audio segment. In some embodiments, to provide a smooth transition of confidence scores between subsequent segments, the present disclosure provides a confidence score that represents the contextual emotion of neighboring audio segments. Thus, in some embodiments, the confidence (or probability) is calculated using a forward-backward algorithm to classify segments s as "negative." i The probability of n divided by the sum of the scores of all possible paths 、 is equal to the sum of the scores of all the paths taken (s i negative).
[0078] FIG. 4 is a flowchart illustrating functional steps in a process for training a machine learning model to classify audio segments containing speech based on acoustic features extracted from the audio segments.
[0079] A plurality of audio segments reflecting speech signals from a verbal interaction are received at step 400. In some embodiments, the interaction reflects, for example, a call or conversation transcript in a contact center context.
[0080] In step 402, each audio segment is labeled or annotated with the emotion expressed by the speaker in the segment, e.g., "neutral" or "negative." In some embodiments, the annotation or labeling process is performed on a temporal sequence of segments that includes parts or entire interactions, enhancing the annotator's ability to perceive speaker-dependent idiosyncrasies in the expression of emotions.
[0081] In step 404, the annotated segments are arranged in a temporal sequence associated with each interaction by the speaker. In some embodiments, a section or portion of the sequence containing 5-15 adjacent segments is extracted from the sequence. In some embodiments, segments representing audio signals from the agent side of the interaction are removed before further processing.
[0082] Acoustic features are extracted from the audio segment in step 406. In some embodiments, frame-level acoustic features are extracted from each frame (e.g., approximately 25 ms in duration).
[0083] In some embodiments, mid-term features (i.e., representing multiple frames) are extracted in step 408. In some embodiments, window boundaries for extracting the mid-term features are determined based, at least in part, on detected phoneme boundaries in the speech signal.
[0084] In some embodiments, at step 410, a training data set is constructed that includes frame-level acoustic features, mid-term acoustic features, and annotations associated with audio segments.
[0085] In some embodiments, in step 412, a machine learning model is trained on the training dataset.
[0086] In some embodiments, in an inference stage, at step 414, the trained machine learning model is applied to one or more target audio segments to detect emotions in the audio segments. In some embodiments, the target audio segments include interactions that reflect conversation recordings, for example, in a call or contact center context. In some embodiments, the inference stage input includes frame-level and / or mid-level acoustic features extracted from the target audio segments.
[0087] The present invention may be a system, a method, and / or a computer program product, which may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0088] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction-execution device. The computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, and mechanically encoded devices having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over wires. Rather, a computer-readable storage medium is a non-transitory (i.e., non-volatile) medium.
[0089] The computer-readable program instructions described herein can be downloaded to each computing / processing device from a computer-readable storage medium or an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to a computer-readable storage medium within the respective computing / processing device for storage.
[0090] Computer-readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including, for example, object-oriented programming languages such as Java, Smalltalk, C++, and conventional processing programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer as a standalone software package, partially on the user's computer, partially on the user's computer, or partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by utilizing state information in the computer-readable program instructions to individualize the electronic circuitry to perform aspects of the present invention.
[0091] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0092] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to generate a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the function(s) / act(s) specified in the flowchart and / or block diagram block(s). These computer-readable program instructions may also be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture that includes instructions that implement aspects of the function(s) / act(s) specified in the flowchart and / or block diagram block(s).
[0093] Computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to perform a series of operational steps on the computer, other programmable apparatus, or other device to create a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the function / act specified in the flowchart and / or block diagram block(s).
[0094] The flowcharts and block diagrams in the figures illustrate the architecture, functions, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may possibly be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs a particular function or that performs or executes a combination of dedicated hardware and computer instructions.
[0095] The recitation of a range of numerical values should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, recitation of a range of 1 to 6 should be considered to have specifically disclosed subranges of 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6, etc. This applies regardless of the breadth of the range.
[0096] The description of various embodiments of the present invention is presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to enable those skilled in the art to understand the principles of the embodiments, practical applications or technical improvements of technologies found in the market, or the embodiments disclosed herein.
[0097] The above-described experiments conducted demonstrate the usefulness and effectiveness of embodiments of the present invention. Some embodiments of the present invention may be constructed based on specific experimental methods and / or experimental results. Therefore, the following experimental methods and / or experimental results should be considered as embodiments of the present invention.
Claims
1. 1. A method comprising: receiving a plurality of audio segments including a speech signal, the audio segments representing a plurality of verbal interactions; receiving a label associated with an emotional state expressed in each of the audio segments; Dividing each of the audio segments into a plurality of frames based on a designated frame duration, the audio segments being arranged in a temporal sequence including 5 to 15 adjacent segments based on their association with a designated interaction of the verbal interaction; extracting frame-level acoustic features from each of the frames; calculating statistics across the acoustic features for a sequence of two or more frames between two adjacent phoneme boundaries in the audio segment, the statistics being based on a plurality of the frame-level acoustic features and representing mid-term acoustic features; During the training phase, (i) the statistics associated with the audio segments; and (ii) training a machine learning model on a training set including the label, the model receiving the statistics as input and generating an output including a classification of the emotional state; and in an inference stage, applying the trained machine learning model to one or more target audio segments comprising a speech signal to detect emotional states expressed in the target audio segments.
2. The method of claim 1 , wherein the temporal sequence boundaries are determined based at least in part on the continuity of the speech signal within the audio segment.
3. The method of claim 1 , wherein the statistics are calculated for the audio segments arranged in the temporal sequence, and the labels are associated with emotional states expressed in the audio segments.
4. The method of claim 1 , wherein the training set further comprises vector representations of the phonemes defined by the phoneme boundaries.
5. The method of claim 1 , wherein the emotional state is one of neutral and negative.
6. 2. The method of claim 1, wherein the acoustic features are selected from the group consisting of Mel-Frequency Cepstral Coefficients (MFCC), Probability of Voicing (POV) features, pitch features, cutoff frequency, signal-to-noise ratio (SNR) characteristics, speech descriptors, vocal tract characteristics, loudness, signal energy, spectral distribution, slope, clarity, spectral flux, chroma features, and zero-crossing rate (ZCR).
7. The method of claim 1 , wherein the statistic is selected from the group consisting of a mean and a standard deviation.
8. The method of claim 1 , wherein the phoneme boundaries are obtained based on applying a speech-to-text machine learning model to the audio segment.
9. 2. The method of claim 1, wherein the extracting further comprises a feature normalization step, and wherein the normalization is performed on at least one of features associated with all of the frames, features associated with frames representing speech by a particular speaker within the verbal interaction interaction, and features associated with the sequence of frames representing phoneme boundaries associated with speech by a particular speaker within the verbal interaction interaction.
10. The method of claim 1 , wherein the verbal interaction represents a conversation between a customer and a call center agent.
11. The method of claim 10 , wherein those of the audio segments that include speech signals representing speech by the agent are removed from the training set.
12. The method of claim 1 , wherein the target audio segments are temporal sequences of audio segments from individual verbal interactions.
13. 1. A system comprising: at least one hardware processor; A non-transitory computer-readable storage medium storing program instructions, the program instructions comprising: receiving a plurality of audio segments including a speech signal, the audio segments representing a plurality of verbal interactions; receiving a label associated with an emotional state expressed in each of the audio segments; Dividing each of the audio segments into a plurality of frames based on a designated frame duration, the audio segments being arranged in a temporal sequence including 5 to 15 adjacent segments based on their association with a designated interaction of the verbal interaction; extracting frame-level acoustic features from each of the frames; calculating statistics across the acoustic features for a sequence of two or more frames between two adjacent phoneme boundaries in the audio segment, the statistics being based on a plurality of the frame-level acoustic features and representing mid-term acoustic features; During the training phase, (i) the statistics associated with the audio segments; and (ii) training a machine learning model on a training set including the label, the model receiving the statistics as input and generating an output including a classification of the emotional state; and a non-transitory computer-readable storage medium executable by the at least one hardware processor to: apply, during an inference stage, the trained machine learning model to one or more target audio segments comprising a speech signal to detect emotional states expressed in the target audio segments.
14. The system of claim 13 , wherein the temporal sequence boundaries are determined based at least in part on continuity of the speech signal within the audio segment.
15. 14. The system of claim 13, wherein the statistics are computed for the audio segments arranged in the temporal sequence, and the labels are associated with emotional states expressed in the audio segments.
16. The system of claim 13 , wherein the training set further comprises vector representations of the phonemes defined by the phoneme boundaries.
17. The system of claim 13 , wherein the emotional state is one of neutral and negative.
18. 14. The system of claim 13, wherein the acoustic features are selected from the group consisting of Mel-Frequency Cepstral Coefficients (MFCC), Probability of Voicing (POV) features, pitch features, cutoff frequency, signal-to-noise ratio (SNR) characteristics, speech descriptors, vocal tract characteristics, loudness, signal energy, spectral distribution, slope, clarity, spectral flux, chroma features, and zero-crossing rate (ZCR).
19. The system of claim 13 , wherein the statistic is selected from the group consisting of a mean and a standard deviation.
20. The system of claim 13 , wherein the phoneme boundaries are obtained based on applying a speech-to-text machine learning model to the audio segment.
21. 14. The system of claim 13, wherein the extracting further comprises a feature normalization step, wherein the normalization is performed on at least one of features associated with all of the frames, features associated with frames representing speech by a particular speaker within the verbal interaction interaction, and features associated with the sequence of frames representing phoneme boundaries associated with speech by a particular speaker within the verbal interaction interaction.
22. The system of claim 13 , wherein the verbal interaction represents a conversation between a customer and a call center agent.
23. 23. The system of claim 22, wherein those of the audio segments that include speech signals representing speech by the agent are removed from the training set.
24. The system of claim 13 , wherein the target audio segment is a temporal sequence of audio segments from an individual verbal interaction.
25. A non-transitory computer-readable storage medium having program instructions embodied therein, the program instructions comprising: receiving a plurality of audio segments including a speech signal, the audio segments representing a plurality of verbal interactions; receiving a label associated with an emotional state expressed in each of the audio segments; Dividing each of the audio segments into a plurality of frames based on a designated frame duration, the audio segments being arranged in a temporal sequence including 5 to 15 adjacent segments based on their association with a designated interaction of the verbal interaction; extracting frame-level acoustic features from each of the frames; calculating statistics across the acoustic features for a sequence of two or more frames between two adjacent phoneme boundaries in the audio segment, the statistics being based on a plurality of the frame-level acoustic features and representing mid-term acoustic features; During the training phase, (i) the statistics associated with the audio segments; and (ii) training a machine learning model on a training set including the label, the model receiving the statistics as input and generating an output including a classification of the emotional state; and (c) applying the trained machine learning model to one or more target audio segments comprising a speech signal to detect emotional states expressed in the target audio segments during an inference stage.
26. 26. The computer-readable storage medium of claim 25, wherein the temporal sequence boundaries are determined based at least in part on a continuity of the speech signal within the audio segment.
27. 26. The computer-readable storage medium of claim 25, wherein the statistics are calculated for the audio segments arranged in the temporal sequence, and the labels are associated with emotional states expressed in the audio segments.
28. 26. The computer-readable storage medium of claim 25, wherein the training set further comprises vector representations of phonemes defined by the phoneme boundaries.
29. 26. The computer-readable storage medium of claim 25, wherein the emotional state is one of neutral and negative.
30. 26. The computer-readable storage medium of claim 25, wherein the acoustic features are selected from the group consisting of Mel-Frequency Cepstral Coefficients (MFCC), Probability of Voicing (POV) features, pitch features, cutoff frequency, signal-to-noise ratio (SNR) characteristics, speech descriptors, vocal tract characteristics, loudness, signal energy, spectral distribution, slope, clarity, spectral flux, chroma features, and zero-crossing rate (ZCR).
31. 26. The computer-readable storage medium of claim 25, wherein the statistic is selected from the group consisting of a mean and a standard deviation.
32. 26. The computer-readable storage medium of claim 25, wherein the phoneme boundaries are obtained based on applying a speech-to-text machine learning model to the audio segment.
33. 26. The computer-readable storage medium of claim 25, wherein the extracting further includes a feature normalization step, and wherein the normalization is performed with respect to at least one of features associated with all of the frames, features associated with frames representing speech by a particular speaker within the verbal interaction interaction, and features associated with the sequence of the frames representing phoneme boundaries associated with speech by a particular speaker within the verbal interaction interaction.
34. 26. The computer-readable storage medium of claim 25, wherein the verbal interaction represents a conversation between a customer and a call center agent.
35. 35. The computer-readable storage medium of claim 34, wherein those of the audio segments that include speech signals representing speech by the agent are removed from the training set.
36. 26. The computer-readable storage medium of claim 25, wherein the target audio segment is a temporal sequence of audio segments from an individual verbal interaction.
Citation Information
Patent Citations
Information transmission device
JP2006113546A
Feeling detection method, feeling detection device, feeling detection program containing the method, and recording medium containing the program
WO2008032787A1
Satisfaction estimation model learning device, satisfaction estimation device, satisfaction estimation model learning method, satisfaction estimation method, and program
WO2019017462A1