Method and system for detecting emotions in audio data
By establishing a machine learning model that combines personalized baselines and environments, the problem of insufficient accuracy of emotion detection in speech recognition systems in different environments is solved, and personalized emotion detection and privacy protection are achieved.
Patent Information
- Application Number
- CN202080047662.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-28
- Filing Date
- 2020-06-08
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2040-06-08
AI Technical Summary
Existing speech recognition systems have difficulty accurately detecting users' emotions and feelings, especially in different environments and contexts, resulting in insufficient accuracy and personalization of emotion detection.
By using machine learning models and baseline technology, the system records the user's audio data in different emotional states, establishes a personalized baseline, and combines environmental and activity information for emotion detection.
It achieves personalized emotion detection in different environments and contexts, improves the accuracy of emotion detection and user privacy protection, and meets legal and user preferences.
Smart Images

Figure CN114051639B_ABST
Abstract
Description
[0001] Cross-references to related application data
[0002] This application claims the benefit of priority to U.S. patent application Ser. No. 16 / 456,158, filed on Jun. 28, 2019, in the name of Daniel Kenneth Bone et al., and entitled “EMOTION DETECTION USING SPEAKER BASELINE.” Background Art
[0003] Speech recognition systems have advanced to the point where humans can interact with computing devices using their voices. Such systems employ techniques to identify words spoken by human users based on various qualities of the received audio input. The audio input can also indicate the user's mood or emotion when the words were spoken.
[0004] Computers, handheld devices, telephone computer systems, kiosks, and a variety of other devices can use speech processing to improve human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] For a more complete understanding of the present disclosure, reference is now made to the following description taken in conjunction with the accompanying drawings.
[0006] Figure 1A A system configured to enroll users for detecting emotions in audio data is shown in accordance with an embodiment of the present disclosure.
[0007] Figure 1B A system configured to detect emotion in audio data is shown according to an embodiment of the present disclosure.
[0008] Figure 2A and Figure 2B is a conceptual diagram of the speech processing components of a system according to an embodiment of the present disclosure.
[0009] Figure 3 is a diagram of an illustrative architecture in which sensor data is combined to identify one or more users according to an embodiment of the present disclosure.
[0010] Figure 4 is a system flow diagram illustrating user identification according to an embodiment of the present disclosure.
[0011] Figure 5 is a conceptual diagram illustrating an emotion detection component for user registration according to an embodiment of the present disclosure.
[0012] Figure 6 is a conceptual diagram illustrating an emotion detection component according to an embodiment of the present disclosure.
[0013] Figure 7is a conceptual diagram of building a training model using training data according to an embodiment of the present disclosure.
[0014] Figure 8 is a conceptual diagram illustrating layers of a training model according to an embodiment of the present disclosure.
[0015] Figure 9 A neural network according to an embodiment of the present disclosure is shown, such as a neural network that can be used for emotion detection.
[0016] Figure 10 A neural network according to an embodiment of the present disclosure is shown, such as a neural network that can be used for emotion detection.
[0017] Figure 11 The operation of an encoder according to an embodiment of the present disclosure is shown.
[0018] Figure 12 is a block diagram conceptually illustrating exemplary components of an apparatus according to an embodiment of the present disclosure.
[0019] Figure 13 is a block diagram conceptually illustrating exemplary components of a server according to an embodiment of the present disclosure.
[0020] Figure 14 An example of a computer network for use with a speech processing system is shown. DETAILED DESCRIPTION
[0021] Automatic speech recognition (ASR) is a field of computer science, artificial intelligence, and linguistics that involves converting audio data associated with speech into text representing that speech. Similarly, natural language understanding (NLU) is a field of computer science, artificial intelligence, and linguistics that involves enabling computers to derive meaning from text input containing natural language. ASR and NLU are often used together as part of speech processing systems. Text-to-speech (TTS) is a field that involves converting text data into audio data that is synthesized to resemble human speech.
[0022] Certain systems can be configured to perform actions in response to user input. For example, in response to the user input "Alexa, play music," the system may output music. Furthermore, in response to the user input "Alexa, what's the weather like?", the system may output synthesized speech representing weather information for the user's geographic location. In another example, in response to the user input "Alexa, send a message to John," the system may capture the content of the voice message and output it via a device registered to "John."
[0023] Using speech to perform sentiment analysis can involve determining a person's emotions, opinions, and / or attitudes toward a situation / topic. A person's emotions can be derived from the words they use to express their opinions. Using speech to perform sentiment analysis can involve determining a person's psychological state, mood, feelings, and / or state of mind. A person's emotions can be derived from how they speak words and the acoustic characteristics of their speech. The system can be configured to classify audio data (e.g., speech from a user) based on emotions and / or feelings derived from the audio data and based on a baseline of users representing neutral emotions / emotions. For example, the system can use a baseline to help capture a speaker's personalized speaking style and associated characteristics. As used herein, speaking style can be represented by acoustic speech attributes such as pitch, speed, rate, accent, tone, stress, rhythm, intonation, volume, etc. As used herein, a baseline can refer to (reference) audio data that represents a speaker's neutral emotional state. By using baseline speech data for an individual speaker, the system can more accurately determine the emotions of an individual user during runtime. For example, a speaker who is typically loud may be perceived by the system as angry until compared to the speaker's baseline. In another example, a normally soft-voiced speaker may be perceived by the system as timid or sad until compared to the speaker's baseline. The system may determine the user's emotion / mood based on an analysis of the user's baseline and runtime input audio data indicating differences between the input audio data and the baseline.
[0024] The system can use the enrollment utterances spoken by the user to determine a baseline. The system can determine whether the enrollment utterances represent the user's neutral emotional state and whether they can be used as a baseline. The system can also be configured to obtain multiple baselines for different environments or activities that the speaker may be involved in. Different environments or activities can cause the speaker to exhibit different neutral emotional states. For example, when the speaker is at work, his or her baseline may be different relative to when the speaker is at home. Similarly, when the speaker is speaking to coworkers relative to when the speaker is speaking to family or children, his or her baseline may be different. As another example, when the speaker speaks in the evening relative to when the speaker speaks in the morning, his or her baseline may be different. Different baselines can capture different acoustic speech properties exhibited by the user in various situations / contexts.
[0025] The system may also be configured to select an appropriate baseline to analyze the input audio data for emotion / emotion detection based on the environment or activity in which the user was engaged when uttering the utterance represented in the runtime input audio data to be analyzed.
[0026] The system may incorporate user permissions and may only perform the functions disclosed herein (such as emotion detection, if approved by the user), and emotion detection may be configured according to user permissions / preferences. As disclosed herein, a user may register with the system for emotion detection by providing speech preferences. For example, the system may perform emotion detection on speech spoken by a user who has opted in and is associated with a capture device (rather than speech captured from other users). Thus, the systems, devices, components, and techniques described herein may restrict processing where appropriate and only process user information in a manner that ensures compliance with all applicable laws, regulations, standards, etc. The systems and techniques may be implemented on a geographic basis to ensure compliance with the laws of various jurisdictions and entities where system components and / or users are located. The system may delete all data related to emotion detection after a certain period of time and / or after the audio data has been analyzed and the output has been presented and / or viewed by the user. The user may also request that the system delete all data related to emotion detection. The system may restrict access to data related to emotion detection based on the user permissions selected by the user.
[0027] The system can process input audio data to determine whether the audio data includes voice activity (e.g., speech) from a human user. The system can then identify portions of the input audio data that represent speech from a specific user. A trained machine learning (ML) model can be used to process portions of the input audio data to predict the emotion category of the audio data. Emotion categories can be used in various applications. For example, emotion categories can be displayed to a user to indicate his or her emotions during interactions with other people and / or to indicate his or her emotions during specific times of the day. Application developers can also use emotion categories with voice-activated systems or smart speaker systems to identify the user's emotions and / or feelings when interacting with the voice-activated system or smart speaker system. Application developers may be able to determine a user's satisfaction with his or her interaction with the voice-activated system or smart speaker system. For example, a gaming application developer can determine a user's emotions while he or she is playing or interacting with a game. As another example, a user's emotions while watching or listening to a commercial can be used for marketing research. In yet another example, a voice-activated system or smart speaker system included in a vehicle can analyze the driver's emotions through audio data and inform the driver if he or she appears anxious, frustrated, or angry, such that his or her emotions / mood may affect his or her driving. Assuming user permission exists, other components can also receive emotion data for different operations.
[0028] In an exemplary embodiment, a user may wear or otherwise carry a device that detects audio data and initiates analysis of the audio data when voice activity is detected. The user may configure the device to monitor their voice interactions with others throughout the day. The system may determine the user's emotional state for various interactions and generate periodic reports for the user. The reports may be stored and / or displayed to the user, such as on a wearable device, phone, tablet, or other device.
[0029] Figure 1A Shown is a system 100 configured to enroll users for emotion detection in audio data according to an embodiment of the present disclosure. Figure 1B A system 100 configured to detect emotions in audio data according to an embodiment of the present disclosure is shown. Although the figures and discussion show specific operating steps of the system in a particular order, the steps described may be performed in a different order (and specific steps may be removed or added) without departing from the intent of the present disclosure. Figure 1A and Figure 1B As shown, the system 100 may include a device 110 local to the user 5 and one or more systems 120 connected across one or more networks 199. Figure 1A As shown, device 110a can communicate with device 110b. Figure 1A The described process may be performed during an enrollment operation (when the system 100 is configured to obtain reference audio data from the user 5 for emotion detection using a baseline) and with respect to Figure 1B The described process may be performed during runtime operation (when the configured system 100 processes input audio data to detect emotions).
[0030] During the enrollment process, one or more systems 120 are configured to obtain audio data representing the user's neutral emotional state. Figure 1A As shown, one or more systems 120 receive (132) audio data representing a first reference utterance. User 5 may speak the utterance represented by the audio data captured by device 110. As part of the registration process, one or more systems 120 may cause device 110 to output audio requesting user 5 to speak a specific sentence. In some embodiments, device 110 may output a specific sentence for user 5 to speak, such as "For registration purposes, please say I like today's weather", and user 5 may say "I like today's weather". One or more systems 120 may store audio data representing the reference utterance spoken by user 5.
[0031] The one or more systems 120 determine (134) whether the audio data can be used as a baseline. The one or more systems 120 may determine whether the audio data represents a neutral emotional state for the user 5. The one or more systems 120 may analyze the audio data to determine whether the corresponding acoustic speech attributes are within a predefined range or meet specific conditions indicating a neutral emotional state. For example, the one or more systems 120 may have identified and stored acoustic speech attributes that represent a neutral emotional state based on analyzing audio data from multiple users / general populations, and may use these acoustic speech attributes to determine whether the audio data (from operation 132) is consistent with these attributes so that the audio data represents a neutral emotional state for the user 5. In some embodiments, the one or more systems 120 may process the audio data using a machine learning (ML) model configured to determine an emotion category corresponding to the audio data. The ML model may determine that the emotion category corresponding to the audio data is neutral. In some embodiments, the one or more systems 120 may also determine whether the quality of the audio data is good enough to be used as a baseline.
[0032] If the one or more systems 120 determine that the audio data cannot be used as a baseline, the one or more systems 120 request (136) the user 5 to speak another utterance. The one or more systems 120 may request the user 5 to repeat the previously presented sentence by outputting, for example, "Please repeat that I like today's weather," or the one or more systems 120 may request the user 5 to speak a different sentence. The one or more systems 120 receive (138) the audio data representing the second reference utterance and return to operation 134 to determine whether the audio data can be used as a baseline (representing a neutral emotional state of the user).
[0033] In some embodiments, one or more systems 120 may make only a few attempts to obtain audio data for a baseline. That is, the operation of step 136 may be performed a limited number of times (e.g., two or three times) before one or more systems 120 recognizes that it is not possible to obtain audio data from user 5 to use as a baseline. One or more systems 120 may cause device 110 to output "something went wrong, let's try to check in at another time." The audio data may be of poor quality due to background noise, or the audio data may not represent a neutral emotional state of the user (e.g., the user may be overly excited or angry during the check-in process).
[0034] If the one or more systems 120 determine that the audio data can be used as a baseline, the one or more systems 120 store (140) the audio data as a baseline associated with the user profile corresponding to the user 5. The one or more systems 120 may store the audio data as a baseline in the profile storage device 270. The one or more systems may determine (142) a first feature vector corresponding to the audio data and may store the first feature vector as the baseline. The first feature vector may represent spectral features derived from the audio data. The one or more systems 120 may use an encoder (e.g., Figure 11 The encoder 1150 of the embodiment of the present invention processes the frames of the audio data and generates a first feature vector. In some embodiments, the first feature vector may represent the acoustic speech attributes (e.g., accent, pitch, prosody, etc.) exhibited by the user 5 in a neutral emotional state.
[0035] One or more systems 120 may be configured to obtain multiple baselines representing the user's neutral emotional state in various situations. Accordingly, profile storage device 270 may include audio data representing multiple baselines for user 5. In some embodiments, to obtain the various baselines, one or more systems 120 may request user 5 to speak while in different environments or while interacting with different people. One or more systems 120 may request the user's permission to record audio (for a limited period of time for registration purposes) while user 5 speaks in different environments or while interacting with different people to capture audio data representing the user's emotional state in different situations. For example, user 5 may exhibit different speaking style / acoustic voice attributes when at home compared to at work. Similarly, user 5 may exhibit different speaking style / acoustic voice attributes when speaking to family members (spouse, significant other, children, pets, etc.) compared to when speaking to coworkers.
[0036] In this case, one or more systems 120 may determine (144) contextual data corresponding to the audio data and may associate the contextual data with a baseline and a user profile (146). As used herein, contextual data corresponding to the audio data / baseline refers to data indicating the environment and / or circumstances associated with the user at the time the audio data was received. For example, one or more systems 120 may determine (e.g., using the location of device 110) the location of user 5 when the reference utterance was spoken. If the appropriate permissions and context are configured to allow operation, one or more systems 120 may determine with whom user 5 was interacting when the utterance was spoken. In some embodiments, one or more systems 120 may receive input data from user 5 indicating contextual data, such as the user's location (e.g., home, work, gym, etc.), with whom the user is interacting (e.g., coworkers, boss, spouse / significant other, children, neighbors, etc.), and the like. As a non-limiting example, one or more systems 120 may store in the profile storage device 270 first audio data representing a first baseline and context data indicating <location: workplace>, second audio data representing a second baseline and context data indicating <location: home>, third audio data representing a third baseline and context data indicating <person: colleague>, fourth audio data representing a fourth baseline and context data indicating <person: daughter>, and so on.
[0037] To obtain a baseline, in some embodiments, device 110 may output a specific sentence for user 5 to say, such as "For check-in purposes, please tell me how I like today's weather," and user 5 may say "I like today's weather." In some embodiments, one or more systems 120 may request user 5 to talk about a topic instead of requesting user 5 to say a specific sentence. For example, device 110 may output "For check-in purposes, please tell me how you feel about today's weather?" and user 5 may say "It's raining today, and I don't like it when it rains." One or more systems 120 may store audio data representing reference utterances spoken by user 5. In some embodiments, one or more systems 120 may request user 5 to say a specific sentence and also discuss a topic in order to capture audio data for both situations, because a user may exhibit different speaking styles / acoustic speech attributes when repeating a sentence than when freely discussing a topic. One or more systems 120 may process audio data representing a user speaking a particular sentence and audio data representing a user freely discussing a topic to determine an appropriate baseline using, for example, the difference in acoustic speech properties of the two cases, a (weighted or unweighted) average of the acoustic speech properties of the two cases, statistical analysis, a machine learning model to process corresponding feature vectors, and / or other methods.
[0038] In this way, one or more systems 120 Figure 1ADuring the illustrated enrollment process, audio data representing a neutral emotional state of user 5 is obtained. One or more systems 120 may perform multiple Figure 1A The operations shown are to obtain multiple baselines representing different situations that a user may be in and in which the user has opted in to emotion detection. Figure 5 Provide a description.
[0039] During runtime, such as Figure 1B As shown, one or more systems 120 receive (150) input audio data. The input audio data may be captured by device 110a and may include speech or sound from user 5 and / or speech and sound from at least one other person. Figure 6 ), voices / sounds from other persons included in the input audio data may be isolated and discarded before further processing. The device 110a may communicate with the device 110b and may send input audio data to the device 110b. Figure 1B Device 110a is shown as a smartwatch, however, device 110a may be any wearable device or any device carried by user 5 and configured to capture audio data when appropriate user permissions are met. Device 110b is shown as a smart phone, however, device 110b may be any mobile device or computing device that communicates with device 110a and is configured to receive data from and send data to device 110a, such as a laptop, tablet, desktop, etc. Alternatively, device 110a may be a voice-activated system or smart speaker and may send input audio data directly to one or more systems 120 rather than forwarding it through device 110b. Alternatively, the operations of devices 110a and 110b may be combined into a single device. Figure 1A The device 110 used may be different from the device 110a used during operation, so the user 5 may register in emotion detection using a device different from the device used to provide input audio.
[0040] The one or more systems 120 identify (152) reference audio data representing a baseline associated with the user profile of user 5. The one or more systems 120 may retrieve the reference audio data from the profile storage device 270.
[0041] As described above, in some embodiments, profile storage device 270 may store multiple baselines for user 5, each of which may correspond to a different context / situation. One or more systems 120 may identify a baseline from the multiple baselines associated with the user profile based on context data associated with the baseline and context data associated with the input audio data. One or more systems 120 may determine context data corresponding to the input audio data, such as the user's location (e.g., the location of device 110a), the person with whom he / she is interacting, and so on. One or more systems 120 may select a baseline with context data similar to the context data of the input audio data, thereby using an appropriate baseline to account for the different speaking styles / acoustic speech attributes exhibited by users in different situations. In other embodiments, one or more systems 120 may analyze features of reference audio data corresponding to the baseline and the input audio data (e.g., using an ML model, statistical analysis, or other methods) to identify a baseline with features similar to the input audio data. In some embodiments, if one or more systems 120 cannot identify a baseline with context data similar to the context data of the input audio data, one or more systems 120 may select the best available baseline based on the quality of the baseline (e.g., audio quality, quality of acoustic features, best representation of a neutral emotional state, etc.).
[0042] The one or more systems 120 may then determine (154) a first feature vector corresponding to the reference audio data, if this operation has not already been performed during the enrollment process (operation 142). The first feature vector may represent spectral features derived from the reference audio data. The one or more systems 120 may use an encoder (e.g., Figure 11 The encoder 1150 of the embodiment of the present invention processes the frame of the reference audio data and generates a first feature vector. In some embodiments, the first feature vector may represent the acoustic speech attributes (e.g., accent, pitch, prosody, etc.) exhibited by the user 5 in a neutral emotional state.
[0043] The one or more systems 120 determine (156) a second feature vector corresponding to the input audio data. The second feature vector may represent spectral features derived from the input audio data. The one or more systems 120 may use an encoder (e.g., Figure 11 The encoder 1150 of the input audio data is used to process the frames of the input audio data and generate a second feature vector. In some embodiments, the second feature vector may represent the acoustic speech attributes (e.g., accent, pitch, prosody, etc.) exhibited by the user 5 when speaking the speech represented by the input audio data.
[0044] The one or more systems 120 process (158) the first feature vector and the second feature vector using a training model. The training model may output one or more scores. The one or more systems 120 determine (160) an emotion category based on the scores generated by the training model. The training model may be an ML model configured to process features of the reference audio data and the input audio data to determine an emotion category corresponding to the input audio data based on a neutral emotional state of the user (represented by the reference audio data). The emotion categories may include broad categories such as positive, neutral, and negative. In other embodiments, the emotion categories may be more specific and may include, for example, anger, happiness, sadness, and neutral. In another embodiment, the emotion categories may include anger, sadness, happiness, surprise, stress, and disgust. As can be appreciated, various emotion categories / indicators are possible depending on the system configuration.
[0045] In some embodiments, one or more systems 120 may determine that the input audio data represents voice activity from a human. One or more systems 120 may identify a voice profile associated with a user profile of device 110. One or more systems 120 may retrieve stored data associated with the user profile. The stored data may include a voice fingerprint or voice biomarker to identify the user using the audio data. In other embodiments, the stored data may include RF data, location data, machine vision data, etc. as described in connection with user identification component 295. One or more systems 120 may identify the voice profile using user identification component 295 as described herein.
[0046] The one or more systems 120 may determine a first portion of the input audio data, wherein the first portion corresponds to a speech profile. For example, the input audio data may capture speech from multiple people, particularly if user 5 is conversing with another person. The one or more systems 120 may isolate the first portion of the input audio data associated with the speech spoken by user 5 and store the first portion for further analysis. The one or more systems 120 may use the first portion of the input audio data to determine a feature vector (in operation 156).
[0047] One or more systems 120 may store association data that associates emotion categories with input audio data and user profiles. In an exemplary embodiment, one or more systems 120 may analyze the input audio data over a certain period of time and determine emotion categories at different time intervals to provide information about the user's emotional state during the period of time or when interacting with others. In another embodiment, one or more systems 120 may analyze the input audio data while the user is interacting with device 110, and the emotion category may indicate the user's satisfaction with the interaction with device 110.
[0048] The one or more systems 120 generate (162) output data that includes at least the emotion category and a portion of the input audio data. The one or more systems 120 may determine text data corresponding to the audio data frame using the ASR processing techniques described below. The one or more systems 120 may also determine time data indicating when the portion of the input audio data was received by the device 110. The output data may include text data corresponding to the portion of the input audio data, time data, and an indicator of the emotion category. The output data may be displayed on the device 110a or the device 110b. The indicator of the emotion category may be text representing the emotion category, an icon representing the emotion category, or other indicator.
[0049] Figure 1B The operations of are generally described herein as being performed by one or more systems 120. However, it should be understood that one or more of the operations may also be performed by device 110a, device 110b, or other devices. Figure 6 Provide a description.
[0050] The entire system of the present disclosure can be operated using various components as shown below. The various components can be located on the same or different physical devices. Communication between the various components can occur directly or across one or more networks 199.
[0051] like Figure 2A and Figure 2B As shown, one or more audio capture components (such as a microphone or microphone array of device 110) capture audio 11. Device 110 processes audio data representing audio 11 to determine whether speech is detected. Device 110 may use various techniques to determine whether the audio data includes speech. In some examples, device 110 may apply voice activity detection (VAD) techniques. Such techniques may determine whether speech is present in the audio data based on various quantitative aspects of the audio data (such as the spectral slope between one or more frames of audio data; the energy level of the audio data in one or more spectral bands; the signal-to-noise ratio of the audio data in one or more spectral bands; or other quantitative aspects). In other examples, device 110 may implement a finite classifier configured to distinguish speech from background noise. The classifier may be implemented using techniques such as linear classifiers, support vector machines, and decision trees. In yet other examples, device 110 may apply hidden Markov model (HMM) or Gaussian mixture model (GMM) techniques to compare the audio data with one or more acoustic models in a storage device, which may include models corresponding to speech, noise (e.g., ambient noise or background noise), or silence. Still other techniques may be used to determine whether speech is present in the audio data.
[0052] When speech is detected in the audio data representing audio 11, device 110 may use wake word detection component 220 to perform wake word detection to determine when a user intends to speak input to device 110. One exemplary wake word is "Alexa."
[0053] Wake word detection is typically performed without performing linguistic analysis, textual analysis, or semantic analysis. Instead, audio data representing audio 11 is analyzed to determine whether specific characteristics of the audio data match a preconfigured acoustic waveform, audio signature, or other data, thereby determining whether the audio data "matches" stored audio data corresponding to the wake word.
[0054] Therefore, the wake-up word detection component 220 can compare the audio data with a stored model or data to detect the wake-up word. One method for wake-up word detection applies a general large vocabulary continuous speech recognition (LVCSR) system to decode the audio signal, where the wake-up word search is performed in the resulting lattice or confusion network. LVCSR decoding may require relatively high computational resources. Another method for wake-up word detection constructs an HMM for each wake-up word speech signal and non-wake-up word speech signal. Non-wake-up word speech includes other spoken words, background noise, etc. One or more HMMs can be constructed to model the characteristics of non-wake-up word speech, which is called a filler model. Viterbi decoding is used to search for the best path in the decoding graph, and the decoded output is further processed to determine the presence of the wake-up word. This method can be extended to include discriminant information by incorporating a hybrid DNN-HMM decoding framework. In another example, the wake-up word detection component 220 can be directly built on a deep neural network (DNN) / recurrent neural network (RNN) structure without involving an HMM. This architecture can estimate the wake-up word posteriors with contextual information by stacking frames within a background window of the DNN or using an RNN. Subsequent a posteriori threshold adjustment or smoothing is used for decision making.Other techniques for wake-up word detection, such as those known in the art, may also be used.
[0055] When the wake word is detected, the device 110 may "wake up" and begin transmitting audio data 211 representing the audio 11 to one or more systems 120, such as Figure 2A As shown. Figure 2BAs shown, device 110a may transmit audio data 211 to device 110b, and device 110b may transmit audio data 211 to one or more systems 120. The audio data 211 may include data corresponding to a wake word, or the portion of the audio corresponding to the wake word may be removed by device 110a before sending the audio data 211 to the one or more systems 120. In some embodiments, device 110a may begin transmitting audio data 211 to one or more systems 120 / device 110b (or otherwise perform additional processing on the audio data) in response to an event occurring or an event detected by device 110a.
[0056] Upon receipt by one or more systems 120, audio data 211 may be sent to coordinator component 230. Coordinator component 230 may include memory and logic that enables coordinator component 230 to transmit various pieces and forms of data to various components of the system and perform other operations as described herein.
[0057] Coordinator component 230 sends audio data 211 to speech processing component 240. ASR component 250 of speech processing component 240 transcribes input audio data 211 into input text data representing one or more hypotheses representing the speech contained in input audio data 211. Therefore, the text data output by ASR component 250 may represent one or more (e.g., in the form of an N-best list) ASR hypotheses representing the speech represented in audio data 211. ASR component 250 interprets the speech in audio data 211 based on similarities between the audio data 211 and pre-established language models. For example, ASR component 250 may compare audio data 211 with acoustic models (e.g., sub-word units such as phonemes) and sound sequences to identify words that match the sound sequences of the speech represented in audio data 211. ASR component 250 outputs text data representing one or more ASR hypotheses. ASR component 250 may also output corresponding scores for the one or more ASR hypotheses. For example, ASR component 250 may output such text data and scores after operating on a language model. Thus, the text data output by the ASR component 250 may include the highest-scoring ASR hypotheses or may include an N-best list of ASR hypotheses. The N-best list may additionally include a corresponding score associated with each ASR hypothesis represented therein. Each score may indicate the confidence level of the ASR processing performed to generate the ASR hypothesis associated with the score. Further details of the ASR processing are included below.
[0058] NLU component 260 receives one or more ASR hypotheses (i.e., text data) and attempts to semantically interpret one or more phrases or one or more sentences represented therein. Specifically, NLU component 260 determines one or more meanings associated with the one or more phrases or one or more sentences represented in the text data based on the words represented in the text data. NLU component 260 determines a segment of the text data that represents an intent for an action desired by a user and allows a device (e.g., device 110, one or more systems 120, skill 290, one or more skill systems 225, etc.) to execute the intent. For example, if the text data corresponds to "play Adele music," NLU component 260 may determine that one or more systems 120 intend to output music and may identify "Adele" as the artist. For another example, if the text data corresponds to "how's the weather," NLU component 260 may determine that one or more systems 120 intend to output weather information associated with the geographic location of device 110. In another example, if the text data corresponds to "turn off the lights," NLU component 260 may determine that one or more systems 120 intend to turn off the lights associated with one or more devices 110 or one or more users 5.
[0059] The NLU component 260 can send the NLU result data (which can include tagged text data, intent indicators, etc.) to the coordinator component 230. The coordinator component 230 can send the NLU result data to one or more skills 290. If the NLU result data includes a single NLU hypothesis, the coordinator component 230 can send the NLU result data to the one or more skills 290 associated with the NLU hypothesis. If the NLU result data includes an N-best list of NLU hypotheses, the coordinator component 230 can send the highest-scoring NLU hypothesis to the one or more skills 290 associated with the highest-scoring NLU hypothesis.
[0060] A "skill" can be software running on one or more systems 120, similar to software applications running on traditional computing devices. That is, a skill 290 can enable one or more systems 120 to perform a specific function to provide data or generate some other requested output. One or more systems 120 can be configured with more than one skill 290. For example, a weather service skill can enable one or more systems 120 to provide weather information, a car service skill can enable one or more systems 120 to book a ride with a taxi or ride-sharing service, and a restaurant skill can enable one or more systems 120 to order a pizza with a restaurant's online ordering system. Skills 290 can operate in coordination between one or more systems 120 and other devices (such as device 110) to perform specific functions. Input to a skill 290 can be obtained from speech processing interactions or through other interactions or input sources. A skill 290 can include hardware, software, firmware, etc. that can be dedicated to a specific skill 290 or shared among different skills 290.
[0061] In addition to or as an alternative to being implemented by one or more systems 120, skills 290 may be implemented by one or more skill systems 225. This may enable one or more skill systems 225 to perform specific functions in order to provide data or perform some other action requested by a user.
[0062] Skill types include home automation skills (e.g., skills that enable a user to control home devices such as lights, door locks, cameras, thermostats, etc.), entertainment device skills (e.g., skills that enable a user to control entertainment devices such as smart TVs), video skills, flash presentation skills, and custom skills that are not related to any pre-configured type of skill.
[0063] One or more systems 120 may be configured with a single skill 290 that is dedicated to interacting with more than one skill system 225 .
[0064] Unless explicitly stated otherwise, references to skills, skill devices, or skill components may include skills 290 operated by one or more systems 120 and / or skills operated by one or more skill systems 225. Furthermore, the functionality described herein as skills may be referred to using many different terms, such as actions, robots, applications, etc.
[0065] One or more systems 120 may include a TTS component 280 that generates audio data (e.g., synthesized speech) from text data using one or more different methods. The text data input to the TTS component 280 may be obtained from a skill 290, a coordinator component 230, or another component of the one or more systems 120.
[0066] In one synthesis method, known as unit selection, the TTS component 280 matches the text data against a database of recorded speech. The TTS component 280 selects matching units of the recorded speech and concatenates these units to form audio data. In another synthesis method, known as parametric synthesis, the TTS component 280 varies parameters such as frequency, volume, and noise to create audio data comprising an artificial speech waveform. Parametric synthesis uses a computerized speech generator, sometimes called a vocoder.
[0067] One or more systems 120 may include a profile storage 270. The profile storage 270 may include various information related to individual users, groups of users, devices, etc., that interact with the one or more systems 120. A "profile" refers to a set of data associated with a user, device, etc. The profile data may include preferences specific to the user, device, etc.; input and output capabilities of the device; internet connection information; user bibliographic information; subscription information; and other information.
[0068] Profile storage device 270 may include one or more user profiles, where each user profile is associated with a different user identifier. Each user profile may include various user identification information. Each user profile may also include user preferences and / or one or more device identifiers, which represent one or more devices registered to the user.
[0069] The profile storage device 270 may include one or more group profiles. Each group profile may be associated with a different group profile identifier. A group profile may be specific to a user group. That is, a group profile may be associated with two or more individual user profiles. For example, a group profile may be a family profile associated with user profiles associated with multiple users of a single family. A group profile may include preferences shared by all user profiles associated with it. Each user profile associated with a group profile may additionally include preferences specific to the user associated with it. That is, each user profile may include preferences that are unique relative to one or more other user profiles associated with the same group profile. A user profile may be an independent profile or may be associated with a group profile. A group profile may include one or more device profiles representing one or more devices associated with the group profile.
[0070] Profile storage 270 may include one or more device profiles. Each device profile may be associated with a different device identifier. Each device profile may include various device identification information. Each device profile may also include one or more user identifiers representing one or more user profiles associated with the device profile. For example, a profile for a family of devices may include user identifiers for family users.
[0071] Profile storage 270 may include audio data representing one or more baselines corresponding to a user's neutral emotional state. Profile storage 270 may include data related to multiple baselines, each baseline associated with different contextual data.
[0072] One or more systems 120 may also include an emotion detection component 275 that may be configured to detect the user's emotion from audio data representing speech / utterance from the user. The emotion detection component 275 may be included in the speech processing component 240 or may be, for example, Figure 2A Individual components are shown. Emotion detection component 275 and other components are generally described as being operated by one or more systems 120. However, device 110 may also operate one or more of the components, including emotion detection component 275.
[0073] The system can be configured to incorporate user permissions and can perform the activities disclosed herein only with the user's approval. Thus, the systems, devices, components, and techniques described herein will generally be configured to limit processing where appropriate and only process user information in a manner that ensures compliance with all applicable laws, regulations, standards, and the like. The systems and techniques can be implemented on a geographic basis to ensure compliance with the laws of various jurisdictions and entities where system components and / or users are located. A user can delete any data stored in the profile storage device 270, for example, data related to one or more baselines (baseline data), emotion detection, and the like.
[0074] One or more systems 120 may include a user identification component 295 that uses various data to identify one or more users. Figure 3As shown, the user identification component 295 may include one or more subcomponents, including a visual component 308, an audio component 310, an identification component 312, a radio frequency (RF) component 314, a machine learning (ML) component 316, and an identification confidence component 318. In some cases, the user identification component 295 may monitor data from one or more subcomponents and determine the identities of one or more users associated with data input to the one or more systems 120. The user identification component 295 may output user identification data 395, which may include a user identifier associated with the user that the user identification component 295 believes to be the origin of the data input to the one or more systems 120. The user identification data 395 may be used to inform processes performed by various components of the one or more systems 120.
[0075] The visual component 308 may receive data from one or more sensors capable of providing images (e.g., a camera) or sensors indicating motion (e.g., a motion sensor). The visual component 308 may perform facial recognition or image analysis to determine the identity of the user and associate that identity with a user profile associated with the user. In some cases, when the user is facing the camera, the visual component 308 may perform facial recognition and identify the user with a high degree of confidence. In other cases, the visual component 308 may have a low degree of confidence in the user's identity, and the user identification component 295 may utilize determinations from other components to determine the user's identity. The visual component 308 may be used in conjunction with other components to determine the user's identity. For example, the user identification component 295 may use data from the visual component 308 in conjunction with data from the audio component 310 to identify what the user's face appears to be saying while the device 110 that the user is facing is capturing audio, for the purpose of identifying the user speaking input to one or more systems 120.
[0076] The overall system 100 of the present disclosure may include a biometric sensor that transmits data to an identification component 312. For example, the identification component 312 may receive data corresponding to a fingerprint, an iris or retina scan, a thermal scan, a user's weight, a user's size, pressure (e.g., within a floor sensor), etc., and may determine a profile corresponding to the user. For example, the identification component 312 may distinguish between a user and the sound from a television. Thus, the identification component 312 may incorporate identification information into a confidence level for determining the identity of the user. The identification information output by the identification component 312 may be associated with specific user profile data so that the identification information uniquely identifies the user's user profile.
[0077] RF component 314 can use RF positioning to track devices that can be carried or worn by users. For example, users (and user profiles associated with users) can be associated with devices. Devices can emit RF signals (e.g., Wi-Fi, The device may detect the signal and indicate the strength of the signal to RF component 314 (e.g., as a received signal strength indication (RSSI)). RF component 314 may use the RSSI to determine the identity of the user (with an associated confidence level). In some cases, RF component 314 may determine that the received RF signal is associated with a mobile device associated with a specific user identifier.
[0078] In some cases, device 110 may include some RF or other detection processing capabilities so that a user speaking an input may scan, tap, or otherwise identify his / her personal device (such as a phone) to device 110. In this way, a user may "register" with one or more systems 120 for the purpose of one or more systems 120 determining who spoke a particular input. Such registration may occur before, during, or after the input is spoken.
[0079] ML component 316 can track various user behaviors as a factor in determining the confidence level of a user's identity. For example, a user may adhere to a regular schedule that causes the user to be in a first location (e.g., at work or at school) during the day. In this example, ML component 316 will consider past behaviors and / or trends when determining the identity of a user providing input to one or more systems 120. Thus, ML component 316 can use historical data and / or usage patterns over time to increase or decrease the confidence level of a user's identity.
[0080] In at least some cases, identification confidence component 318 receives determinations from various components 308, 310, 312, 314, and 316 and can determine a final confidence level associated with the user's identity. In some cases, the confidence level can determine whether to perform an action in response to the user input. For example, if the user input includes a request to unlock a door, the confidence level may need to meet or exceed a threshold, which may be higher than a threshold confidence level required to perform a user request associated with playing a playlist or sending a message. The confidence level or other score data may be included in user identification data 395.
[0081] The audio component 310 may receive data from one or more sensors (e.g., one or more microphones) capable of providing audio signals to facilitate identification of the user. The audio component 310 may perform audio recognition on the audio signals to determine the identity of the user and an associated user identifier. In some cases, aspects of one or more systems 120 may be configured at a computing device (e.g., a local server). Thus, in some cases, the audio component 310 operating on the computing device may analyze all sounds to facilitate identification of the user. In some cases, the audio component 310 may perform voice recognition to determine the identity of the user.
[0082] The audio component 310 may also perform user identification based on the audio data 211 input into the one or more systems 120 for speech processing. The audio component 310 may determine a score indicating whether the speech in the audio data 211 originates from a particular user. For example, a first score may indicate a likelihood that the speech in the audio data 211 originates from a first user associated with a first user identifier, a second score may indicate a likelihood that the speech in the audio data 211 originates from a second user associated with a second user identifier, and so on. The audio component 310 may perform user identification by comparing speech characteristics represented in the audio data 211 with stored speech characteristics of the user (e.g., a stored voice profile associated with the device 110 that captured the spoken user input).
[0083] like Figure 2B As shown, emotion detection component 275 and user identification component 295 may be included in device 110b. Device 110a may transmit audio data 211 to device 110b. Upon receipt, device 110b may send audio data 211 to user identification component 295 to perform the operations described herein with respect to component 295, for example, including identifying a user profile corresponding to audio data 211. User identification component 295 may send data to emotion detection component 275 to perform the operations described herein.
[0084] Figure 4 4. ASR component 250 performs ASR processing on ASR feature vector data 450. ASR confidence data 407 may be passed to user identification component 295.
[0085] The user identification component 295 performs user identification using various data, including user identification feature vector data 440, feature vectors 405 representing voice profiles of users of one or more systems 120, ASR confidence data 407, and other data 409. The user identification component 295 can output user identification data 395 that reflects a specific confidence level that the user input is spoken by one or more specific users. The user identification data 395 can include one or more user identifiers (e.g., corresponding to one or more voice profiles). Each user identifier in the user identification data 395 can be associated with a corresponding confidence value representing the likelihood that the user input corresponds to the user identifier. The confidence value can be a number or a binned value.
[0086] The one or more feature vectors 405 input to the user identification component 295 may correspond to one or more voice profiles. The user identification component 295 may use the one or more feature vectors 405 to compare with a user identification feature vector 440 representing the current user input to determine whether the user identification feature vector 440 corresponds to one or more feature vectors in the feature vectors 405 of the voice profile. Each feature vector 405 may be the same size as the user identification feature vector 440.
[0087] To perform user identification, user identification component 295 may determine the device 110 from which audio data 211 originated. For example, audio data 211 may be associated with metadata including a device identifier representing device 110. Device 110 or one or more systems 120 may generate the metadata. One or more systems 120 may determine a group profile identifier associated with the device identifier, may determine a user identifier associated with the group profile identifier, and may include the group profile identifier and / or user identifier in the metadata. One or more systems 120 may associate the metadata with a user identification feature vector 440 generated from audio data 211. User identification component 295 may send a signal to voice profile storage device 485 requesting only audio data and / or feature vectors 405 associated with the device identifier, group profile identifier, and / or user identifier represented in the metadata (depending on whether the audio data and / or corresponding feature vectors are stored). This limits the range of possible feature vectors 405 considered by user identification component 295 at runtime, thereby reducing the amount of time required to perform user identification processing by reducing the number of feature vectors 405 that need to be processed. Alternatively, the user identification component 295 may access all (or some other subset) of the audio data and / or feature vectors 405 available to the user identification component 295. However, accessing all of the audio data and / or feature vectors 405 will likely increase the amount of time required to perform the user identification process based on the amount of audio data and / or feature vectors 405 to be processed.
[0088] If the user identification component 295 receives audio data from the voice profile storage device 485, the user identification component 295 may generate one or more feature vectors 405 corresponding to the received audio data.
[0089] The user identification component 295 may attempt to identify the user who spoke the speech represented in the audio data 211 by comparing the user identification feature vector 440 to one or more feature vectors 405. The user identification component 295 may include a scoring component 422 that determines a corresponding score indicating whether the user input (represented by the user identification feature vector 440) was spoken by one or more specific users (represented by one or more feature vectors 405). The user identification component 295 may also include a confidence component 424 that determines the overall accuracy of the user identification process (such as that of the scoring component 422) and / or an individual confidence value for each user that may be identified by the scoring component 422. The output from the scoring component 422 may include a different confidence value for each received feature vector 405. For example, the output may include a first confidence value for a first feature vector 405a (representing a first voice profile), a second confidence value for a second feature vector 405b (representing a second voice profile), and so on. Although shown as two separate components, the score component 422 and the confidence component 424 can be combined into a single component or can be separated into more than two components.
[0090] Scoring component 422 and confidence component 424 can implement one or more training machine learning models known in the art (such as neural networks, classifiers, etc.). For example, scoring component 422 can use probabilistic linear discriminant analysis (PLDA) technology. PLDA scoring determines how likely it is that a user identification feature vector 440 corresponds to a specific feature vector 405. PLDA scoring can generate a confidence value for each feature vector 405 considered and can output a list of confidence values associated with the corresponding user identifier. Scoring component 422 can also use other techniques (such as GMM, generating Bayesian models, etc.) to determine confidence values.
[0091] Confidence component 424 may input various data (including information about ASR confidence 407, speech length (e.g., number of frames or other measured length entered by the user), audio condition / quality data (such as signal interference data or other metric data), fingerprint data, image data, or other factors) to consider how confident user identification component 295 is about linking the user to the confidence value entered by the user. Confidence component 424 may also consider the confidence values and associated identifiers output by scoring component 422. For example, confidence component 424 may determine that a lower ASR confidence 407, or poorer audio quality, or other factors may result in a lower confidence for user identification component 295. Whereas, a higher ASR confidence 407, or better audio quality, or other factors may result in a higher confidence for user identification component 295. The precise determination of confidence may depend on the configuration and training of confidence component 424 and the one or more models implemented thereby. Confidence component 424 may operate using a number of different machine learning models / techniques (such as GMM, neural networks, etc.). For example, confidence component 424 can be a classifier configured to map the scores output by scoring component 422 to confidence values.
[0092] The user identification component 295 can output user identification data 395 specific to one or more user identifiers. For example, the user identification component 295 can output user identification data 395 relative to each received feature vector 405. The user identification data 395 can include a numerical confidence value (e.g., 0.0-1.0, 0-1000, or any scale that the system is configured to operate). Thus, the user identification data 395 can output an n-best list of potential users with numerical confidence values (e.g., user identifier 123 -0.2, user identifier 234 -0.8). Alternatively or additionally, the user identification data 395 can include binned confidence values. For example, a calculated recognition score for a first range (e.g., 0.0-0.33) can be output as "low," a calculated recognition score for a second range (e.g., 0.34-0.66) can be output as "medium," and a calculated recognition score for a third range (e.g., 0.67-1.0) can be output as "high." The user identification component 295 may output an n-best list of user identifiers with binned confidence values (e.g., user identifier 123 - low, user identifier 234 - high). A combined binning and numeric confidence value output is also possible. The user identification data 395 may include only information related to the highest-scoring identifiers determined by the user identification component 295, without including a list of identifiers and their corresponding confidence values. The user identification component 295 may also output an overall confidence value for the individual confidence values being correct, where the overall confidence value indicates how confident the user identification component 295 is in the output result. The confidence component 424 may determine the overall confidence value.
[0093] Confidence component 424 can determine differences between individual confidence values when determining user identification data 395. For example, if the difference between a first confidence value and a second confidence value is large, and the first confidence value is above a threshold confidence value, user identification component 295 can identify the first user (associated with feature vector 405 associated with the first confidence value) as the user who spoke the user input with a higher confidence than if the difference between the confidence values is smaller.
[0094] User identification component 295 can perform threshold processing to avoid outputting incorrect user identification data 395.For example, user identification component 295 can compare the confidence value outputted by confidence component 424 with the threshold confidence value.If the confidence value does not meet (for example, does not meet or exceeds) threshold confidence value, then user identification component 295 can not output user identification data 395, or can only include the indicator of the user who can't identify the user input in these data 395.In addition, user identification component 295 can not output user identification data 395, until enough user identification feature vector data 440 are accumulated and processed to verify that the user is higher than the threshold confidence value.Therefore, user identification component 295 can wait until the enough threshold amount of the audio data of the user input has been processed before outputting user identification data 395.Confidence component 424 also can consider the amount of received audio data.
[0095] The user identification component 295 may default to outputting binned (e.g., low, medium, high) user identification confidence values. However, this may be problematic in certain situations. For example, if the user identification component 295 calculates a single binned confidence value for multiple feature vectors 405, the system may be unable to determine which specific user is the origin of the user input. In this case, the user identification component 295 may override its default and output a numeric confidence value. This allows the system to determine that the user associated with the highest numeric confidence value is the origin of the user input.
[0096] User identification component 295 can use other data 409 to inform the user identification process. One or more training models or other components of user identification component 295 can be trained to use other data 409 as input features when performing user identification processing. Other data 409 can include a variety of data types depending on the system configuration and can be obtained from other sensors, devices, or storage devices. Other data 409 can include the time of day when audio data 211 is generated or received from device 110, the day of the week when audio data 211 is generated or received from device 110, and the like.
[0097] Other data 409 may include image data or video data. For example, facial recognition may be performed on the image data or video data received from device 110 (or another device) that received audio data 211. Facial recognition may be performed by user identification component 295. The output of the facial recognition process may be used by user identification component 295. That is, the facial recognition output data may be used in conjunction with a comparison of user identification feature vector 440 with one or more feature vectors 405 to perform a more accurate user identification process.
[0098] Other data 409 may include location data for device 110. The location data may be specific to the building in which device 110 is located. For example, if device 110 is located in user A's bedroom, such location may increase the user identification confidence value associated with user A and / or decrease the user identification confidence value associated with user B.
[0099] Other data 409 may include data indicating the type of device 110. Different types of devices may include, for example, smart watches, smart phones, tablets, and vehicles. The type of device 110 may be indicated in a configuration file associated with the device 110. For example, if the device 110 from which the audio data 211 is received is a smart watch or a vehicle belonging to user A, the fact that the device 110 belongs to user A may increase the user identification confidence value associated with user A and / or decrease the user identification confidence value associated with user B.
[0100] Other data 409 may include geographic coordinate data associated with device 110. For example, a group profile associated with a vehicle may indicate multiple users (e.g., user A and user B). The vehicle may include a global positioning system (GPS) that indicates the latitude and longitude coordinates of the vehicle when the vehicle generates audio data 211. Thus, if the vehicle is located at coordinates corresponding to user A's work location / building, this may increase the user identification confidence value associated with user A and / or decrease the user identification confidence values of all other users indicated in the group profile associated with the vehicle. A profile associated with device 110 may indicate global coordinates and an associated location (e.g., work, home, etc.). One or more user profiles may also or alternatively indicate global coordinates.
[0101] Other data 409 may include data representing activities of a specific user that can be used to perform user identification processing. For example, a user may have recently entered a code to disable a home security alarm. Devices 110 represented in the group profile associated with the home may have generated audio data 211. Other data 409 may reflect signals from the home security alarm regarding the disabled user, the disabled time, etc. If a mobile device (such as a smartphone, Tile tracker, dongle, or other device known to be associated with a specific user is detected as being in proximity to device 110 (e.g., physically close to the device, connected to the same WiFi network as the device, or otherwise located near the device), this may be reflected in other data 409 and considered by user identification component 295.
[0102] Depending on the system configuration, the other data 409 may be configured to be included in the user identification feature vector data 440 so that all data related to the user input to be processed by the scoring component 422 can be included in a single feature vector. Alternatively, the other data 409 may be reflected in one or more different data structures to be processed by the scoring component 422.
[0103] Figure 5 is a conceptual diagram illustrating an emotion detection component including components for user registration according to an embodiment of the present disclosure. In some embodiments, the emotion detection component 275 may include a registration component 505 and a background component 515.
[0104] Registration component 505 may be configured to obtain audio data from the user representing the user's neutral emotional state. Registration component 505 may be configured to cause device 110 to request the user to speak one or more sentences. For example, registration component 505 may cause device 110 to output "For registration purposes, please say I like today's weather," and the user may say "I like today's weather," which may be represented by audio data 211. Registration component 505 may process audio data 211 representing reference utterances spoken by the user. In some cases, audio data 211 may include multiple utterances, and reference audio data 510 may correspond to multiple utterances.
[0105] The registration component 505 may also be configured to determine whether the audio data 211 can be used as a baseline for representing the user's neutral emotional state. If the audio data 211 is determined to be a good / valid baseline, the registration component 505 may store the audio data 211 as reference audio data 510 in the profile storage device 270 and associate the reference audio data 510 with the user's profile as a baseline for emotion detection.
[0106] Registration component 505 can analyze audio data 211 to determine whether corresponding acoustic speech attributes are within a predefined range or meet specific conditions indicating a neutral emotional state for the user. As used herein, acoustic speech attributes refer to features such as accent, pitch, prosody (intonation, tone, stress, rhythm), voice, etc. that can be derived from audio data. Registration component 505 can have identified and stored acoustic speech attributes that indicate a neutral emotional state based on analyzing audio data from multiple users representing a general population or a specific population (to account for accents, cultural differences, and other factors that affect speech based on geographic location), and can use these acoustic speech attributes to determine whether audio data 211 indicates a neutral emotional state for the user.
[0107] In some embodiments, the enrollment component 505 may employ an ML model to process the audio data 211 to determine an emotion category corresponding to the audio data. If the ML model determines that the emotion category corresponding to the audio data 211 is neutral, the enrollment component 505 may store the audio data 211 as reference audio data 510. If the ML model determines that the emotion category corresponding to the audio data 211 is not neutral (angry, happy, etc.), the audio data 211 may be discarded and not used as a baseline for emotion detection. The audio data 211 may be input to an encoder (not shown) to determine one or more frame feature vectors (not shown). The one or more frame feature vectors may represent audio frame-level features extracted from the audio data 211. A frame feature vector may represent audio frame-level features of a 20ms audio frame of the audio data 211. The one or more frame feature vectors may be derived through spectral analysis of the audio data 211. In an exemplary embodiment, the emotion component 275 may determine that the audio data 211 includes an entire utterance, and the one or more frame feature vectors may be used to determine one or more utterance feature vectors representing utterance-level features of the one or more utterances represented in the audio data 211. One or more utterance feature vectors may be determined by performing statistical calculations, incremental calculations, and other processing on one or more frame feature vectors of an audio frame corresponding to an utterance of interest. An ML model (not shown) employed by enrollment component 505 may process the one or more frame feature vectors to determine one or more scores indicating the user's emotion when uttering the utterance represented by the one or more frame feature vectors. In another embodiment, the ML model may process utterance-level feature vectors to determine one or more scores indicating the user's emotion when uttering the utterance represented by the one or more frame feature vectors. The ML model may be trained using a training dataset to process audio frame features and / or utterance-level features to determine the user's emotion. In some embodiments, the ML model may be trained to output a score indicating a confidence level of neutrality for the user's emotion. For example, a score of 1-2 may indicate a low confidence level, a score of 3 may indicate a medium confidence level, and a score of 4-5 may indicate a high confidence level. In other embodiments, the ML model may be trained to output a low, medium, or high indication of a neutral emotion category. In exemplary embodiments, the ML model may be a neural network machine learning model (recurrent neural network, deep learning neural network, convolutional neural network, etc.), a statistical model, a probabilistic model, or another type of model.
[0108] Registration component 505 can be configured to: if audio data 211 does not represent a good baseline for emotion detection, then request the user to repeat a sentence or say another sentence. Registration component 505 can cause device 110 to output, for example, "Please repeat that I like today's weather." Registration component 505 can process the audio data received from the user in response to determine whether the audio data can be used as a baseline. In some embodiments, registration component 505 can only make a few attempts to obtain audio data for a baseline. After trying two or three times and not being able to obtain data that can be used for a baseline, registration component 505 can cause device 110 to output audio to inform the user that the system will not continue the registration process and the user should try again at another time. Audio data 211 may be of poor quality due to background noise, or audio data 211 may not represent the user's neutral emotional state (e.g., the user may be too excited or angry during the registration process).
[0109] In some embodiments, registration component 505 may request the user to say a specific sentence. In other embodiments, registration component 505 may request the user to talk about a topic rather than saying a specific sentence. In some embodiments, registration component 505 may request the user to say a specific sentence and also discuss a topic to capture audio data for both situations, because the user may exhibit different speaking styles / acoustic voice attributes when repeating a sentence relative to when freely discussing a topic. Registration component 505 may process the audio data representing the user saying a specific sentence and the audio data representing the user freely discussing a topic to, for example, use the difference in the acoustic voice attributes of the two situations, the (weighted or unweighted) average value of the acoustic voice attributes of the two situations, statistical analysis, a machine learning model to process corresponding feature vectors and / or use other methods to determine an appropriate baseline.
[0110] Emotion detection component 275 can be configured to obtain reference audio data (for multiple baselines) from the user in different situations. Doing so allows the system to account for the different speaking styles / acoustic speech attributes exhibited by users in different situations. Context component 515 can be configured to determine data (e.g., context data 520) representing the user's environment, situation, location, context, or other contextual data corresponding to the user when he or she spoke the audio used for the baseline. For example, context component 515 can determine the user's location when the reference utterance was spoken by using the location of device 110 or other information associated with the user's profile. Context component 515 can determine the type of interaction, including who the user was interacting with when the utterance was spoken, the context in which the user was speaking (e.g., a work meeting, a family / friends gathering, a sporting event, a concert, etc.), the time of day (e.g., morning, afternoon, evening, day of the week, etc.), any actions the user was performing when speaking (e.g., driving, walking, watching TV, etc.), and the like. Contextual data 520 may also include data representing other contextual information corresponding to when the user uttered the audio, such as weather information, physiological data associated with the user (e.g., heart rate, blood pressure, body temperature, etc.), season of the year, month of the year, etc. Context component 515 may determine contextual data 520 by retrieving data from user profile storage 270, other data storage devices, and / or other systems / applications. Context component 515 may derive contextual data 520 by processing audio data and determining, based on the audio data, characteristics or features that indicate specific contextual data. In some embodiments, the system may receive input data from the user that indicates contextual data, such as the user's location (e.g., home, work, gym, etc.), who the user is interacting with (e.g., coworkers, boss, partner / significant other, children, neighbors, etc.), the context in which the user is located (e.g., work meeting, social gathering, etc.), the action the user is performing (e.g., driving, walking, etc.), etc.
[0111] The emotion detection component 275 may store a plurality of baselines and corresponding contextual data in the profile storage device 270. For example, the emotion detection component 275 may store in the profile storage device 270 first audio data (e.g., 510a) representing a first baseline and contextual data (e.g., 520a) indicating <location: work place>, second audio data (e.g., 510b) representing a second baseline and contextual data (e.g., 520b) indicating <location: home>, third audio data (e.g., 510c) representing a third baseline and contextual data (e.g., 520c) indicating <person: coworker>, fourth audio data (e.g., 510d) representing a fourth baseline and contextual data (e.g., 520d) indicating <person: daughter>, and so on.
[0112] In some embodiments, before the enrollment component 505 processes the audio data 211, the emotion detection component 275 may determine that the audio data 211 includes speech from one or more persons other than the user enrolled in emotion detection. For example, as part of the enrollment process, the system may receive permission from the user to record his or her speech for a limited period of time, thereby obtaining audio representing the user's interactions in a variety of situations and contexts, so that the system can determine a baseline for different contexts. As described above, this is beneficial because a user may exhibit different speaking styles / acoustic voice attributes in different situations based on who he or she is interacting with, where he or she is speaking, and / or what he or she is doing. Therefore, the audio data 211 may include speech from one or more persons other than the user. In this case, the emotion detection component 275 may identify the one or more users using the user identification component 295, as in combination with Figure 3 and Figure 4 If a portion of the audio data 211 is determined to be from a person other than the user, that portion of the audio data 211 is discarded, and only the portion of the audio data 211 corresponding to the user is stored for further processing and to register the user for emotion detection.
[0113] Figure 6 is a conceptual diagram showing an emotion detection component according to an embodiment of the present disclosure. Figure 5 In addition to the components shown, the emotion detection component 275 may also include a voice activity detection (VAD) component 605, a training model 615, and a baseline selection component 620. The audio data 211 captured by the device 110 may be input into the VAD component 605. The emotion detection component 275 may reside with the device 110a, with another device (such as the device 110b) that is close to and in communication with the device 110, or with a remote device (such as the one or more systems 120). If the emotion detection component 275 is not resident on the device 110a that captures the audio, the emotion detection component 275 may not necessarily include the VAD component 605 (or may not necessarily include other components) and may or may not include other components. The exact composition of the emotion detection component 275 depends on the system configuration.
[0114] The VAD component 605 can determine whether the audio data 211 includes speech spoken by a human or voice activity performed by a human, and can determine that the audio data 211 includes a portion of speech or voice activity. The VAD component 605 can send the portion of the audio data 211 including speech or voice activity to the user identification component 295. The VAD component 605 can use voice activity detection technology. Such technology can determine whether speech is present in the audio data based on various quantitative aspects of the audio data (such as the spectral slope between one or more frames of audio data; the energy level of the audio data in one or more spectral bands; the signal-to-noise ratio of the audio data in one or more spectral bands; or other quantitative aspects). In other examples, the VAD component 605 can implement a finite classifier configured to distinguish speech from background noise. The classifier can be implemented using technologies such as linear classifiers, support vector machines, and decision trees. In yet other examples, device 110 may apply Hidden Markov Model (HMM) or Gaussian Mixture Model (GMM) techniques to compare the audio data with one or more acoustic models in a storage device, wherein the acoustic models may include models corresponding to speech, noise (e.g., ambient noise or background noise), or silence. Still other techniques may be used to determine whether speech is present in the audio data.
[0115] The user identification component 295 (which may be located on the same or different device as the emotion detection component 275) may communicate with the emotion detection component 275 to determine the user audio data 610 corresponding to a specific user profile. The user identification component 295 may identify one or more users, such as in conjunction with Figure 3 and Figure 4 As described. For example, the user identification component 295 may identify stored data corresponding to a voice profile associated with the user profile and, based on analyzing the stored data, determine a confidence level that a portion of the input audio data corresponds to the voice profile. The user identification component 295 may determine whether the confidence level meets / satisfies a threshold. If the confidence level of a portion of the input audio data is below the threshold, the corresponding portion of the input audio is discarded because it does not represent the speech of the user associated with the user profile. If the confidence level of a portion of the input audio data meets / satisfies the threshold, the corresponding portion of the input audio data is stored as user audio data 610.
[0116] User audio data 610 may be a portion of audio data 211 that includes the speech or one or more utterances of a specific user associated with a user profile. In other words, audio data representing the speech of a specific user may be isolated and stored as user audio data 610 for further analysis. In an exemplary embodiment, a user may be associated with or use device 110 and may have provided permission to one or more systems 120 to record and analyze his or her speech / conversation to determine an emotion category corresponding to the conversation.
[0117] Before performing further analysis on the user audio data 610 , the emotion detection component 275 can confirm that the user has granted permission to analyze the user's spoken speech for emotion detection.
[0118] The user audio data 610 may be input to the encoder 1150 ( Figure 11 1150 ). The one or more frame feature vectors 612 may be used to determine the audio frame level features extracted from the user audio data 610. One frame feature vector 612 may represent features extracted from a 25ms window of audio, where the window slides or moves in increments of 10ms to extract features represented by the next frame feature vector. In other embodiments, one frame feature vector 612 may represent features corresponding to a single word in the speech. The emotion detection component 275 may determine the portions of the user audio data 610 corresponding to the respective words and extract features from the respective portions of the audio using the encoder 1150. One or more frame feature vectors 612 may be derived from a spectral analysis of the user audio data 610 and may indicate acoustic speech attributes such as accent, pitch, intonation, tone, stress, rhythm, speed, and the like.
[0119] Baseline selection component 620 can be configured to identify or select a baseline for emotion detection. In some embodiments, profile storage device 270 can store reference audio data corresponding to multiple baselines associated with different contextual data. Baseline selection component 620 can determine which baseline to use during runtime to analyze specific input audio data 211. Baseline selection component 620 can select a baseline from the multiple baselines based on the contextual data associated with the baseline and the contextual data associated with audio data 211. Baseline selection component 620 requests context component 515 to determine contextual data corresponding to audio data 211, such as the user's location (e.g., the location of the user using device 110), the people they are interacting with, etc. Baseline selection component 620 can select a baseline with contextual data similar to the contextual data of audio data 211 used for emotion detection, thereby using an appropriate baseline to account for the fact that users exhibit different speaking styles / acoustic speech attributes in different situations. In other embodiments, baseline selection component 620 can analyze the reference audio data corresponding to the baseline and features of audio data 211 (e.g., using an ML model, statistical analysis, or other methods) to identify a baseline with features similar to audio data 211. In some embodiments, if baseline selection component 620 is unable to identify a baseline with contextual data similar to that of the audio data, baseline selection component 620 may select the best available baseline based on the quality of the baselines (e.g., audio quality, quality of acoustic features, best representation of a neutral emotional state, etc.) In some embodiments, the system may determine an average baseline using features of all or some of the baselines associated with the user profile.
[0120] In some embodiments where the profile storage 270 includes only one baseline, the baseline selection component 620 may be disabled and may not perform any action.
[0121] The baseline selection component 620 can retrieve reference audio data 602 corresponding to the baseline to be used for emotion detection. The reference audio data 602 can be input to the encoder 1150 ( Figure 111150 ). The one or more frame feature vectors 614 may be used to determine the one or more frame feature vectors 614. The one or more frame feature vectors 614 may represent audio frame-level features extracted from the reference audio data 602. One frame feature vector 614 may represent features extracted for a 25ms window of audio, where the window slides or moves in increments of 10ms to extract features represented by the next frame feature vector. In other embodiments, one frame feature vector 614 may represent features corresponding to a single word in the speech. The emotion detection component 275 may determine that the reference audio data 602 corresponds to a portion of a single word and extract features from the corresponding portion of the audio using the encoder 1150. The one or more frame feature vectors 614 may be derived by spectral analysis of the reference audio data 602 and may indicate acoustic speech attributes corresponding to a neutral emotional state of the user, such as accent, pitch, intonation, tone, stress, rhythm, speed, etc.
[0122] The training model 615 can process one or more frame feature vectors 612 and one or more frame feature vectors 614. The training model 615 can be configured to process the features of the reference audio data 602 and the input audio data 211 to determine the emotion category corresponding to the audio data 211 based on the neutral emotional state of the user (represented by the reference audio data 602). The training model 615 can output one or more scores 630, which indicate the emotion category 640 corresponding to the audio data 211. The emotion category can include broad categories, such as positive, neutral and negative. In other embodiments, the emotion category can be more specific and can include, for example, anger, happiness, sadness and neutral. In another embodiment, the emotion category can include anger, sadness, happiness, surprise, stress and disgust. As can be understood, various emotion categories / indicators are possible according to the system configuration. In some embodiments, the training model 615 can be configured to determine the background data corresponding to the input audio data 211.
[0123] In some embodiments, the system may be configured to further process the audio data 211 / user audio data 610 using one or more other trained models to detect the user's emotions derived from the words the user says to express his or her opinions / views.
[0124] The training model 615 can be a neural network, such as a deep learning neural network (DNN). Figure 8 As shown, a neural network may include multiple layers from input layer 1 810 to output layer N 820. Each layer includes one or more nodes and is configured to input data of a specific type and output data of another type. The layers may be represented by data structures that represent the connections between the layers and the operations within the layers. Figure 8The neural network shown is configured to input data of type data A 802 (which is the input to layer 1 810) and output data of type data Z 808 (which is the output from the last layer N 820). The output from one layer is then used as the input to the next layer. For example, the output data from layer 1 810 (data B 804) is the input data for layer 2 812, and so on, so that the input to layer N 820 is data Y 806 output from the second-to-last layer (not shown).
[0125] The data describing the neural network describes the structure and operation of the neural network layers when the values of the input data / output data of a specific layer are unknown before the neural network actually operates during runtime.
[0126] Machine learning (ML) is a valuable computing technique that allows computing systems to learn techniques for solving complex problems without requiring them to follow explicit algorithms. ML uses trained models, which consist of internally configured operations that manipulate specific types of input data to determine a desired outcome. Trained models are used in many computing tasks, such as computer vision, speech processing, predictive analytics, and more.
[0127] Training models comes in many forms, including training classifiers, support vector machines (SVMs), neural networks (such as deep neural networks (DNNs), recurrent neural networks (RNNs), or convolutional neural networks (CNNs)), etc. For example, a neural network typically includes an input layer, an output layer, and one or more intermediate hidden layers, where the input layer is configured to receive a specific type of data, and the output layer is configured to output a desired type of data to be derived from the network, and the one or more hidden layers perform various functions to generate output data from the input data.
[0128] Various machine learning techniques can be used to train and operate the model to perform the various steps described herein, such as user identification feature extraction, encoding, user identification scoring, user identification confidence determination, etc. The model can be trained and operated according to various machine learning techniques. Such techniques may include, for example, neural networks (such as deep neural networks and / or recurrent neural networks), inference engines, trained classifiers, etc. Examples of trained classifiers include support vector machines (SVMs), neural networks, decision trees, AdaBoost (abbreviation for "Adaptive Boosting") combined with decision trees and random forests. For example, focusing on SVMs, SVMs are a supervised learning model with an associated learning algorithm that analyzes data and identifies patterns in the data, typically used for classification and regression analysis. Given a set of training examples, each labeled as belonging to one of two categories, the SVM training algorithm builds a model that assigns new examples to one category or the other, making it a non-probabilistic binary linear classifier. More complex SVM models can be constructed using training sets that identify more than two categories, where the SVM determines which category is most similar to the input data. The SVM model can be mapped so that examples of different categories are separated by a clear gap. New examples are then mapped into this same space and predicted to belong to a category based on which side of the gap they fall. The classifier can publish a "score" indicating the category to which the data most closely matches. The score provides an indication of how closely the data matches the category.
[0129] In order to apply machine learning techniques, the machine learning process itself needs to be trained. Training a machine learning component (in this case, such as one of the first or second models) requires establishing a "ground truth" for the training examples. In machine learning, the term "ground truth" refers to the accuracy of the classification of the training set for supervised learning techniques. A variety of techniques can be used to train the model, including backpropagation, statistical learning, supervised learning, semi-supervised learning, stochastic learning, or other known techniques.
[0130] Figure 7 Components for training an ML model for emotion detection using a baseline are conceptually illustrated. The emotion component 275 can include a model building component 710. The model building component 710 can be a separate component included in one or more systems 120.
[0131] The model building component 710 can train one or more machine learning models to determine an emotion corresponding to a user input based on the user's neutral emotional state represented by the baseline / reference audio data. The model building component 710 can train the one or more machine learning models during offline operation. The model building component 710 can train the one or more machine learning models using a training dataset.
[0132] The training data set may include a pair of audio data, one audio data representing a neutral emotional state of a speaker and the other audio data representing a non-neutral emotional state of the speaker. For example, reference audio data 702a may represent a neutral emotional state of a first speaker, and test audio data 704a may represent a non-neutral (e.g., angry) emotional state of the first speaker. Reference audio data 702b may represent a neutral emotional state of a second speaker and test audio data 704b may represent a non-neutral (e.g., happy) emotional state of the second speaker. The pair of audio data 702 and 704 may constitute a training data set used by the model building component 710 to train an ML model to detect emotions using a baseline. The test audio data 704 may be annotated or labeled with the emotion category corresponding to the test audio data.
[0133] In some embodiments, the training dataset may also include contextual data 706 corresponding to the reference audio data 702 and / or the test audio data 704. The contextual data 706a may, for example, represent the environment, situation, location, occasion, or other contextual information corresponding to the first speaker when the first speaker uttered the reference audio data 702a and / or the test audio data 704a. The contextual data 706a may also represent the type of interaction, including who the first speaker was interacting with when the utterance was uttered, the occasion the first speaker was in when the utterance was uttered (e.g., a work meeting, a family / friend gathering, a sporting event, a concert, etc.), the time of day (e.g., morning, afternoon, evening, day of the week, etc.), any actions the first speaker was performing when the utterance was uttered (e.g., driving, walking, watching TV, etc.), etc. The contextual data 520 may also include data representing other contextual information corresponding to when the first speaker uttered the audio, such as weather information, physiological data associated with the user, the season of the year, the month of the year, etc. The contextual data 706a may represent the context corresponding to the reference audio data 702a and the test audio data 704a, where both have similar / identical contexts. In other embodiments, the background data 706a may represent only the background corresponding to the reference audio data 702a, and the training data set may optionally include additional background data (not shown) corresponding to the test audio data 704a. Thus, the training model 615 can be configured using the background data 706 to determine / identify background data corresponding to the input audio data during runtime operation.
[0134] As part of the training process, model building component 710 can determine weights and parameters for various layers of training model 615. The weights and parameters corresponding to the final state of training model 615 can be stored as stored data 712.
[0135] exist Figure 9An exemplary neural network for training model 615 is shown in FIG. The neural network may be composed of an input layer 902, one or more intermediate layers 904, and an output layer 906. The one or more intermediate layers may also be referred to as one or more hidden layers. Each node in the hidden layer is connected to each node in the input layer and each node in the output layer. Although Figure 9 Although shown with a single hidden layer, a neural network can include multiple intermediate layers. In this case, each node in a hidden layer is connected to every node in the next higher layer and the next lower layer. Each node in the input layer represents a potential input to the neural network, and each node in the output layer represents a potential output of the neural network. Each connection from one node to another node in the next layer can be associated with a weight, or score. The neural network can output a single output or a weighted set of possible outputs.
[0136] In one aspect, a neural network can be constructed with recurrent connections so that the output of a hidden layer of the network is fed back into the hidden layer for the next set of inputs. Figure 10 Such a neural network is shown in . Each node in the input layer 1002 is connected to each node in the hidden layer 1004. Each node in the hidden layer 1004 is connected to each node in the output layer 1006. As shown in the figure, the output of the hidden layer 1004 is fed back to the hidden layer to be used to process the next set of inputs. A neural network that incorporates recurrent connections may be referred to as a recurrent neural network (RNN).
[0137] Neural networks can also be used to perform ASR processing, including acoustic model processing and language model processing. In the case of an acoustic model using a neural network, each node of the input layer of the neural network can represent the acoustic features of a feature vector of acoustic features, such as those that can be output after a first pass of speech recognition, and each node of the output layer represents a score corresponding to a sub-word unit (such as a phoneme, triphone, etc.) and / or an associated state that can correspond to the sound represented by the feature vector. For a given input to the neural network, it outputs multiple potential outputs, each of which has an assigned score representing the probability that the specific output is the correct output given the specific input. The highest-scoring output of the acoustic model neural network can then be fed back into the HMM, which can determine the transitions between sounds before passing the results to the language model.
[0138] In the case where the language model uses a neural network, each node in the input layer of the neural network can represent the previous word, and each node in the output layer can represent the potential next word determined by the trained neural network language model. Because the language model can be configured as a recurrent neural network that incorporates some history of words processed by the neural network, such as Figure 10The network shown. Predictions of potential next words can be based on previous words in the utterance rather than just the most recent word. The language model neural network can also output a weighted prediction of the next word.
[0139] The processing of a neural network is determined by the learned weights of each node input and the network structure. Given a specific input, the neural network determines the output one layer at a time until the output layer of the entire network is calculated.
[0140] The connection weights can initially be learned by the neural network during training, where a given input is associated with a known output. In a set of training data, various training examples are fed back into the network. Each example typically sets the weight of the correct connection from the input to the output to 1, and assigns a weight of 0 to all connections. As the examples in the training data are processed by the neural network, inputs can be sent to the network and compared with the associated outputs to determine how well the network performs compared to the target performance. Using training techniques (such as backpropagation), the weights of the neural network can be updated to reduce the error generated by the neural network when processing the training data. In some cases, the neural network can be trained using the entire grid to improve speech recognition when processing the entire grid.
[0141] Figure 11 1, 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, n ,...× N , where × n Is a D-dimensional vector, encoder E(×1,...× N )=y projects the feature sequence onto y, where y is an F-dimensional vector. F is a fixed length of the vector and can be configured based on the user and other system configurations of the encoding vector. Any specific encoder 1150 will be configured to output vectors of the same size, thereby ensuring continuity in the size of the output encoding vectors from any specific encoder 1150 (although different encoders may output vectors of different fixed sizes). The value y may be referred to as a sequence ×1, ... × N Embedded. × n The lengths of and y are fixed and known a priori, but the feature sequence ×1, ... × NThe length of N is not necessarily known a priori. The encoder can be implemented as a recurrent neural network (RNN), such as a long short-term memory RNN (LSTM-RNN) or a gated recurrent unit RNN (GRU-RNN). RNN is a tool by which a network of nodes can be represented digitally, and each node representation includes information about the previous part of the network. For example, an RNN performs a linear transformation on a sequence of feature vectors, which converts the sequence into a fixed-size vector. The resulting vector maintains the characteristics of the sequence, which could originally be arbitrarily long, in a reduced vector space. The output of the RNN after consuming the sequence of feature data values is the encoder output. There are many ways for an RNN encoder to consume the encoder output, including but not limited to:
[0142] Linear, one direction (forward or backward),
[0143] ● Bilinear, which is essentially a concatenation of forward and backward embeddings, or
[0144] ●Tree, sequence-based parse tree.
[0145] Additionally, an attention model can be used, which is another RNN or DNN that learns to "draw" attention to specific parts of the input. The attention model can be used in conjunction with the above-mentioned method of consuming the input.
[0146] Figure 11 The operation of the encoder 1150 is shown. The input feature value sequence (starting with feature value ×11102, through feature value × n 1104 continues and takes the eigenvalue × N1106 ends) is input into encoder 1150. Encoder 1150 can process the input feature values as described above. Encoder 1150 outputs encoded feature vector y 1110, which is a fixed-length feature vector of length F. One or more encoders (such as 1150) can be used together with emotion detection component 275. For example, audio data 211 / user audio data 610 can be processed using encoder 1150a to determine one or more feature vectors 612, and reference audio data 602 can be processed using encoder 1150b to determine one or more feature vectors 614. In some embodiments, encoders 1150a and 1150b can both be LSTMs, but can have different weights and parameters configured to encode input audio data and reference audio data respectively. In other embodiments, encoders 1150a and 1150b can have the same weights and parameters. In another embodiment, encoder 1150a (for processing input audio data) and encoder 1150b (for processing reference audio data) can share their weights and parameters for specific layers. For example, the emotion detection component 275 may employ a shared or stacked LSTM to process the input audio data and the reference audio data. One or more layers of the encoder 1150b (e.g., layer 1 810, layer 812) may share their weights and parameters with one or more layers of the encoder 1150a, and vice versa.
[0147] Figure 12 is a block diagram conceptually illustrating device 110a and device 110b that may be used with the system. Figure 13 is a block diagram conceptually illustrating exemplary components of a remote device, such as one or more systems 120 and one or more skill systems 225 that may assist in ASR processing, NLU processing, and the like. The system (120 / 225) may include one or more servers. As used herein, "server" may refer to a traditional server as understood in a server / client computing architecture, but may also refer to many different computing components that may assist in the operations discussed herein. For example, a server may include one or more physical computing components (such as rack servers) that are physically and / or network-connected to other devices / components and are capable of performing computing operations. A server may also include one or more virtual machines that simulate a computer system and run on one device or across multiple devices. A server may also include other combinations of hardware, software, firmware, etc. to perform the operations discussed herein. One or more servers may be configured to operate using one or more of the following: a client-server model, a computer bureau model, grid computing technology, fog computing technology, mainframe technology, utility computing technology, a peer-to-peer model, sandbox technology, or other computing technology.
[0148] Multiple systems (100 / 120 / 225) may be included in the overall system of the present disclosure, such as one or more systems 120 for performing ASR processing, one or more systems 120 for performing NLU processing, one or more skill systems 225 for performing actions in response to user input, etc. In operation, each of these systems may include computer-readable instructions and computer-executable instructions residing on a corresponding device (120 / 225), as will be discussed further below.
[0149] Each of these devices (100 / 110 / 120 / 225) may include one or more controllers / processors (1204 / 1304), each of which may include a central processing unit (CPU) for processing data and computer-readable instructions, and memory (1206 / 1306) for storing data and instructions for the corresponding device. The memory (1206 / 1306) may individually include volatile random access memory (RAM), non-volatile read-only memory (ROM), non-volatile magnetoresistive memory (MRAM), and / or other types of memory. Each device (100 / 110 / 120 / 225) may also include a data storage device component (1208 / 1308) for storing data and controller / processor executable instructions. Each data storage device component (1208 / 1308) may individually include one or more types of non-volatile storage devices, such as magnetic storage devices, optical storage devices, solid-state storage devices, etc. Each device (100 / 110 / 120 / 225) may also be connected to removable or external non-volatile memory and / or storage devices (such as removable memory cards, storage key drives, network storage devices, etc.) through a corresponding input / output device interface (1202 / 1302).
[0150] Computer instructions for operating each device (100 / 110 / 120 / 225) and its various components may be executed by one or more controllers / processors (1204 / 1304) of the respective device, using memory (1206 / 1306) as temporary "working" storage during runtime. The computer instructions for a device may be stored in a non-transitory manner in non-volatile memory (1206 / 1306), storage devices (1208 / 1308), or one or more external devices. Alternatively, in addition to or in lieu of software, some or all of the executable instructions may be in hardware or firmware embedded on the respective device.
[0151] Each device (100 / 110 / 120 / 225) includes an input / output device interface (1202 / 1302). Various components can be connected via the input / output device interface (1202 / 1302), as will be discussed further below. In addition, each device (100 / 110 / 120 / 225) can include an address / data bus (1224 / 1324) for transferring data between components of the respective device. Each component within the device (100 / 110 / 120 / 225) can also be directly connected to other components, in addition to (or instead of) connecting to other components across the bus (1224 / 1324).
[0152] Reference Figure 12 , the device 110 may include an input / output device interface 1202 that is connected to various components, such as an audio output component, such as a speaker 1212, a wired or wireless headset (not shown), or other components capable of outputting audio. The device 110 may also include an audio capture component. The audio capture component may be, for example, a microphone 1220 or an array of microphones 1220, a wired or wireless headset (not shown), or the like. If an array of microphones 1220 is included, the approximate distance from the origin of the sound can be determined by acoustic localization based on the time and amplitude differences between the sounds captured by different microphones of the array. The device 110 may additionally include a display 1216 for displaying content. The device 110 may also include a camera 1218.
[0153] Via one or more antennas 1214, the I / O device interface 1202 can connect to one or more networks 199 via a wireless local area network (WLAN) such as a WiFi radio, Bluetooth, and / or a wireless network radio such as a radio capable of communicating with a wireless communication network such as a Long Term Evolution (LTE) network, a WiMAX network, a 3G network, a 4G network, a 5G network, etc. Wired connections such as Ethernet can also be supported. Through one or more networks 199, the system can be distributed across a networked environment. The I / O device interface (1202 / 1302) can also include communication components that allow data to be exchanged between devices such as different physical servers or other components in a server collection.
[0154] Components of one or more devices 110, one or more systems 100, one or more systems 120, or one or more skill systems 225 may include their own dedicated processors, memory, and / or storage. Alternatively, one or more of the components of one or more devices 110, one or more systems 120, or one or more skill systems 225 may utilize the I / O device interfaces (1202 / 1302), one or more processors (1204 / 1304), memory (1206 / 1306), and / or storage (1208 / 1308) of one or more devices 110, one or more systems 120, or one or more skill systems 225, respectively. Thus, the ASR component 250 may have its own I / O device interfaces (1202 / 1302), one or more processors, memory, and / or storage; the NLU component 260 may have its own I / O interfaces (1202 / 1302), one or more processors, memory, and / or storage; and so on for the various components discussed herein.
[0155] As described above, multiple devices may be employed in a single system. In such a multi-device system, each of the devices may include different components for performing different aspects of the system's processing. Multiple devices may include overlapping components. As described herein, the components of device 110, one or more systems 100, one or more systems 120, and one or more skill systems 225 are illustrative and may be positioned as standalone devices or may be included in whole or in part as components of a larger device or system.
[0156] like Figure 14As shown, multiple devices (110a-110k, 120, 225) may comprise components of the system, and the devices may be connected via one or more networks 199. The one or more networks 199 may include local or private networks or may include wide-area networks such as the Internet. Devices may connect to the one or more networks 199 via wired or wireless connections. For example, smartwatch 110a, smartphone 110b, voice detection device 110c, tablet computer 110d, vehicle 110e, display device 110f, smart TV 110g, washer / dryer 110h, refrigerator 110i, toaster 110j, and / or microwave oven 110k may connect to the one or more networks 199 via a wireless service provider, via WiFi or cellular network connection, etc. Other devices may be included as networked support devices, such as one or more systems 120, one or more skill systems 225, and / or other items. Support devices may connect to the one or more networks 199 via wired or wireless connections. A networked device may capture audio using one or more built-in or connected microphones or other audio capture devices, where processing is performed by an ASR component, NLU component, or other component (such as an ASR component 250, NLU component 260, etc. of one or more systems 120) of the same device or another device connected via one or more networks 199.
[0157] The foregoing may also be understood in light of the following clauses.
[0158] 1. A computer-implemented method comprising:
[0159] During the registration period:
[0160] receiving first audio data representing a first reference utterance spoken by a user;
[0161] processing the first audio data to determine that the first audio data represents a neutral emotional state of the user;
[0162] determining first contextual data corresponding to the first audio data, the first contextual data representing at least one of a first location or a first interaction type associated with the first audio data;
[0163] determining a first feature vector corresponding to the first audio data, the first feature vector representing an acoustic speech property corresponding to the first audio data; and
[0164] associating the first feature vector with the first contextual data and a user profile associated with the user;
[0165] receiving second audio data representing a second reference utterance spoken by the user;
[0166] processing the second audio data to determine that the second audio data represents a neutral emotional state of the user;
[0167] determining second contextual data corresponding to the second audio data, the second contextual data representing at least one of a second location or a second interaction type associated with the second audio data;
[0168] determining a second feature vector corresponding to the second audio data, the second feature vector representing an acoustic speech property corresponding to the second audio data; and
[0169] associating the second feature vector with the second contextual data and the user profile;
[0170] During a certain period of time after the registration period:
[0171] receiving third audio data representing an input utterance spoken by the user;
[0172] determining a third eigenvector corresponding to the third audio data, the third eigenvector representing an acoustic speech attribute corresponding to the third audio data;
[0173] determining third background data corresponding to the third audio data;
[0174] selecting the first feature vector based on that the third background data corresponds to the first background data;
[0175] processing the first feature vector and the third feature vector using a training model to determine a score, the training model configured to compare reference audio data to input audio data to determine an emotion associated with the third audio data;
[0176] determining a sentiment category using the scores; and
[0177] The emotion category is associated with the third audio data and the user profile.
[0178] 2. The computer-implemented method of clause 1, further comprising:
[0179] During the Registration Period:
[0180] receiving second audio data representing a second reference utterance spoken by the user;
[0181] processing the second audio data using an emotion detection model to determine first emotion data representing an emotion of the user when the second reference utterance was spoken;
[0182] determining that the first emotion data indicates an emotion other than neutral; and
[0183] generating output audio data requesting the user to speak another utterance;
[0184] receiving the first audio data in response to the output audio data, and
[0185] Wherein processing the first audio data to determine that the first audio data represents the neutral emotional state of the user comprises:
[0186] processing the first audio data using the emotion detection model to determine second emotion data representing an emotion of the user when the first reference utterance was spoken; and
[0187] It is determined that the second emotion data indicates a neutral emotion class.
[0188] 3. The computer-implemented method of clause 1 or 2, wherein determining the first eigenvector comprises:
[0189] processing the first audio data using a first encoder having at least a first processing layer corresponding to first model data and a second processing layer corresponding to second model data to determine the first feature vector, wherein the first model data and the second model data are associated with a neutral emotional state of the user, and
[0190] Wherein determining the third eigenvector comprises:
[0191] The third audio data is processed using a second encoder having a third processing layer corresponding to at least third model data to determine the third feature vector, wherein the third model data includes a portion of the first model data.
[0192] 4. The computer-implemented method of clause 1, 2, or 3, further comprising:
[0193] determining text data corresponding to the second audio data using text-to-speech processing;
[0194] determining a timestamp corresponding to the second audio data, the timestamp indicating when a device associated with the user received the first audio data;
[0195] generating output data, the output data comprising the emotion category, the text data, and the timestamp; and
[0196] The output data is displayed via the device.
[0197] 5. A computer-implemented method comprising:
[0198] Receive input audio data;
[0199] determining that the input audio data represents speech spoken by a user associated with a user profile;
[0200] receiving first background data corresponding to the input audio data;
[0201] selecting reference audio data from a plurality of reference audio data associated with the user profile, wherein the reference audio data is selected based on the first context data corresponding to second context data associated with the reference audio data, and the reference audio data represents a neutral emotional state of the user;
[0202] determining first feature data representing acoustic speech properties corresponding to the reference audio data;
[0203] determining second feature data representing acoustic speech properties corresponding to the input audio data;
[0204] Processing the first feature data and the second feature data using a training model to determine an emotion category corresponding to the input audio data; and
[0205] Association data is stored that associates the emotion category with the user profile and the input audio data.
[0206] 6. The computer-implemented method of clause 5, further comprising:
[0207] receiving first audio data representing a first reference utterance;
[0208] storing a first position corresponding to the first audio data as the second background data;
[0209] associating the first audio data with the user profile and the second contextual data;
[0210] receiving second audio data representing a second reference utterance;
[0211] storing a second position corresponding to the second audio data as third background data; and
[0212] associating the second audio data with the user profile and the third contextual data,
[0213] Wherein selecting the reference audio data further comprises:
[0214] determining that the first context data includes a third position associated with the input audio data;
[0215] determining that the third position corresponds to the first position; and
[0216] The first audio data is selected as the reference audio data based on the third position corresponding to the first position.
[0217] 7. The computer-implemented method of clause 5 or 6, further comprising:
[0218] receiving first audio data representing a first reference utterance;
[0219] receiving second audio data representing a second reference utterance;
[0220] processing the first audio data using an emotion detection model to determine a first score;
[0221] processing the second audio data using the emotion detection model to determine a second score;
[0222] determining that the first score corresponds to a neutral sentiment category; and
[0223] The first audio data is stored as the reference audio data.
[0224] 8. The computer-implemented method of clause 5, 6, or 7, further comprising:
[0225] receiving first audio data representing a first reference utterance;
[0226] determining the second context data corresponding to the first audio data, the second context data representing at least one of a first location and a first interaction type corresponding to the first audio data;
[0227] associating the first audio data with the user profile and the second contextual data;
[0228] receiving second audio data representing a second reference utterance;
[0229] determining third context data corresponding to the second audio data, the third context data representing at least one of a second location and a second interaction type corresponding to the second audio data;
[0230] associating the second audio data with the user profile and the third contextual data;
[0231] Wherein selecting the reference audio data further comprises:
[0232] determining that the first context data corresponds to the second context data; and
[0233] The first audio data is selected as the reference audio data.
[0234] 9. The computer-implemented method of clause 5, 6, 7, or 8, wherein determining the first feature data and determining the second feature data comprises:
[0235] processing the reference audio data using a first encoder having at least a first processing layer and a second processing layer to determine the first feature data; and
[0236] The input audio data is processed using a second encoder and data corresponding to the second processing layer to determine the second feature data.
[0237] 10. The computer-implemented method of clause 5, 6, 7, 8, or 9, further comprising, at a first time period prior to receiving the input audio data:
[0238] determining a first set of utterances, the first set of utterances comprising a first utterance representing a neutral emotional state of a second user and a second utterance representing a non-neutral emotional state of the second user;
[0239] determining a second set of utterances, the second set of utterances comprising a third utterance representing a neutral emotional state of a third user and a fourth utterance representing a non-neutral emotional state of the third user;
[0240] storing the first set of utterances and the second set of utterances as training data;
[0241] processing the training data to determine model data; and
[0242] The training model is determined using the model data, the training model being configured to compare reference audio with input audio to determine an emotion of the user corresponding to the reference audio and the input audio.
[0243] 11. The computer-implemented method of clause 5, 6, 7, 8, 9, or 10, wherein receiving the input audio data comprises receiving a first utterance spoken by the user and receiving a second utterance spoken by a further user, and the method further comprises:
[0244] determining a first confidence level that the first utterance corresponds to the user profile;
[0245] determining that the first confidence level satisfies a threshold;
[0246] storing a first portion of the input audio data corresponding to the first utterance as user audio data;
[0247] determining a second confidence level that the second utterance corresponds to the user profile;
[0248] determining that the second confidence level fails to meet the threshold;
[0249] discarding a second portion of the input audio data corresponding to the second utterance; and
[0250] The second feature is determined using the first portion of the input audio data.
[0251] 12. The computer-implemented method of clause 5, 6, 7, 8, 9, 10, or 11, further comprising:
[0252] determining text data corresponding to the input audio data using text-to-speech processing;
[0253] determining time data indicating when the device received the input audio data;
[0254] generating output data comprising the textual data, the temporal data, and an indicator of the emotion category; and
[0255] The output data is displayed using the device.
[0256] 13. A system comprising:
[0257] at least one processor; and
[0258] at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
[0259] Receive input audio data;
[0260] determining that the input audio data represents speech spoken by a user associated with a user profile;
[0261] receiving first background data corresponding to the input audio data;
[0262] selecting reference audio data from a plurality of reference audio data associated with the user profile, wherein the reference audio data is selected based on the first context data corresponding to second context data associated with the reference audio data, and the reference audio data represents a neutral emotional state of the user;
[0263] determining first feature data representing acoustic speech properties corresponding to the reference audio data;
[0264] determining second feature data representing acoustic speech properties corresponding to the input audio data;
[0265] Processing the first feature data and the second feature data using a training model to determine an emotion category corresponding to the input audio data; and
[0266] Association data is stored that associates the emotion category with the user profile and the input audio data.
[0267] 14. The system of clause 13, wherein the instructions, when executed by the at least one processor, further cause the system to:
[0268] receiving first audio data representing a first reference utterance;
[0269] storing a first position corresponding to the first audio data as the second background data;
[0270] associating the first audio data with the user profile and the second contextual data;
[0271] receiving second audio data representing a second reference utterance;
[0272] storing a second position corresponding to the second audio data as third background data; and
[0273] associating the second audio data with the user profile and the third contextual data,
[0274] wherein the instructions that cause the system to select the reference audio data further cause the system to:
[0275] determining that the first context data includes a third position associated with the input audio data;
[0276] determining that the third position corresponds to the first position; and
[0277] The first audio data is selected as the reference audio data based on the third position corresponding to the first position.
[0278] 15. The system of clause 13 or 14, wherein the instructions, when executed by the at least one processor, further cause the system to:
[0279] receiving first audio data representing a first reference utterance;
[0280] receiving second audio data representing a second reference utterance;
[0281] processing the first audio data using an emotion detection model to determine a first score;
[0282] processing the second audio data using the emotion detection model to determine a second score; determining that the first score corresponds to a neutral emotion category; and
[0283] The first audio data is stored as the reference audio data.
[0284] 16. The system of clause 13, 14, or 15, wherein the instructions, when executed by the at least one processor, further cause the system to:
[0285] receiving first audio data representing a first reference utterance;
[0286] determining the second context data corresponding to the first audio data, the second context data representing at least one of a first location and a first interaction type corresponding to the first audio data;
[0287] associating the first audio data with the user profile and the second contextual data;
[0288] receiving second audio data representing a second reference utterance;
[0289] determining third context data corresponding to the second audio data, the third context data representing at least one of a second location and a second interaction type corresponding to the second audio data;
[0290] associating the second audio data with the user profile and the third contextual data;
[0291] wherein the instructions that cause the system to select the reference audio data further cause the system to:
[0292] determining that the first context data corresponds to the second context data; and
[0293] The first audio data is selected as the reference audio data.
[0294] 17. The system of clause 13, 14, 15 or 16, wherein the instructions causing the system to determine the first feature data and to determine the second feature data further cause the system to:
[0295] processing the reference audio data using a first encoder having at least a first processing layer and a second processing layer to determine the first feature data; and
[0296] The input audio data is processed using a second encoder and data corresponding to the second processing layer to determine the second feature data.
[0297] 18. The system of clause 13, 14, 15, 16, or 17, wherein the instructions, when executed by the at least one processor, further cause the system, during a first time period prior to receiving the input audio data:
[0298] determining a first set of utterances, the first set of utterances comprising a first utterance representing a neutral emotional state of a second user and a second utterance representing a non-neutral emotional state of the second user;
[0299] determining a second set of utterances, the second set of utterances comprising a third utterance representing a neutral emotional state of a third user and a fourth utterance representing a non-neutral emotional state of the third user;
[0300] storing the first set of utterances and the second set of utterances as training data;
[0301] processing the training data to determine model data; and
[0302] The training model is determined using the model data, the training model being configured to compare reference audio with input audio to determine an emotion of the user corresponding to the reference audio and the input audio.
[0303] 19. The system of clause 13, 14, 15, 16, 17, or 18, wherein the instructions that cause the system to receive the input audio data further cause the system to receive a first utterance spoken by the user and receive a second utterance spoken by a further user, and the instructions further cause the system to:
[0304] determining a first confidence level that the first utterance corresponds to the user profile;
[0305] determining that the first confidence level satisfies a threshold;
[0306] storing a first portion of the input audio data corresponding to the first utterance as user audio data;
[0307] determining a second confidence level that the second utterance corresponds to the user profile;
[0308] determining that the second confidence level fails to meet the threshold;
[0309] discarding a second portion of the input audio data corresponding to the second utterance; and
[0310] The second feature is determined using the first portion of the input audio data.
[0311] 20. The system of clause 13, 14, 15, 16, 17, 18, or 19, wherein the instructions, when executed by the at least one processor, further cause the system to:
[0312] determining text data corresponding to the input audio data using text-to-speech processing;
[0313] determining time data indicating when the device received the input audio data;
[0314] generating output data comprising the textual data, the temporal data, and an indicator of the emotion category; and
[0315] The output data is displayed using the device.
[0316] The concepts disclosed herein may be employed within a variety of devices and computer systems, including, for example, general-purpose computing systems, speech processing systems, and distributed computing environments.
[0317] The above aspects of the present disclosure are intended to be illustrative. They are selected to explain the principles and applications of the present disclosure and are not intended to be exhaustive or limit the present disclosure. Many modifications and variations of the disclosed aspects may be apparent to those skilled in the art. Those of ordinary skill in the art of computer and speech processing will recognize that the components and process steps described herein may be interchangeable with other components or steps or combinations of components or steps and still achieve the benefits and advantages of the present disclosure. In addition, those skilled in the art will appreciate that the present disclosure may be practiced without some or all of the specific details and steps disclosed herein.
[0318] Aspects of the disclosed system may be implemented as a computer method or an article of manufacture such as a memory device or a non-transitory computer-readable storage medium. The computer-readable storage medium may be computer-readable and may include instructions for causing a computer or other device to perform the processes described in the present disclosure. The computer-readable storage medium may be implemented by volatile computer memory, non-volatile computer memory, a hard drive, a solid-state memory, a flash drive, a removable disk, and / or other media. In addition, components of the system may be implemented in firmware or hardware, such as an acoustic front end (AFE), which includes, among other things, analog and / or digital filters (e.g., filters configured as firmware of a digital signal processor (DSP)).
[0319] Unless otherwise specifically stated or otherwise understood in the context of use, conditional language used herein, such as "can," "may," "could," "may," "for example," etc., among others, is generally intended to convey that a particular embodiment includes particular features, elements, and / or steps, although other embodiments do not. Therefore, such conditional language is generally not intended to imply that the features, elements, and / or steps are necessary for one or more embodiments in any way, or to imply that one or more embodiments must include logic for deciding whether to include these features, elements, and / or steps or whether to perform these features, elements, and / or steps in any specific embodiment with or without other input or prompts. The terms "comprise," "include," "have," etc. are synonymous and are used inclusively in an open-ended manner and do not exclude additional elements, features, actions, operations, etc. In addition, the term "or" is used in its inclusive sense (rather than exclusive sense) so that, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list.
[0320] Unless specifically stated otherwise, disjunctive language, such as the phrase "at least one of X, Y, Z," is understood to be generally used to present the context that an item, term, etc. can be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that a particular embodiment requires that each of at least one of X, at least one of Y, or at least one of Z be present.
[0321] As used in this disclosure, unless specifically stated otherwise, the terms "a" or "an" may include one or more items. In addition, unless specifically stated otherwise, the phrase "based on" is intended to mean "based at least in part on."
Claims
1. A computer-implemented method for detecting emotion in audio data, comprising: Receive input audio data; determining that the input audio data represents speech spoken by a user associated with a user profile; receiving first background data corresponding to the input audio data; selecting reference audio data from a plurality of reference audio data associated with the user profile, wherein the reference audio data is selected based on the first context data corresponding to second context data associated with the reference audio data, and the reference audio data represents a neutral emotional state of the user; determining first feature data representing acoustic speech properties corresponding to the reference audio data; determining second feature data representing acoustic speech properties corresponding to the input audio data; processing the first feature data and the second feature data using a trained model to determine an emotion category corresponding to the input audio data; as well as Association data is stored that associates the emotion category with the user profile and the input audio data.
2. The computer-implemented method for detecting emotion in audio data of claim 1 , further comprising: receiving first audio data representing a first reference utterance; storing a first position corresponding to the first audio data as the second background data; associating the first audio data with the user profile and the second contextual data; receiving second audio data representing a second reference utterance; storing a second position corresponding to the second audio data as third background data; as well as associating the second audio data with the user profile and the third contextual data, Wherein selecting the reference audio data further comprises: determining that the first context data includes a third position associated with the input audio data; determining that the third position corresponds to the first position; and The first audio data is selected as the reference audio data based on the third position corresponding to the first position.
3. The computer-implemented method for detecting emotion in audio data of claim 1 , further comprising: receiving first audio data representing a first reference utterance; receiving second audio data representing a second reference utterance; processing the first audio data using an emotion detection model to determine a first score; processing the second audio data using the emotion detection model to determine a second score; determining that the first score corresponds to a neutral emotion category; as well as The first audio data is stored as the reference audio data.
4. The computer-implemented method for detecting emotion in audio data of claim 1 , further comprising: receiving first audio data representing a first reference utterance; determining the second context data corresponding to the first audio data, the second context data representing at least one of a first location and a first interaction type corresponding to the first audio data; associating the first audio data with the user profile and the second contextual data; receiving second audio data representing a second reference utterance; determining third context data corresponding to the second audio data, the third context data representing at least one of a second location and a second interaction type corresponding to the second audio data; associating the second audio data with the user profile and the third contextual data; Wherein selecting the reference audio data further comprises: determining that the first background data corresponds to the second background data; as well as The first audio data is selected as the reference audio data.
5. The computer-implemented method for detecting emotion in audio data according to any one of claims 1 to 4, wherein determining the first feature data and determining the second feature data comprises: processing the reference audio data using a first encoder having at least a first processing layer and a second processing layer to determine the first feature data; as well as The input audio data is processed using a second encoder and data corresponding to the second processing layer to determine the second feature data.
6. The computer-implemented method for detecting emotion in audio data according to any one of claims 1 to 4, further comprising, at a first time period prior to receiving the input audio data: determining a first set of utterances, the first set of utterances comprising a first utterance representing a neutral emotional state of a second user and a second utterance representing a non-neutral emotional state of the second user; determining a second set of utterances, the second set of utterances comprising a third utterance representing a neutral emotional state of a third user and a fourth utterance representing a non-neutral emotional state of the third user; storing the first set of utterances and the second set of utterances as training data; processing the training data to determine model data; as well as The training model is determined using the model data, the training model being configured to compare reference audio with input audio to determine an emotion of the user corresponding to the reference audio and the input audio.
7. The computer-implemented method for detecting emotion in audio data of any one of claims 1 to 4, wherein receiving the input audio data comprises receiving a first utterance spoken by the user and receiving a second utterance spoken by another user, and the method further comprises: determining a first confidence level that the first utterance corresponds to the user profile; determining that the first confidence level satisfies a threshold; storing a first portion of the input audio data corresponding to the first utterance as user audio data; determining a second confidence level that the second utterance corresponds to the user profile; determining that the second confidence level fails to meet the threshold; discarding a second portion of the input audio data corresponding to the second utterance; as well as The second feature is determined using the first portion of the input audio data.
8. The computer-implemented method for detecting emotions in audio data according to any one of claims 1 to 4, further comprising: determining text data corresponding to the input audio data using text-to-speech processing; determining time data indicating when the device received the input audio data; generating output data, the output data comprising the textual data, the temporal data, and an indicator of the emotion category; as well as The output data is displayed using the device.
9. A system for detecting emotion in audio data, comprising: at least one processor; as well as at least one memory comprising instructions that, when executed by the at least one processor, cause the system to: Receive input audio data; determining that the input audio data represents speech spoken by a user associated with a user profile; receiving first background data corresponding to the input audio data; selecting reference audio data from a plurality of reference audio data associated with the user profile, wherein the reference audio data is selected based on the first context data corresponding to second context data associated with the reference audio data, and the reference audio data represents a neutral emotional state of the user; determining first feature data representing acoustic speech properties corresponding to the reference audio data; determining second feature data representing acoustic speech properties corresponding to the input audio data; processing the first feature data and the second feature data using a trained model to determine an emotion category corresponding to the input audio data; and Association data is stored that associates the emotion category with the user profile and the input audio data.
10. The system for detecting emotion in audio data of claim 9, wherein the instructions, when executed by the at least one processor, further cause the system to: receiving first audio data representing a first reference utterance; storing a first position corresponding to the first audio data as the second background data; associating the first audio data with the user profile and the second contextual data; receiving second audio data representing a second reference utterance; storing a second position corresponding to the second audio data as third background data; and associating the second audio data with the user profile and the third contextual data, wherein the instructions that cause the system to select the reference audio data further cause the system to: determining that the first context data includes a third position associated with the input audio data; determining that the third position corresponds to the first position; and The first audio data is selected as the reference audio data based on the third position corresponding to the first position.
11. The system for detecting emotion in audio data of claim 9, wherein the instructions, when executed by the at least one processor, further cause the system to: receiving first audio data representing a first reference utterance; receiving second audio data representing a second reference utterance; processing the first audio data using an emotion detection model to determine a first score; processing the second audio data using the emotion detection model to determine a second score; determining that the first score corresponds to a neutral emotion category; and The first audio data is stored as the reference audio data.
12. The system for detecting emotion in audio data of claim 9, wherein the instructions, when executed by the at least one processor, further cause the system to: receiving first audio data representing a first reference utterance; determining the second context data corresponding to the first audio data, the second context data representing at least one of a first location and a first interaction type corresponding to the first audio data; associating the first audio data with the user profile and the second contextual data; receiving second audio data representing a second reference utterance; determining third context data corresponding to the second audio data, the third context data representing at least one of a second location and a second interaction type corresponding to the second audio data; associating the second audio data with the user profile and the third contextual data; wherein the instructions that cause the system to select the reference audio data further cause the system to: determining that the first context data corresponds to the second context data; and The first audio data is selected as the reference audio data.
13. A system for detecting emotion in audio data according to any one of claims 9 to 12, wherein the instructions causing the system to determine the first feature data and to determine the second feature data further cause the system to: processing the reference audio data using a first encoder having at least a first processing layer and a second processing layer to determine the first feature data; and The input audio data is processed using a second encoder and data corresponding to the second processing layer to determine the second feature data.
14. The system for detecting emotion in audio data according to any one of claims 9 to 12, wherein the instructions, when executed by the at least one processor, further cause the system to: determining a first set of utterances, the first set of utterances comprising a first utterance representing a neutral emotional state of a second user and a second utterance representing a non-neutral emotional state of the second user; determining a second set of utterances, the second set of utterances comprising a third utterance representing a neutral emotional state of a third user and a fourth utterance representing a non-neutral emotional state of the third user; storing the first set of utterances and the second set of utterances as training data; processing the training data to determine model data; and The training model is determined using the model data, the training model being configured to compare reference audio with input audio to determine an emotion of the user corresponding to the reference audio and the input audio.
15. The system for detecting emotion in audio data according to any one of claims 9 to 12, wherein the instructions, when executed by the at least one processor, further cause the system to: determining text data corresponding to the input audio data using text-to-speech processing; determining time data indicating when the device received the input audio data; generating output data, the output data comprising the textual data, the temporal data, and an indicator of the emotion category; and The output data is displayed using the device.
Citation Information
Patent Citations
Data processing method and device and device for data processing
CN108734096A
Apparatus and Methods for the Detection of Emotions in Audio Interactions
US20080040110A1