EXPRESSION OF EMOTIONS IN THE SPEECH OUTPUT OF CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
Machine learning models integrating user and character data improve emotional state determination in character animation, addressing the limitations of text-based systems by enhancing emotional expression in conversational AI.
Patent Information
- Application Number
- DE102025104331
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-20
- Filing Date
- 2025-02-06
- Publication Date
- 2025-08-07
AI Technical Summary
Conventional systems for animating characters in applications struggle to accurately determine emotional states from voice output due to reliance on text analysis alone, leading to inappropriate or unnatural user experiences.
Utilizing machine learning models that incorporate user information, character information, visual information, and dialog history to determine emotion attributes and generate speech output, allowing for a more nuanced expression of emotions.
Enhances the ability to accurately determine and express a broad range of emotional states in character interactions, providing a more realistic and natural user experience.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Many applications, such as gaming applications, interactive applications, communication applications, multimedia applications, video conferencing applications, in-vehicle infotainment applications and / or similar, use animated characters or digital avatars that interact with users of the applications / machines / devices and / or interact with other animated characters within the applications (e.g., non-player characters (NPCs)). To provide users with a more realistic experience, the systems may attempt to animate the characters by having them express emotions when interacting with the users. For example, when determining speech output, hereinafter referred to as “speech,” that an animated character is to make to a user, a system may also determine an emotional state of the animated character, e.g., based on an analysis of the text of the speech output.The emotional state can then be used to cause the animated character to perform speech in a manner that expresses the emotional state. For example, the animated character's voice used to perform speech can reflect the animated character's emotional state.
[0002] However, if the systems only use the text of the speech to determine emotional states, they may incorrectly determine the emotional states due to the circumstances of the interaction. For example, people can express the same text such as "I wish you a good day" with different emotional states, e.g., happy or indifferent. If the text is only associated with an emotional state, which is then later used by the animated characters when outputting the speech corresponding to the text, the animated characters may express their speech in an inappropriate or incorrect emotional state, which can lead to an undesirable or unnatural user experience. Since the systems only use certain emotional states for animated characters, such asHappy or sad, they are unable to get the animated characters to express a wide range or spectrum of emotional states with speech. For example, people may express the same emotional state differently at different times, e.g., when a person is somewhat cheerful or very cheerful. When the same emotional state is expressed differently, the user's speech may also change, e.g., the characteristics (e.g., pitch, speed, etc.) of the user's speech. SUMMARY
[0003] Embodiments of the present disclosure address the expression of emotions in speech output from conversational AI systems and applications. Systems and methods are described that use one or more machine learning models (e.g., one or more language models—such as a Large Language Model (LLM) and / or a Vision Language Model (VLM)) to determine one or more attributes associated with the speech, such as one or more emotion attributes and / or one or more reaction attributes, and then use the attribute(s) to generate speech output that expresses emotions. In some examples, the machine learning models may use various types of information to determine the attribute(s), such as user information, character information, dialogue history, input text (e.g., a prompt), visual information, audio information, and / or so on.For example, the machine learning models may use the information to determine one or more tags associated with emotional states and / or voice characteristics, with the tag(s) then being used to generate speech output with a voice associated with the emotion.
[0004] Unlike conventional systems, the systems of the present disclosure are capable of determining the emotions associated with the speech output by using additional inputs associated with the speech output text, such as user information, character information, visual information, audio information, and / or dialogue history. As described in detail herein, by using the additional inputs, the current systems can better determine the actual emotion of the speech output—e.g., because the same text may be associated with different emotions due to different circumstances associated with the speech output. Furthermore, unlike the conventional systems, in some embodiments, the current systems are capable of determining additional values for variables, attributes, and / or tags associated with the emotion and / or the speech output.As described in more detail here, by determining the additional values for the variables, attributes and / or tags, the current systems are in turn able to better determine the actual emotion of the speech output - e.g. because the variables can change the emotions and / or voice characteristics associated with the speech output.
[0005] The invention is defined by the claims. To illustrate the invention, aspects and embodiments are described herein, which may or may not fall within the scope of the claims.
[0006] The disclosure extends to all novel aspects or features described and / or illustrated herein.
[0007] Further features of the disclosure are characterized by the independent and dependent claims.
[0008] Any feature of one aspect of the disclosure may be applied to other aspects of the disclosure, in any suitable combination. In particular, method aspects may be applied to device or system aspects, and vice versa.
[0009] Furthermore, functions implemented in hardware may also be implemented in software, and vice versa. Any reference to software and hardware functions herein should be interpreted accordingly.
[0010] Any system or device feature described here may also be provided as a method feature, and vice versa. System and / or device aspects described functionally (including means plus functional features) may alternatively be expressed by their corresponding structure, e.g., by an appropriately programmed processor and associated memory.
[0011] It should also be appreciated that certain combinations of the various features described and defined in each aspect of the disclosure may be implemented and / or provided and / or used independently of one another.
[0012] The disclosure also provides computer programs and computer program products comprising software code that, when executed on a data processing device, is adapted to perform any of the methods described herein and / or embody any of the device and system features described herein, including any or all of the component steps of each method.
[0013] The disclosure also provides a computer or computer system (including networked or distributed systems) having an operating system that supports a computer program for performing the methods described herein and / or for embodying the device or system features described herein.
[0014] The disclosure also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0015] The disclosure also provides a signal carrying one or more of the aforementioned computer programs.
[0016] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0017] Aspects and embodiments of the disclosure will now be described, by way of example only, with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The present systems and methods for expressing emotions in the speech output of conversational AI systems and applications are described in detail below with reference to the attached figures, where: Fig. 1 shows an example data flow diagram of a process for using one or more machine learning models to generate speech output expressing emotions, in accordance with some embodiments of the present disclosure; Fig. 2 shows an example of one or more language models that determine attributes associated with input text in accordance with some embodiments of the present disclosure; Fig. 3 shows an example of one or more language models determining emotions and output text in association with input text, in accordance with some embodiments of the present disclosure; Fig. 4 shows an example of generating speech output expressing emotions in accordance with some embodiments of the present disclosure; Fig. 5A shows a data flow diagram illustrating a process for training one or more language models to generate speech output expressing emotions during a first training stage, in accordance with some embodiments of the present disclosure; Fig. 5B shows a data flow diagram illustrating a process for training one or more language models to generate speech output expressing emotions during a second training stage, in accordance with some embodiments of the present disclosure; Fig. 6A illustrates an example of generating response attributes for training one or more language models in accordance with some embodiments of the present disclosure; Fig. 6B illustrates an example of generating emotion attributes for training one or more language models in accordance with some embodiments of the present disclosure; Fig. 6C illustrates an example of generating ground truth data for training one or more language models in accordance with some embodiments of the present disclosure; Fig. 7 is a flowchart illustrating a method for using models to generate speech output expressing emotions, in accordance with some embodiments of the present disclosure; Fig. 8 is a flowchart illustrating a method for using one or more emotion control models to generate speech output in accordance with some embodiments of the present disclosure; Fig. 9 is a flowchart illustrating a method for using one or more models that manage emotion dialogue to generate speech output, in accordance with some embodiments of the present disclosure; Fig. 10 is a block diagram of an example computing device suitable for use in implementing some embodiments of the present disclosure; and Fig. 11 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] Systems and methods are described that relate to the expression of emotions in speech output from conversational AI systems and applications. For example, a system may receive input data associated with a user or a character being animated. For the user, the input data may include, for example, first text data representing the text entered by the user (e.g., a prompt) (or converted from audio data), second text data representing emotional information (e.g., one or more emotional states) associated with the user, audio data representing the user's speech input (e.g., in the form of a spectrogram), image or video data representing one or more images or videos showing the user, profile data representing information about the user, and / or any other type of data.As described herein, text may represent one or more letters, words, symbols, numbers, characters, punctuation marks, tokens, and / or the like. Additionally, input data for the character may represent information about the character, such as characteristics of the character (e.g., occupation, relationships, personality traits, etc.), past communications, current circumstances (e.g., current interactions with other characters, current location, current goals, etc.), and / or other information associated with the character. In these examples, the input data is described as associated with the user and / or the character; in other examples, the input data may include any other type of input data.
[0020] In some examples, the system(s) may then be configured to process at least a portion of the input data (e.g., the input data associated with the user, the information about the character, etc.) and / or data (referred to in some examples as "history data") representing a dialogue flow between the user and the character using one or more first machine learning models (referred to in some examples as "first model(s)" and / or "state model(s)"), such as one or more language models. As described herein, the first model(s) may be configured to manage an emotional dialogue flow associated with a conversation by predicting the emotions associated with the speech output.For example, the first model(s) may generate and / or output a first output representing one or more attributes associated with emotions (also referred to as “emotion attribute(s)”) and / or one or more attributes associated with a response (also referred to as “generic response attribute(s)”), at least based on processing the at least part of the input data and / or the historical data.
[0021] In some examples, the emotion attributes may be associated with one or more labels associated with one or more emotional states, one or more values (e.g., intensity values) associated with the label(s), and / or the like. As described herein, an emotional state may include, but is not limited to, monotonous, indifferent, anger, disgust, fear, joyful, sad, anticipation, trust, surprise, and / or any other emotional state. Furthermore, the emotion attribute(s) may be associated with the user (e.g., also referred to as “user emotion attribute(s)”) who provided the input and / or a response (e.g., also referred to as “response emotion attribute(s)”) to the input. The user's emotion attribute(s) may, for example,be linked to the user's emotion, where the user's emotion attribute(s) may contain one or more values associated with one or more emotional state labels. Furthermore, the response's emotion attributes may be linked to the character's response's emotion, where the response's emotion attributes may contain one or more values associated with one or more emotional state labels.
[0022] Furthermore, in some examples, the generic attributes of the response may be associated with one or more labels associated with a character and / or the response, one or more values (e.g., intensity values) associated with the label(s), and / or the like. As described herein, the label(s) associated with the character and / or the response may include, but are not limited to, quality, toxicity, creativity, helpfulness, humor, accuracy, coherence, wordiness, and / or any other label. Thus, the labels associated with the character and / or the response may be linked to various attributes of the character that are entered, e.g., based on a biography and / or personality of the character, and / or updated during the conversation.
[0023] In some examples, the system(s) may then be configured to process at least a portion of the input data (e.g., the first text data, the character information, etc.) and / or at least a portion of the first outputs of the first model(s) using one or more second machine learning models (referred to in some examples as "second model(s)" and / or "control model(s)"), such as one or more language models. As described herein, the second model(s) may be configured to control the emotions associated with the outputs, e.g., by controlling the emotions of the character that issues the outputs.For example, based at least on processing the at least part of the input data and / or the at least part of the first output, the second model(s) may generate and / or output a second output representing (1) output text (e.g., speech synthesis markup language, etc.) and / or (2) one or more voice tags associated with emotions (e.g., one or more emotional states) and / or one or more speech features. For example, if the input text includes a prompt, the output text may represent a response to the prompt. As described herein, a feature associated with speech may include, but is not limited to, pitch, volume, tone, rhythm, timbre, intensity, resonance, tempo, and / or any other feature.
[0024] In some examples, the system(s) may then be configured to process at least a portion of the second output with one or more third machine learning models (referred to in some examples as "third model(s)"), e.g., one or more text-to-speech (TTS) models. For example, the third model(s) may generate and / or output audio data representing speech output corresponding to the output text based at least on the processing of the at least a portion of the second output. For example, if the output text includes a response to the prompt, the speech output may include one or more words corresponding to the response. As described herein, the third model(s) may further generate the speech output to express the emotions corresponding to the voice tag(s) of the second output. For example, if the second outputindicates that the emotion is calm with medium volume, then the speech output can be played with a calm voice at medium volume.
[0025] In some examples, the system(s) may continue these processes if the system(s) receive further input (e.g., additional prompts) from the user. For example, if the user continues conversing with the character, the system may continue these processes to generate speech output corresponding to responses to the user's prompts. As described herein, by performing these processes, one or more parts of a response and / or one or more of the responses (e.g., each response) may be associated with a corresponding emotion such that the emotion of the speech output and / or the character changes during the conversation. In this way, the system(s) may better animate the character during the conversation than the conventional systems described above.
[0026] Although the use of two or more models in a pipeline is described, in some embodiments a single model may be used to accomplish each of these tasks, such as an LLM or a VLM, without limitations.
[0027] In some examples, the system(s) may train one or more of the models described herein using one or more techniques. For example, during a first training stage, the system(s) may train the second model(s) to produce outputs representing (1) instances of output text (e.g., responses) associated with instances of input text (e.g., prompts) and / or (2) voice tags associated with emotions and / or voice characteristics corresponding to the instances of output text. During the first training stage, the second model(s) may be trained using training input data representing the input text instances, attributes (e.g., response attributes, emotion attributes, etc.), and / or character information, as well as corresponding ground truth data representing the output text instances and the voice tags.
[0028] Before, during, and / or after the first training phase, the system may train the first model(s) and / or further train the second model(s) in a second training phase. In a first example, the system may train the first model(s) to generate outputs representing the attributes (e.g., the response attributes, the emotion attributes, etc.). In such an example, the first model(s) may be trained using training input data representing user data (e.g., text, images, audio, etc.), instances of input text, character information, and / or dialogue history, and corresponding ground-truth data representing the attributes.In a second example, the system(s) may train the first model(s) to produce outputs representing the attributes, while also training the second model(s) to produce outputs representing (1) the instances of output text (e.g., responses) associated with the instances of input text (e.g., prompts) and / or (2) the voice tags corresponding to the instances of output text. In this second example, the system(s) may train the first model(s) and / or the second model(s) using training input data representing the user data (e.g., text, images, audio, etc.), the instances of input text, the character information, and / or the dialogue history, as well as corresponding ground truth data representing the instances of output text and the voice tags.
[0029] As described in more detail herein, the system(s) may use one or more techniques to generate the training data, such as the training input data and / or the ground truth data. In some examples, the system may generate at least a portion of the training data using inputs from users, where the inputs specify the information (e.g., the instances of input text, the attributes, the character information, the instances of output text, the voice tags, etc.) represented by the training data. Additionally or alternatively, in some examples, the system(s) may generate at least a portion of the training data using one or more additional machine learning models.For example, for an instance of training data associated with an instance of input text, the system may use the additional model(s) to generate at least the emotion attribute(s), the reaction attribute(s), the instance of output text, and / or the voice tag(s) associated with the instance of input text. Furthermore, the system(s) may perform similar processes to generate any number of instances of training data for training the first model(s) and / or the second model(s).
[0030] While the examples herein describe the first model(s) being used to generate the first output, the second model(s) being used to generate the second output, and the third model(s) being used to generate the audio data, in other examples, one or more of the first model(s), the second model(s), or the third model(s) may be combined. In a first example, the first model(s) and the second model(s) may be combined such that the combined model(s) receive the inputs described herein with respect to the first model(s) and then output the second outputs described herein with respect to the second model(s).In a second example, the first model(s), the second model(s), and the third model(s) may be combined such that the combined model(s) receive the inputs described herein with respect to the first model(s) and then output the outputs described herein with respect to the third model(s). For a third example, the second model(s) and the third model(s) may be combined such that the combined model(s) receive the inputs described herein with respect to the second model(s) and then output the audio data described herein with respect to the third model(s).
[0031] The systems and methods described herein may be used, without limitation, by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, guided and unguided robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, shuttles, emergency vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Furthermore, the systems and methods described herein may be used for a variety of purposes, including, without limitation, machine control, machine locomotion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.
[0032] The described embodiments may be included in a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented with a robot, aviation systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented with an edge device, systems implementing large language models (LLMs), systems implementing one or more vision language models (VLMs), systems containing one or more virtual machines (VMs), systems for performing operations for generating synthetic data, systems implemented at least partially in a data center,Systems for performing conversational AI operations, systems for performing light transport simulations, systems for collaborative content creation for 3D assets, systems for performing generative AI operations, systems implemented at least in part using cloud computing resources, and / or other types of systems.
[0033] Fig. 1 shows an example data flow diagram of a process 100 for using one or more machine learning models to generate speech output expressing emotions, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are only examples. Other arrangements and elements (e.g., engines, interfaces, functions, orderings, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as individual or distributed components or in conjunction with other components, and in any suitable combination and location.Various functions described herein performed by devices may be performed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.
[0034] The process 100 may include one or more processing components 102 that receive user data 104 associated with a user. As described herein, the processing component(s) 102 may include one or more machine learning models (e.g., one or more audio-to-emotion models, one or more audio-to-gesture models, etc.), one or more neural networks, one or more algorithms, one or more modules, and / or any other type of processing component configured to perform the processes described herein with respect to the processing component(s) 102. Furthermore, the user data 104 may include, among other things, first text data (e.g., a portion of the input data 106) that may be input text (e.g.,a request (such as a question, a request, an instruction, a demand, a note, and / or any other type of speech) from the user (or converted from audio), audio data representing the user's speech (e.g., in the form of a spectrogram), image data or video data representing one or more images or videos in which the user is depicted, profile data representing information about the user (e.g., when the user gave consent to receive and / or store such information), and / or any other type of data. In some examples, at least a portion of the user data 104 may be generated using one or more client devices (e.g., one or more computing devices 1000).
[0035] The process 100 may then include the processing component(s) 102 that process at least a portion of the user data 104 and, at least based on the processing, generate user attribute data 108 associated with the user. In some examples, the user attribute data 108 may represent one or more labels (e.g., one or more tags) associated with one or more emotional states of the user. As described herein, an emotional state may include, but is not limited to, dull, indifferent, angry, disgusted, afraid, happy, sad, expectant, trusting, surprised, and / or any other emotional state. In some examples, the user attribute data 108 may also represent one or more values (e.g., one or more intensity scales) associated with the label(s).As described herein, a value associated with a label may be within a range, e.g., between 0 and 5 (and / or another range). For example, the user attribute data 108 may represent a first value associated with a first emotion tag, a second value associated with a second emotion tag, a third value associated with a third emotion tag, and / or so on.
[0036] As described herein, the processing component(s) 102 may perform various types of processing to generate the user attribute data 108. In a first example, e.g., if the user data 104 includes audio data representing the user's speech, the processing component(s) 102 may process the audio data using one or more audio processing techniques to determine the information represented by the user attribute data 108. In a second example, e.g., if the user data 104 includes image data representing the user's image(s), the processing component(s) 102 may process the image data using one or more image processing techniques to determine the information represented by the user attribute data 108. While the example in Fig. 1 shows that the processing component(s) 102 is (are) separate from one or more language models 110, in other examples the processing component(s) 102 may be included as part of the language model(s) 110 (e.g., as one or more layers thereof).
[0037] The process 100 may then include the language model(s) 110 receiving data as input, such as at least a portion of the user attribute data 108, at least a portion of the input data 106, and / or at least a portion of the dialogue data 112. In some examples, the input data 106 for the language model(s) 110 may represent text input from the user, character information associated with a character configured for speech output, and / or any other type of information. Furthermore, the character information may include, but is not limited to, characteristics of the character (e.g., occupation, relationships, personality traits, etc.), past communications, current circumstances (e.g., current interactions with other characters, current location, current goals, etc.), and / or other information associated with the character and / or the application (e.g., game) associated with the character.Additionally, the dialogue data 112 may represent one or more prior communications between the user and the character, such as one or more prior instances of input text (e.g., one or more prior requests) and / or one or more prior instances of output text (e.g., one or more prior responses) associated with the prior instance(s) of input text.
[0038] The process 100 may then include the language model(s) 110 processing the input data and, at least based on the processing, generating and / or outputting response attribute data 114 associated with one or more generic response attributes and / or emotion attribute data 116 associated with one or more emotion attributes. As described herein, in some examples, the emotion attributes may be associated with one or more labels corresponding to one or more emotional states, one or more values (e.g., intensity values) associated with the label(s), and / or any other information. Furthermore, the emotion attributes may be linked to the user (e.g., also referred to as "user emotion attribute(s)") who provided the input and / or a response (e.g., also referred to as "response emotion attribute(s)") to the input.In addition, a value associated with a label may lie within a range, e.g., between 0 and 5 (and / or another range).
[0039] In a first example, the emotion attribute data 116 may represent emotion attributes of the user, including at least a first value for a first emotional state, a second value for a second emotional state, a third value for a third emotional state, and / or so on for the user. In a second example, the emotion attributes 116 may represent the emotion attributes of the response, including at least a first value for a first emotional state, a second value for a second emotional state, a third value for a third emotional state, and / or so on for the response. In some examples, only one label may be associated with a value greater than zero (e.g., for the emotion attribute(s) of the response). In some examples, more than one label may be associated with a corresponding value greater than zero.
[0040] In some examples, the generic attributes of the response may also be associated with one or more labels associated with a character and / or the response, one or more values (e.g., intensity values) of the label(s), and / or the like. As described herein, the label(s) associated with the character and / or the response may include, but are not limited to, quality, toxicity, creativity, helpfulness, humor, correctness, coherence, wordiness, and / or any other label. For example, the label(s) associated with the character and / or the response may be associated with various attributes of the character that are entered and / or updated during the conversation. Furthermore, a value associated with a label may be within a range, e.g., within a range between 0 and 5 (and / or another range).For example, the response attribute data 114 may include at least a first value for a first identifier, a second value for a second identifier, a third value for a third identifier, and / or so on.
[0041] Fig. For example, Figure 2 shows an example of the language model(s) 110 that determine(s) attributes associated with the input text, in accordance with some embodiments of the present disclosure. As shown, the language model(s) 110 may receive as input user data 202 (which may resemble and / or include at least a portion of the user data 104), user attribute data 204 (which may resemble and / or include at least a portion of the user attribute data 108), dialog data 206 (which may resemble and / or include at least a portion of the dialog data 112), and / or input data 208 (which may resemble and / or include at least a portion of the input data 106). In the example of Fig. 2, the input data 208 represents input text, such as a request containing "Good morning, can you tell me if we have reached the village?" The language model(s) 110 may then process the data and, at least based on the processing, generate and / or output response attribute data 210 (which may resemble and / or include at least a portion of the response attribute data 114), emotion attribute data 212 (which may resemble and / or include at least a portion of the emotion attribute data 116), and emotion attribute data 212 (which may resemble and / or include at least a portion of the emotion attribute data 116).
[0042] As shown, the response attribute data represents 210 values for various generic response attributes, such as 4 associated with quality, 0 associated with humor, 0 associated with toxicity, 2 associated with creativity, 4 associated with helpfulness, 4 associated with correctness, 4 associated with coherence, and 1 associated with complexity. In addition, the emotion attributes represent 212 values associated with response emotion attributes, such as 0 associated with monotonous, 0 associated with angry, 0 associated with calm, 0 associated with disgust, 0 associated with fearful, 5 associated with happy, 0 associated with humor, and 0 associated with sad. In addition, the emotion attributes represent 214 values associated with the user's emotion attributes, such as0 is associated with monotonous, 0 is associated with angry, 4 is associated with calm, 0 is associated with disgust, 0 is associated with fear, 4 is associated with happy, 0 is associated with humor, and 0 is associated with sad. These are just a few response attributes and a few emotion attributes. In other examples, response attribute data 210 may include one or more additional and / or alternative response attributes, and / or emotion attribute data 212 and 214 may include one or more additional and / or alternative emotion attributes.
[0043] Going back to the example of Fig. 1, the process 100 may include one or more language models 118 that receive data as input, such as at least a portion of the input data 106, at least a portion of the response attribute data 114, and at least a portion of the emotion attribute data 116. The process 100 may then include the language model(s) 118 processing the data and, at least based on the processing, generating output data 120. As in the example of Fig. 1, the output data 120 may include at least text data 122 and tag data 124. In some examples, the text data 122 may represent output text (e.g., speech synthesis markup language, etc.) associated with the input text, such as a response (e.g., information, instructions, a question, a request, etc.) to the prompt. In some examples, the tag data 124 may also represent one or more voice tags associated with at least one emotion (e.g., one or more emotional states) or one or more characteristics associated with speech. As described herein, a characteristic associated with speech may include, but is not limited to, pitch, volume, tone, rhythm, resonance, tempo, and / or any other characteristic.
[0044] Fig. For example, Figure 3 shows an example of the language model(s) 118 determining the emotions and output text associated with the input text, in accordance with some embodiments of the present disclosure. As shown, the language model(s) 118 may receive as input at least the input data 208, the response attribute data 210, the emotion attribute data 212, and the emotion attribute data 214. The language model(s) 118 may then process the data and, at least based on the processing, generate output data 302 (which may resemble and / or include at least a portion of the output data 120), including text data 304 (which may resemble and / or include at least a portion of the text data 122), and tag data 306 (which may resemble and / or include at least a portion of the tag data 124).
[0045] In the example of Fig. 3, the text data 122 may represent a response to the prompt associated with the input data 208. For example, the text data 122 may include the text "Good morning, traveler, you have reached the village." Furthermore, the tag data 306 may represent a prosody emotion tag with 3 for "calm" and 4 for "cheerful," a pitch tag with "medium," and a volume tag with "medium." These are just a few examples of voice tags that may be represented by the tag data 306. In other examples, the tag data 306 may represent one or more additional and / or alternative voice tags.
[0046] Returning to the example of Fig. 1, the process 100 may include one or more language models 126 receiving at least a portion of the output data 120 as input. The process 100 may then include the language model(s) 126 processing the data and, at least based on the processing, generating audio data 128 representing speech. As described herein, the speech represented by the audio data 128 may be associated with the text represented by the text data 122 (e.g., include the words of the text). Furthermore, the speech may be expressed based at least on the emotion and / or feature(s) associated with the speech as represented by the tag data 124. For example, the audio data 128 may cause the speech to be spoken using one or more emotional states associated with the tag(s) represented by the tag data 124.Furthermore, the audio data 128 may cause the speech to be spoken using the values of the features associated with the speech, such as volume, pitch, speed, any accentuation, and / or the like, as represented by the tag(s). In other words, the speech model(s) 126 may be configured to generate the audio data 128 such that the character delivers the speech in a manner that expresses the emotion.
[0047] Fig. For example, Figure 4 shows an example of generating speech expressing emotions in accordance with some embodiments of the present disclosure. As shown, the language model(s) 126 may receive at least a portion of the output data 302 as input. The language model(s) 126 may then process at least the portion of the output data 302 and, at least based on the processing, generate audio data 402 (which may be similar to and / or include the audio data 128) representing speech. As shown, the speech includes the text "I'm fine, it's nice to see you today." The audio data 402 may then be used to cause a character 404 to output the speech, which may be represented by 406.For example, the character 404 may output the speech in a manner that emphasizes the emotion and / or feature(s) associated with the speech represented by the output data 302.
[0048] Back to the example of Fig. 1: While the example of Fig. 1 depicts the processing component(s) 102, the language model(s) 110, the language model(s) 118, and the language model(s) 126 as separate, in other examples, the processing component(s) 102, the language model(s) 110, the language model(s) 118, and the language model(s) 126 may be combined into one or more models. For example, the language model(s) 110 and the language model(s) 118 may be combined into a single language model that performs one or more of the processes described herein with respect to the language model(s) 110 and the language model(s) 118.
[0049] Furthermore, a language model as described herein may include, but is not limited to, a statistical language model, a neural language model, a probabilistic language model, a large-scale language model, a vision language model, and / or any other type of language model. While the examples herein illustrate the use of language models to perform the process 100 of Fig. 1, in other examples, any other type of model may be used to describe at least part of the process 100 of Fig. 1. For example, the language model(s) 110, the language model(s) 118, and / or the language model(s) 126 may include any other type of model configured to perform at least a portion of the processes described herein.
[0050] Furthermore, the process 100 may be performed with a single computing device, such as a computing device 1000 and / or a data center 1100, that includes the various components and / or language models, and / or with multiple computing devices, such as multiple computing devices 1000 and / or data centers 1100, each including at least a portion of the components and / or language models and communicating with each other over one or more networks.
[0051] As described herein, in some examples, the language model(s) 110 and / or the language model(s) 118 may be trained at different training levels. Fig. For example, Figure 5A shows a data flow diagram illustrating a process 600 for training one or more language models to generate speech output expressing emotions during a first training stage, in accordance with some embodiments of the present disclosure. As shown, the language model(s) 118 may be trained using training input data 502. In some examples, the training input data 502 may include response attribute data 504 that is similar to and / or includes response attribute data 114 and emotion attribute data 506 that is similar to and / or includes emotion attribute data 116. Additionally, the training input data 502 may include text data 508 representing one or more instances of input text, such as one or more prompts. The training input data 502 may be synthetic (e.g., generated based on computer models or renderings), real (e.g.,designed and produced based on real data), machine-annotated (e.g., using feature analysis and learning to extract features from data and then generate labels), human-annotated (e.g., a labeler or annotation expert determines the position of the labels), and / or a combination thereof.
[0052] The language model(s) 118 may be trained with the training input data 502 as well as the corresponding ground truth data 510. The ground truth data 510 may include annotations, labels, masks, and / or the like. As shown, the ground truth data 510 may include, for example, at least text data 512, which may be similar to and / or include the text data 122, and tag data 514, which may be similar to and / or include the tag data 124. For example, the text data 512 may represent one or more instances of output text (e.g., speech synthesis markup language, etc.) associated with the instance(s) of input text, such as one or more responses to the one or more prompts. Furthermore, in some examples, the tag data 124 may represent one or more tags associated with at least one emotion (e.g.,one or more emotional states) or one or more features associated with language. The ground truth data 510 may be synthetic (e.g., generated from computer models or renderings), real (e.g., designed and manufactured from real data), machine-generated (e.g., using feature analysis and learning to extract features from data and then generate labels), human-annotated (e.g., a labeler or annotator determines the position of the labels), and / or a combination thereof. In some examples, there may be corresponding ground truth data 510 for each instance of the training input data 502.
[0053] As in Fig. 5A further illustrates, a training engine 516 may use one or more loss functions that measure the loss (e.g., error) in the outputs 518 compared to the ground truth data 510. In some examples, the outputs 518 may be similar to and / or include the output data 120. Any type of loss function may be used, such as cross-entropy loss, mean squared error, mean absolute error, mean bias error, and / or other loss function types. In some examples, different outputs 518 may have different loss functions. For example, the output text instance(s) may have a first loss function and the voice tag(s) may have a second loss function. In such examples, the loss functions may be combined to form an overall loss, and the overall loss may be used to train the language model(s) 118 (e.g., to update the parameters).In each example, back-calculations may be performed to recursively calculate the gradients of the loss function(s) with respect to the training parameters. In some examples, the weights and bias values of the language model(s) 118 may be used to calculate these gradients.
[0054] Next, Fig. 5B is a data flow diagram illustrating a process 520 for training one or more language models to generate speech expressing emotions during a second training stage, in accordance with some embodiments of the present disclosure. As shown, the language model(s) 110 of FIG. 1 and / or the language model(s) 118 may be trained using training input data 522. In some examples, the input training data 522 may be similar to the data generated during the process 100 of FIG. Fig. 1 input to the language model(s) 110, be similar to and / or include at least a portion thereof. For example, the training input data 522 may be similar to and / or include at least a portion of the user data 104, the input data 106, the user attribute data 108, and / or the dialogue data 112. The input training data 522 may be synthetic (e.g., generated based on computer models or renderings), real (e.g., designed and manufactured based on real data), machine-generated (e.g., using feature analysis and learning to extract features from data and then generate labels), human-annotated (e.g., a labeler or annotator determines the position of the labels), and / or a combination thereof.
[0055] The language model(s) 110 and / or the language model(s) 118 may be trained with the input training data 522 and the corresponding ground truth data 524. The ground truth data 524 may include annotations, labels, masks, and / or the like. For example, the ground truth data 524 may include at least the response attribute data 504, the emotion attribute data 506, the text data 512, and / or the tag data 514 (see figure). As described herein, the ground truth data 524 may be synthetic (e.g., generated based on computer models or renderings), real (e.g., designed and manufactured based on real data), machine-generated (e.g., using feature analysis and learning to extract features from data and then generate labels), human-annotated (e.g., a labeler or annotator determines the position of the labels), and / or a combination thereof.In some examples, for each instance of the input training data 522, there may be corresponding ground truth data 524.
[0056] As described herein, in some examples, the second stage of training may be performed using various techniques. For example, in a first technique involving only training the language model(s) 110, the training engine 516 may use one or more loss functions that measure the loss (e.g., error) in the outputs 526 compared to the ground truth data 524. In such examples, the outputs 526 may resemble and / or include the response attribute data 114 and / or the emotion attribute data 116. Any type of loss function may be used, such as cross entropy loss, mean squared error, mean absolute error, mean bias error, and / or other types of loss functions. In some examples, different outputs 526 may have different loss functions.For example, the generic response attributes may have a first loss function, the emotional response attributes a second loss function, and the emotional user attributes a third loss function. In such examples, the loss functions may be combined to form an overall loss, and the overall loss may be used to train the language model(s) 110 (e.g., to update the parameters). In each example, backward computations may be performed to recursively calculate the gradients of the loss function(s) with respect to the training parameters. In some examples, the weights and bias values of the language model(s) 110 may be used to calculate these gradients.
[0057] Additionally or alternatively, in some examples, the second stage of training may comprise a second technique that includes training the language model(s) 110 and / or the language model(s) 118. For example, the training engine 516 may use one or more loss functions that measure the loss (e.g., error) in the outputs 528 compared to the ground truth data 524, where the outputs 528 may be similar to the output data 120. Any type of loss function may be used, such as cross-entropy loss, mean square error, mean absolute error, mean bias error, and / or other loss function types. In some examples, different outputs 528 may have different loss functions. For example, the output text instance(s) may have a first loss function, and the voice tag(s) may have a second loss function.In such examples, the loss functions may be combined to form an overall loss, and the overall loss may be used to train the language model(s) 110 and / or the language model(s) 118 (e.g., to update the parameters). In each example, backward computations may be performed to recursively calculate the gradients of the loss function(s) with respect to the training parameters. In some examples, the weights and biases of the language model(s) 110 and / or the language model(s) 118 may be used to calculate these gradients.
[0058] As described herein, in some examples, at least a portion of the training input data and / or the ground truth data used to train the language model(s) 110 and / or the language model(s) 118 may be generated using one or more models. Fig. For example, Figure 6A illustrates an example of generating response attributes for training one or more language models in accordance with some embodiments of the present disclosure. As illustrated, one or more models 602 may receive input data 604. As described herein, in some examples, the input data 604 may represent one or more emotional responses, similar to the text data 122, a dialogue history, similar to the dialogue data 112, and / or other data. Additionally, the outputs of the model(s) 602 may include the response attribute data 504 used to train the language model(s) 110 and / or the language model(s) 118. In some examples, the model(s) 602 may include a regression layer that uses a language model token output embedding to predict the response attribute data 504.However, in other examples, the model(s) 602 may include any other type of model.
[0059] Fig. 6B also illustrates an example of generating emotion attributes for training one or more language models in accordance with some embodiments of the present disclosure. As illustrated, one or more models 606 may receive input data 608. As described herein, in some examples, the input data 608 may represent one or more emotional responses, similar to the text data 122, a dialogue history, similar to the dialogue data 112, and / or other data. Furthermore, the output of the model(s) 606 may include the emotion attribute data 506 used to train the language model(s) 110 and / or the language model(s) 118. In some examples, the model(s) 606 may include a regression layer that uses a language model token output embedding to predict the emotion attribute data 506.However, in other examples, the model(s) 606 may include any other type of model.
[0060] In addition, Fig. 6C illustrates an example of generating ground truth data for training one or more language models in accordance with some embodiments of the present disclosure. As illustrated, one or more models 610 may receive input data 612. As described herein, in some examples, the input data 612 may represent one or more emotional responses, similar to the text data 122, a dialogue history, similar to the dialogue data 112, and / or other data. Additionally, the outputs of the model(s) 610 may include the ground truth data 510 used to train the language model(s) 110 and / or the language model(s) 118. In some examples, the model(s) 610 may be an autoregressive decoder-only transformer model that sequentially generates SSML-tagged responses. In other examples, the model(s) 610 mayHowever, the 610 models can also include any other type of model.
[0061] In some examples, at least a portion of the input data 604 may be from the example of Fig. 6A, at least a portion of the input data 608 from the example of Fig. 6B and / or at least part of the input data 612 from the example of Fig. 6C may contain the same input data. Additionally, in some examples, at least a portion of the input data 604 from the example of Fig. 6A, at least part of the input data 608 from the example of Fig. 6B and / or at least part of the input data 612 from the example of Fig. 6C at least part of the training input data 502 from the example of Fig. 5A and / or at least part of the training input data 522 from the example of Fig. 5B. In this way, the model(s) 602, the model(s) 606, and / or the model(s) 610 may process the same input data to generate instances of training data for training the language model(s) 110 and / or the language model(s) 118.
[0062] Each block of the methods 700, 800, and 900 described herein comprises a computational process that may be performed using any combination of hardware, firmware, and / or software (see Fig. 7-9). Various functions may be performed, for example, by a processor executing instructions stored in memory. The methods 700, 800, and 900 may also be embodied in the form of computer-usable instructions stored on computer storage media. The methods 700, 800, and 900 may be provided by a standalone application, a service, a hosted service (standalone or in combination with another hosted service), or a plug-in for another product, to name a few. Furthermore, the methods 700, 800, and 900 are exemplified with reference to Fig. 1. However, these methods 700, 800, and 900 may additionally or alternatively be performed by any system or combination of systems, including, but not limited to, the systems described herein.
[0063] Fig. 7 illustrates a flowchart showing a method 700 for using models to generate speech output expressing emotions, in accordance with some embodiments of the present disclosure. The method 700 may include, at block B702, generating a first output based on one or more first models processing user data that represents one or more attributes associated with emotions. For example, the language model(s) 110 may receive the user data 104 and / or the user attribute data 108 as input. In some examples, the language model(s) 110 may also receive the input data 106 and / or the dialogue data 112 as input. The language model(s) 110 may then process the data to generate the first output that includes the emotion attribute data 116.As described herein, in some examples, the emotion attribute data 116 may represent the attribute(s) associated with emotions, such as one or more response emotion attributes and / or one or more user emotion attributes. In some examples, the first output may also include the response attribute data 114.
[0064] The method 700 may include, at block B704, generating a second output representing one or more voice tags and second text related to the first text, at least based on one or more second models that process the first output and the first text. For example, the language model(s) 118 may receive as input the emotion attributes 116 and the input data 106 representing the first text. In some examples, the language model(s) 118 may also receive as input the response attribute data 114 and / or the input data 106 representing information associated with the character and / or the application. The language model(s) 118 may then process the data to generate the output data 120.As described herein, the output data 120 may include at least the text data 122 representing the second text and the tag data 124 representing the voice tag(s) associated with the second text. In some examples, the voice tags may include one or more emotion tags and / or one or more speech characteristic tags.
[0065] The method 700 may include, at block B706, generating audio data representing speech based at least on one or more third models that process the second output corresponding to the second text and present in a voice associated with the one or more voice tags. For example, the language model(s) 126 may receive the output data 120 from the language model(s) 118 as input. The language model(s) 126 may then process the output data 120 to generate the audio data 128 representing the speech corresponding to the second text present in the voice associated with the voice tag(s). The voice may, for example, be represented by emotions associated with the emotion tags and / or may include one or more features associated with the speech feature tags.
[0066] The method 700 may cause the output of the speech represented by the audio data at block B708. For example, the speech represented by the audio data 128 may be output, e.g., by one or more characters. In some examples, by performing the processes described herein, the characters may output the speech such that the characters exhibit an emotion.
[0067] Fig. 8 shows a flowchart illustrating a method 800 for using one or more emotion control models to generate speech output according to some embodiments of the present disclosure. The method 800 may, at block B802, generate a second output representative of second text and information associated with a voice associated with the emotion based at least on one or more language models processing first data representative of one or more attributes associated with emotions and first text. For example, the language model(s) 118 may receive as input the emotion attribute(s) 116 data representing the attribute(s) associated with the emotion and the input data 106 representing the first text.In some examples, the language model(s) 118 may also receive as input the response attribute data 114 and / or the input data representing information associated with the character and / or the application. The language model(s) 118 may then process the data to generate the output data 120. As described herein, the output data 120 may include at least the text data 122 representing the second text and the tag data 124 representing the information, such as the voice tags associated with the second text.
[0068] The method 800 may include, at block B804, generating audio data representing speech based at least on the one or more language models that process the second data corresponding to the second text and that are based at least on the information. For example, the language model(s) 126 may receive the output data 120 of the language model(s) 118 as input. The language model(s) 126 may then process the output data 120 to generate the audio data 128 representing the speech corresponding to the second text contained in the voice associated with the information. The voice may, for example, be represented by an emotion associated with the emotion tag(s) and / or include one or more features associated with the speech feature tag(s).
[0069] Method 800 may cause the output of the speech represented by the audio data at block B806. For example, the speech represented by the audio data 128 may be output, e.g., by one or more characters. In some examples, by performing the processes described herein, the characters may output the speech such that the characters exhibit an emotion.
[0070] Fig. 9 illustrates a flowchart showing a method 900 for using one or more models that manage emotion dialogue to generate speech, in accordance with some embodiments of the present disclosure. The method 900 may include, at block B902, obtaining data associated with at least one user or character of an application. For example, the user data 104 associated with the user, the user attribute data 108 associated with the user, and / or the input data 106 representing character information may be obtained. As described herein, the user may have a conversation with the character. As such, and in some examples, additional data may be obtained, such as the dialogue data 112 representing the conversation between the user and the character.
[0071] The method 900 may include, at block B904, generating an output representing one or more attributes associated with emotions, at least based on one or more language models that at least process the data. For example, the language model(s) 110 may process the user data 104, the user attribute data 108, and / or the input data. At least based on the processing, the language model(s) 110 may generate and / or output the emotion attribute data 116 representing the attribute(s) associated with the emotion. As described herein, the attribute(s) may include one or more response emotion attributes and / or one or more user emotion attributes. In some examples, the language model(s) maythe language models) 110 may further generate and / or output, at least based on the processing, the response attribute data 114 representing one or more generic response attributes.
[0072] The method 900 may, at block B906, generate audio data representing speech corresponding to a voice associated with the emotion based at least on the outputs. For example, the language model(s) 118 may process the response attribute data 114 and / or the emotion attribute data 116 to generate the output data 120. The language model(s) 126 may then process the output data 120 to generate the audio data 126 representing the speech. As described herein, the speech may be in a voice associated with the emotion.
[0073] Method 900 may cause the output of the speech represented by the audio data at block B908. For example, the speech represented by the audio data 128 may be output, e.g., by one or more characters. In some examples, by performing the processes described herein, the characters may output the speech such that the characters exhibit an emotion. EXAMPLE COMPUTER DEVICE
[0074] Fig. 8 is a block diagram of an example computing device 1000 suitable for use in implementing some embodiments of the present disclosure. Computing device 1000 may include an interconnect system 1002 that directly or indirectly interconnects the following devices: memory 1004, one or more central processing units (CPUs) 1006, one or more graphics processing units (GPUs) 1008, a communications interface 1010, input / output (I / O) ports 1012, input / output components 1014, a power supply 1016, one or more presentation components 1018 (e.g., display(s)), and one or more logic units 1020. In at least one embodiment, computing device(s) 1000 may include one or more virtual machines (VMs), and / or each of the components thereof may include virtual components (e.g., virtual hardware components).As non-limiting examples, one or more GPUs 1008 may include one or more vGPUs, one or more CPUs 1006 may include one or more vCPUs, and / or one or more logic units 1020 may include one or more virtual logic units. A computing device (or multiple computing devices) 1000 may include discrete components (e.g., a full GPU for computing device 1000), virtual components (e.g., a portion of a GPU for computing device 1000), or a combination thereof.
[0075] Although the different blocks in Fig. 8 are shown as being connected with lines via the interconnect system 1002, this is not to be understood as a limitation and is for clarity only. For example, in some embodiments, a presentation component 1018, such as a display, may be considered an I / O component 1014 (e.g., if the display is a touchscreen). Another example is that the CPUs 1006 and / or the GPUs 1008 may include memory (e.g., the memory 1004 may represent any memory device in addition to the memory of the GPUs 1008, the CPUs 1006, and / or other components). In other words, the computing device of Fig. 8 is merely illustrative. It does not distinguish between categories such as ‘workstation’, ‘server’, ‘laptop’, ‘desktop’, ‘tablet’, ‘client device’, ‘mobile device’, ‘handheld device’, ‘game console’, ‘electronic control unit (ECU)’, ‘virtual reality system’ and / or other device or system types, since all components of the computing device are Fig. 8 should be considered.
[0076] The interconnect system 1002 may represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 1002 may include one or more bus or connection types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another bus or connection type. In some embodiments, there are direct connections between the components. For example, the CPU 1006 may be directly connected to the memory 1004. Additionally, the CPU 1006 may be directly connected to the GPU 1008. For direct or point-to-point connections between components, the interconnect system 1002 may include a PCIe connection to establish the connection.In these examples, computing device 1000 may not necessarily include a PCI bus.
[0077] The memory 1004 may consist of a variety of computer-readable media. The computer-readable media may be any available media accessible by the computing device 1000. The computer-readable media may include both volatile and non-volatile media, as well as removable and non-removable media. By way of example and without limitation, the computer-readable media may include computer storage media and communication media.
[0078] The computer storage media may include both volatile and non-volatile media and / or removable and non-removable media, as implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other types of data. For example, the memory 1004 may store computer-readable instructions (e.g., programs and / or program elements, such as an operating system). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, Digital Versatile Disks (DVD) or other optical disk storage, magnetic cartridges, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 1000. As used herein, computer storage media does not per se include signals.
[0079] Computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any media for transmitting information. The term "modulated data signal" may refer to a signal in which one or more of its properties are adjusted or altered to encode information in the signal. Examples of computer storage media include wired media, such as a wired network or a direct wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the above media should also be considered computer-readable media.
[0080] The CPU(s) 1006 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 and to perform one or more of the methods and / or processes described herein. The CPU(s) 1006 may each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of concurrently executing a plurality of software threads. The CPU(s) 1006 may include any type of processor and may include different types of processors depending on the type of computing device 1000 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 1000, the processor may be, for example, an Advanced RISC Machines (ARM) processor using Reduced Instruction Set Computing (RISC) or an x86 processor using Complex Instruction Set Computing (CISC). Computing device 1000 may include one or more CPUs 1006 in addition to one or more microprocessors or additional coprocessors, such as math coprocessors.
[0081] In addition to or alternatively to the CPU(s) 1006, the GPU(s) 1008 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the methods and / or processes described herein. One or more of the GPU(s) 1008 may be an integrated GPU (e.g., with one or more of the CPU(s) 1006 and / or one or more of the GPU(s) 1008 may be a discrete GPU. In certain embodiments, one or more of the GPU(s) 1008 may be a coprocessor of one or more of the CPU(s) 1006. The GPU(s) 1008 may be used by the computing device 1000 to render graphics (e.g., 3D graphics) or to perform general purpose computations. The GPU(s) 1008 may be used, for example, for General-Purpose Computing on GPUs (GPGPU).The GPU(s) 1008 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU(s) 1008 may generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 1006 received via a host interface). The GPU(s) 1008 may include graphics memory, e.g., display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory may be part of the memory 1004. The GPU(s) 1008 may include two or more GPUs operating in parallel (e.g., via a link). The link may connect the GPUs directly (e.g., using NVLINK) or connect the GPUs via a switch (e.g., using NVSwitch). When used together, each GPU can generate 1008 pixel data or GPGPU data for different parts of an output or for different outputs (e.g.a first GPU for a first image and a second GPU for a second image). Each GPU can have its own memory or share memory with other GPUs.
[0082] In addition to or alternatively to the CPU(s) 1006 and / or the GPU(s) 1008, the logic unit(s) 1020 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 1006, the GPU(s) 1008, and / or the logic unit(s) 1020 may discretely or jointly execute any combination of the methods, processes, and / or portions thereof. One or more of the logic units 1020 may be part of one or more of the CPU(s) 1006 and / or the GPU(s) 1008, and / or one or more of the logic units 1020 may be discrete components or otherwise external to the CPU(s) 1006 and / or the GPU(s) 1008.In certain embodiments, one or more of the logic units 1020 may be a coprocessor of one or more of the CPU(s) 1006 and / or one or more of the GPU(s) 1008.
[0083] Examples of the logical unit(s) 1020 include one or more processing cores and / or components thereof, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel visual cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like.
[0084] The communication interface 1010 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1000 to communicate with other computing devices over an electronic communication network, including wired and / or wireless communication. The communication interface 1010 may include components and functions that enable communication over a variety of networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the logic unit(s) 1020 and / or the communication interface 1010 may include one or more data processing units (DPUs) to transfer the data received over a network and / or via the interconnect system 1002 directly to one or more GPU(s) 1008 (e.g., a memory).
[0085] Through the I / O ports 1012, the computing device 1000 can be logically connected to other devices, including the I / O components 1014, the presentation component(s) 1018, and / or other components, some of which may be built into (e.g., integrated) the computing device 1000. Example I / O components 1014 include a microphone, a mouse, a keyboard, a joystick, a gamepad, a game controller, a satellite dish, a scanner, a printer, a wireless device, etc. The I / O components 1014 can provide a natural user interface (NUI) that processes air gestures, voice inputs, or other physiological inputs from a user. In some cases, the inputs can be transmitted to an appropriate network element for further processing.An NUI may implement any combination of speech recognition, pen recognition, facial recognition, biometric recognition, both on-screen and off-screen gesture recognition, air gestures, head and eye tracking, and touch detection (as described in more detail below) in conjunction with a display of computing device 1000. Computing device 1000 may include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture recognition and capture. Additionally, computing device 1000 may include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable the detection of motion. In some examples, the output of the accelerometers or gyroscopes from computing device 1000 may be used to present immersive augmented reality or virtual reality.
[0086] Power supply 1016 may be a hardwired power supply, a battery power supply, or a combination thereof. Power supply 1016 may supply power to computing device 1000 so that components of computing device 1000 can operate.
[0087] The presentation component(s) 1018 may include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 1018 may receive data from other components (e.g., the GPU(s) 1008, the CPU(s) 1006, DPUs, etc.) and output the data (e.g., as an image, video, audio, etc.). EXAMPLE OF A DATA CENTER
[0088] Fig. Figure 9 shows an example of a data center 1100 that may be used in at least one embodiment of the present disclosure. The data center 1100 may include a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and / or an application layer 1140.
[0089] As in Fig. 9, the infrastructure layer of the data center 1110 may include a resource orchestrator 1112, clustered computing resources 1114, and node computing resources (“KKR”) 1116(1)-1116(N), where “N” represents any natural number. In at least one embodiment, the KRRs 1116(1)-1116(N) may be any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output devices (NW I / O), network switches, virtual machines (VMs), power modules and / or cooling modules, etc. In some embodiments, one or more KKRs among the KKRs 1116(1)-1116(N) may correspond to a server having one or more of the above-mentioned computing resources.Furthermore, in some embodiments, the KRRs 1116(1)-1116(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of the KRRs 1116(1)-1116(N) may correspond to a virtual machine (VM).
[0090] In at least one embodiment, the grouped computing resources 1114 may include separate groupings of KRRs 1116 housed in one or more racks (not shown) or in many racks in data centers in different geographical locations (also not shown). Separate groupings of KRRs 1116 within the grouped computing resources 1114 may include grouped computing, networking, and / or memory resources that may be configured or allocated to support one or more workloads. In at least one embodiment, multiple KRRs 1116 with CPUs, GPUs, DPUs, and / or other processors may be grouped in one or more racks to provide computing resources to support one or more workloads. The one or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0091] Resource orchestrator 1112 may configure or otherwise control one or more KRRs 1116(1)-1116(N) and / or grouped computing resources 1114. In at least one embodiment, resource orchestrator 1112 may be a management entity for the software design infrastructure (SDI) of data center 1100. Resource orchestrator 1112 may be hardware, software, or a combination thereof.
[0092] In at least one embodiment, as in Fig. 9, the framework layer 1120 may include a job scheduler 1128, a configuration manager 1134, a resource manager 1136, and / or a distributed file system 1138. The framework layer 1120 may include a framework to support the software 1132 of the software layer 1130 and / or one or more applications 1142 of the application layer 1140. The software 1132 or the application(s) 1142 may include web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1120 may be some type of free and open source software web application framework such as, but not limited to, Apache Spark™ (hereinafter "Spark"), which may utilize the distributed file system 1138 for processing large amounts of data (e.g., "Big Data").In at least one embodiment, the job scheduler 1128 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 1100. The configuration manager 1134 may be capable of configuring various layers, such as the software layer 1130 and the framework layer 1120, including Spark and the distributed file system 1138, to support the processing of large amounts of data. The resource manager 1136 may manage clustered or grouped computing resources allocated to support the distributed file system 1138 and the job scheduler 1128. In at least one embodiment, the clustered or grouped computing resources may include the clustered computing resources 1114 at the infrastructure layer 1110 of the data center.The resource manager 1136 may coordinate with the resource orchestrator 1112 to manage these allocated or assigned computing resources.
[0093] In at least one embodiment, the software 1132 included in software layer 1130 may include software used by at least portions of KRRs 1116(1)-1116(N), clustered computing resources 1114, and / or distributed file system 1138 of framework layer 1120. One or more types of software may include, but are not limited to, web searching software, email virus scanning software, database software, and video content streaming software.
[0094] In at least one embodiment, the application(s) 1142 included in the application layer 1140 may include one or more types of applications used by at least portions of the KKRs 1116(1)-1116(N), the clustered compute resources 1114, and / or the distributed file system 1138 of the framework layer 1120. One or more types of applications may include any number of genomic applications, cognitive computation, and machine learning applications, including, but not limited to, training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in connection with one or more embodiments.
[0095] In at least one embodiment, the configuration manager 1134, the resource manager 1136, and the resource orchestrator 1112 may perform any number and type of modifying actions based on any amount and type of data collected in any technically feasible manner. Self-modifying actions may relieve the operator of a data center 1100 from potentially making poor configuration decisions and potentially avoid underutilized and / or malfunctioning parts of a data center.
[0096] Data center 1100 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weighting parameters according to a neural network architecture using software and / or computing resources described above with respect to data center 1100.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1100 by using weighting parameters calculated by one or more training techniques, such as, but not limited to, those described herein.
[0097] In at least one embodiment, data center 1100 may utilize CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference with the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service that enables the user to train or infer information, such as image recognition, speech recognition, or other artificial intelligence services. EXAMPLE OF NETWORK ENVIRONMENTS
[0098] Network environments suitable for implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., any single device) may reside on one or more instances of the computing device(s) 1000 of Fig. 8 - e.g., each device may include similar components, features, and / or functions of the computing device(s) 1000. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may also be part of a data center 1100, an example of which is shown in Fig. 9 is described in more detail.
[0099] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can comprise multiple networks or a network of networks. For example, the network can comprise one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network comprises a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0100] Compatible network environments include one or more peer-to-peer network environments—in which case, a server may not be included in a network environment—and one or more client-server network environments—in which case, one or more servers may be included in a network environment. In peer-to-peer network environments, the functionality described here can be implemented with respect to one or more servers on any number of client devices.
[0101] In at least one embodiment, a network environment may include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer may include a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or application(s) may each include web-based service software or applications. In certain embodiments, one or more of the client devices may utilize the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be, among other things, a type of free and open-source software framework for web applications that, for example, uses a distributed file system for processing large amounts of data (e.g., "Big Data").
[0102] A cloud-based network environment may provide cloud computing and / or cloud storage performing any combination of the computing and / or data storage functions described herein (or one or more portions thereof). Each of these various functions may be distributed across multiple locations from central servers or core servers (e.g., from one or more data centers that may be located across a state, region, country, globe, etc.). When a connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers may delegate at least some functionality to the edge server(s). A cloud-based network environment may be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0103] The client device(s) may include at least some of the components, features, and functions of the Fig.8. A client device may be, for example, a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a video camera, a surveillance device or system, a vehicle, a boat, an aircraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computing system, an embedded system controller, a remote control, an appliance, a consumer electronics device, a workstation, an edge device, any combination of these described devices, or any other suitable device.
[0104] The disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules executed by a computer or other machine, such as a personal data assistant or other handheld device. In general, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements specific abstract data types. The disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The disclosure may also be applied in distributed computing environments where tasks are performed by remotely controlled devices interconnected via a communications network.
[0105] As used herein, any reference to "and / or" in reference to two or more elements should be interpreted to mean only one element or a combination of elements. For example, "Element A, Element B, and / or Element C" may include only Element A, only Element B, only Element C, Element A and Element B, Element A and Element C, Element B and Element C, or both Elements A, B, and C. Furthermore, "at least one of Element A or Element B" may include at least one of Element A, at least one of Element B, or at least one of Element A and at least one of Element B.
[0106] The subject matter of the present disclosure is described herein with a certain degree of particularity to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors contemplated that the claimed subject matter may be embodied in other ways to incorporate various steps or combinations of steps similar to those described herein in connection with other present or future technologies. Although the terms "step" and / or "block" are used herein to refer to various elements of the methods employed, the terms should not be interpreted to imply any particular ordering among or between various steps described herein unless the order of each step is expressly described. EXAMPLE PARAGRAPHS A: A method comprising: generating, at least based on one or more first models processing user data, a first output representing one or more attributes associated with one or more emotional states; generating, at least based on one or more second models processing the first output and a prompt, a second output representing a response to the prompt and one or more tags corresponding to a voice associated with the one or more attributes; generating, at least based on the processing of the second output, audio data representing speech corresponding to the response and expressed using the voice; and causing output of the speech represented by the audio data. B: The method of paragraph A, comprising: generating a third output representing one or more second attributes associated with a character based at least on the one or more first models that process character data, wherein generating the second output is further based at least on the one or more second models that process the third output. C: The method of either paragraph A or paragraph B, comprising: obtaining character data representative of one or more second attributes associated with a character to output the speech, wherein generating the second output is further based on at least the one or more second models processing the character data. D: The method of any one of paragraphs AC, further comprising: obtaining data representative of at least one of one or more prior requests or one or more prior responses associated with the one or more prior requests, wherein generating the first output is further based on at least the one or more first models processing the data representative of the at least one of the one or more prior requests or the one or more prior responses. E: A method according to any one of paragraphs AD, wherein the user data comprises at least one of the following elements: text data representing text describing one or more second emotional states associated with a user; audio data representing user speech associated with the user; video data representing one or more videos associated with the user; or image data representing one or more images associated with the user. F: A system comprising: one or more processors for: generating second data, at least based on one or more language models that process the first data representing one or more emotional states and a first text, wherein the second data represents a second text and information associated with a voice related to the one or more emotional states; generating audio data, at least based on the second data, wherein the audio data represents speech corresponding to the second text and the audio data is expressed using the voice; and causing output of the speech represented by the audio data. G: The system of paragraph F, wherein: the one or more emotional states comprise at least a first emotional state associated with issuing a response corresponding to the second text and a second emotional state associated with issuing the response corresponding to the second text; and the first data further represents a first value associated with the first emotional state and a second value associated with the second emotional state. H: The system of either paragraph F or paragraph G, wherein: the one or more emotional states comprise at least a first emotional state associated with a user and a second emotional state associated with the user; and the first data further represents a first value associated with the first emotional state and a second value associated with the second emotional state. I: System according to any one of paragraphs FH, wherein the second data is further generated on the basis of at least one or more language models processing third data representative of one or more attributes associated with a character to output the language. J: The system of paragraph I, wherein the second data is representative of at least one of the following: one or more labels describing the one or more attributes associated with the character; or one or more intensity values associated with the one or more labels. K: The system of any one of paragraphs FJ, wherein the information associated with the voice relating to the one or more emotional states comprises at least one of the following: one or more second emotional states associated with the voice; one or more first values associated with the one or more second emotional states; one or more voice characteristics associated with the voice; or one or more second values associated with the one or more voice characteristics. L: The system of any one of paragraphs FK, wherein the one or more processors are further operable to: obtain third data associated with a user; and generate, at least based on the one or more language models processing the third data, the first data representing the one or more emotional states. M: The system of paragraph L, wherein the third data comprises at least one of the following: text data representing text describing one or more second emotional states associated with the user; audio data representing the user's speech associated with the user; or image data representing one or more images associated with the user. N: The system of any one of paragraphs FM, wherein the one or more processors are further operable to: generate, at least based on the one or more language models processing the third data representative of one or more second emotional states and third text, fourth data representative of fourth text and second information associated with a second voice related to the one or more second emotional states; generate, at least based on the fourth data, second audio data representative of second audio corresponding to the fourth audio and expressed by the second voice; and cause a second output of the second audio represented by the second audio data. O: The system of clause N, wherein the one or more processors are further operable to generate the third data representative of the one or more second emotional states based at least on the one or more language models processing fifth data associated with a user, the first text, and the second text. P: The system of any one of paragraphs FO, wherein: one or more first language models of the one or more language models generate the first data representative of the one or more emotional states; one or more second language models of the one or more language models generate the second data representative of the second text and the information associated with the voice relating to the one or more emotional states; and one or more third language models of the one or more language models generate the audio data representative of the speech. Q: System according to any one of paragraphs FP, wherein the system comprises at least one of the following elements: control system for an autonomous or semi-autonomous machine; perception system for an autonomous or semi-autonomous machine; system for performing one or more simulation operations; system for performing one or more digital twin operations; system for performing a light transport simulation; system for performing collaborative content creation for 3D assets; system for performing one or more deep learning operations; system implemented using an edge device; system implemented using a robot; system for performing one or more generative AI operations; system for performing operations using one or more large language models (LLMs); system for performing operations using one or more visual language models (VLMs);System for performing one or more conversational AI operations; System for generating synthetic data; System for presenting virtual reality, augmented reality, or mixed reality content; System including one or more virtual machines (VMs); System implemented at least in part in a data center; or System implemented at least in part using cloud computing resources. R: One or more processors comprising: processing circuitry for generating audio data representing speech, wherein the speech is associated with one or more emotional states, the audio data being generated based on at least one or more language models comprising first data associated with a user providing a prompt associated with a response, and for processing second data representative of one or more attributes associated with a character to output the speech. S: One or more processors according to paragraph R, wherein the processing circuitry is further arranged to: generate, at least based on the one or more language models comprising at least one of the first data or the second data, third data representative of the one or more emotional states or the one or more second emotional states associated with the user, wherein the audio data is generated at least based on the one or more language models further processing the second data and the third data. T: One or more processors according to either paragraph R or paragraph S, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs);A system for performing operations using one or more visual language models (VLMs); a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting virtual reality, augmented reality, or mixed reality content; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0107] It is to be understood that the aspects and embodiments described above are only exemplary and that changes in detail may be made within the scope of the claims.
[0108] Each apparatus, method, and feature described in the specification and (where appropriate) in the claims and drawings may be provided independently or in any suitable combination.
[0109] The reference signs contained in the claims are for illustrative purposes only and do not limit the scope of the claims.
Claims
[1] Procedure comprising: Generating, at least based on one or more first models processing user data, a first output representing one or more attributes associated with one or more emotional states; generating, at least based on one or more second models that process the first output and a prompt, a second output representing a response to the prompt and one or more tags corresponding to a voice related to the one or more attributes; generating, at least based on processing of the second output, audio data representing speech corresponding to the response and expressed by the voice; and Output the speech represented by the audio data. [2] A method according to claim 1, comprising: Generating, at least based on the one or more first models that process character data, a third output representing one or more second attributes associated with a character, wherein generating the second output is based at least on the one or more second models processing the third output. [3] Method according to one of the preceding claims, comprising: Obtaining character data representing one or more second attributes associated with a character that is to output the speech, wherein generating the second output is based at least on the one or more second models processing the character data. [4] Method according to one of the preceding claims, comprising: Obtaining data representing at least one of one or more previous requests or one or more previous responses associated with the one or more previous requests, wherein generating the first output is further based on at least the one or more first models processing the data representing the one or more prior prompts or the one or more prior responses. [5] Method according to one of the preceding claims, wherein the user data comprises at least one of the following elements: Text data representing text describing one or more second emotional states of a user; Audio data representing the user's speech that corresponds to the user Video data representing one or more videos corresponding to the user; or Image data representing one or more images corresponding to the user. [6] System comprising: one or more processors for: Generating, at least based on one or more language models that process first data representing one or more emotional states and first text, second data representing second text and information associated with a voice related to the one or more emotional states; generating, at least on the basis of the second data, audio data representing a speech corresponding to the second text and expressed by the voice; and Output the speech represented by the audio data. [7] System according to claim 6, wherein: the one or more emotional states comprise at least a first emotional state associated with the output of a response corresponding to the second text and a second emotional state associated with the output of the response corresponding to the second text; and the first data further represent a first value associated with the first emotional state and a second value associated with the second emotional state. [8] System according to one of claims 6 or 7, wherein: the one or more emotional states comprise at least a first emotional state associated with a user and a second emotional state associated with the user; and the first data are also representative of a first value associated with the first emotional state and a second value associated with the second emotional state. [9] The system of any one of claims 6 to 8, wherein the second data is further generated based on at least one or more language models that process third data representative of one or more attributes associated with a character to output the speech. [10] The system of claim 9, wherein the second data represents at least one of the following features: one or more labels describing the one or more attributes associated with the character; or one or more intensity values associated with the one or more labels. [11] The system of any of claims 6-10, wherein the information associated with the voice relating to the one or more emotional states comprises at least one of the following elements: one or more second emotional states associated with the voice; one or more first values associated with the one or more second emotional states; one or more vocal characteristics associated with the voice; or one or more second values associated with the one or more voice characteristics. [12] The system of any one of claims 6 to 11, wherein the one or more processors further serve to: Obtaining third-party data associated with a user; and Generating, at least based on the one or more language models processing the third data, from the first data representing the one or more emotional states. [13] The system of claim 12, wherein the third data comprises at least one of the following elements: Text data representing text describing one or more second emotional states of the user; Audio data representing the user’s language corresponding to the user; or Image data representing one or more images corresponding to the user. [14] The system of any one of claims 6 to 13, wherein the one or more processors further serve to: Generating, at least based on the one or more language models that process third data representing one or more second emotional states and third text, fourth data representing fourth text and second information associated with a second voice related to the one or more second emotional states; generating, at least on the basis of the fourth data, second audio data representing a second language corresponding to the fourth text and expressed with the second voice; and Initiating a second output of the second language represented by the second audio data. [15] The system of claim 14, wherein the one or more processors are further operable to generate the third data representative of the one or more second emotional states based at least on the one or more language models processing fifth data associated with a user, the first text, and the second text. [16] System according to any one of claims 6-15, wherein: one or more first language models of the one or more language models generate the first data representing the one or more emotional states; one or more second language models of the one or more language models generate the second data representing the second text and the information associated with the voice relating to the one or more emotional states; and one or more third language models of the one or more language models generate the audio data representing the speech. [17] System according to any one of claims 6-16, wherein the system comprises at least one of the following elements: Control system for an autonomous or semi-autonomous machine; Perception system for an autonomous or semi-autonomous machine; System for performing one or more simulation operations; System for performing one or more digital twin operations; System for performing light transport simulations; System for collaborative content creation for 3D assets; System for performing one or more deep learning operations; System implemented with an edge device; System implemented with the help of a robot; System for performing one or more generative AI operations; System for performing operations using one or more large language models (LLMs); System for performing operations using one or more visual language models (VLMs); System for performing one or more conversational AI operations; System for generating synthetic data; System for displaying virtual reality, augmented reality or mixed reality content; System that contains one or more virtual machines (VMs); System that is at least partially implemented in a data center; or System implemented at least in part using cloud computing resources. [18] One or more processors comprising: a processing circuit to generate audio data representing speech in a voice associated with one or more emotional states, wherein the audio data is generated based at least on one or more language models that process first data associated with a user providing a prompt associated with a response and process second data representative of one or more attributes associated with a character to output the speech. [19] One or more processors according to claim 18, wherein the processing circuitry is further to Generating, at least on the basis of the one or more language models that process at least one of the first data or the second data, third data representative of the one or more emotional states or the one or more second emotional states of the user, wherein the audio data is generated at least on the basis of the one or more speech models that further process the second data and the third data. [20] One or more processors according to one of claims 18 or 19, wherein the one or more processors are included in at least one of the following elements: Control system for an autonomous or semi-autonomous machine; Perception system for an autonomous or semi-autonomous machine; System for performing one or more simulation operations; System for performing one or more digital twin operations; System for performing light transport simulations; System for collaborative content creation for 3D assets; System for performing one or more deep learning operations; System implemented with an edge device; System implemented with the help of a robot; System for performing one or more generative AI operations; System for performing operations using one or more large language models (LLMs); System for performing operations using one or more visual language models (VLMs); System for performing one or more conversational AI operations; System for generating synthetic data; System for displaying virtual reality, augmented reality or mixed reality content; System that contains one or more virtual machines (VMs); System that is at least partially implemented in a data center; or System implemented at least in part using cloud computing resources.