Controllable Natural Paralanguage for Text-to-Speech Synthesis
By training machine learning models to annotate and generate speech with paralinguistic markers, the solution addresses the lack of paralinguistic cue utilization in dialogue systems, resulting in more natural and understandable interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-03-10
AI Technical Summary
Conventional computer dialogue systems fail to utilize and control paralinguistic cues, such as variations in pitch, duration, and stress, to convey additional meaning beyond the words themselves, leading to monotonous and confusing interactions.
Implementing machine learning models trained on paralinguistic effects to annotate and generate speech with markup markers, guiding a speech generation module to produce pronunciations that convey intended meanings through prosodic variations.
Enhances the naturalness and understandability of dialogue systems by incorporating nuanced paralinguistic cues, allowing for more human-like interactions and improved user comprehension.
Smart Images

Figure 2026041757000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 63 / 042,085, filed June 22, 2020, and entitled "Controllable, natural paralinguistics for text to speech synthesis," the disclosure of which is incorporated herein by reference in its entirety.
[0002] FIELD Embodiments of the present design generally relate to voices generated by computing devices. Embodiments of the present design also generally relate to computing devices that understand a user's speech input. [Background technology]
[0003] The same or similar words in a spoken sentence can convey different meanings depending on the type of acoustic cue used for those words within the spoken sentence. For example, "My loving husband" spoken with a cheerful acoustic cue can convey a sentiment of affection. "My loving husband" spoken with a long, slow pronunciation of the words "love" and "husband" and in a shouted voice can convey a sentiment of "dislike." In another example, "Yeah" spoken with a normal pronunciation can convey a sentiment of "yes, that's right." In contrast, "Yeah, Yeah, Yeah" spoken with a lengthening of each subsequent "Yeah" compared to the first "Yeah" and with a falling intonation of the phrase as it progresses through the "Yeahs" conveys a sentiment of sarcasm and a sense of "not very likely." People use a variety of acoustic cues (overlaid as additional information sources in addition to the words themselves) to convey or clarify meaning. For example, extremely common nonwords like most "normal" words, "huh" and "mm," can have multiple meanings, such as i) skepticism, ii) lack of understanding, iii) annoyance, and iv) others, often clarified by the speaker using acoustic cues indicating paralinguistic effects. A typical teenager may express familiar expressions, such as sarcasm or sarcasm, that fail to use paralinguistic effects through prosodic changes to convey secondary information when speaking with a voice digital assistant like Alexa or a non-native speaker of the language because they must be very verbatim, slow, and deliberate with words in order to successfully communicate with the other party. Limited verbatim computer-to-computer conversational interaction with computer systems now appears to be seen as the expected norm, not a problem.
[0004] Thus, many conventional computer dialogue systems have not made specific use of lower-level speech phenomena, such as word stress within a sentence for a particular discourse meaning, changes in the duration and / or pitch of pronounced words, etc., to convey additional information beyond the words themselves. Many conventional computer dialogue systems simply analyze pronounced words and / or non-words verbatim and ignore any paralinguistic effects conveyed along with the pronounced words and / or non-words. Similarly, many conventional dialogue systems simply speak to users in a monotonous manner, pronouncing words or non-words without any paralinguistic effects. Also, some dialogue systems often generate prosody in an arbitrary manner unrelated to the meaning or state of the discourse, which can be confusing and / or make some utterances difficult to understand.
[0005] Paralinguistic effects can often be highly nuanced and complex. Different patterns of similar lower-level speech phenomena, and where paralinguistic effects occur positionally within a word or sentence, can all potentially change the secondary meaning conveyed using a particular paralinguistic effect.
[0006] There are several types of artificial intelligence disciplines, such as robot artificial intelligence, natural language processing artificial intelligence, machine learning including deep learning artificial intelligence, and fuzzy logic artificial intelligence. There are at least three types of machine learning: i) supervised learning including semi-supervised learning, ii) unsupervised learning, and iii) reinforcement learning. In addition, there are many machine learning algorithms within the machine learning algorithm category, such as: 1- clustering algorithms (e.g., hierarchical clustering, k-means, etc.), 2- association rule learning algorithms (e.g., apriori algorithm, eicra algorithm, etc.), 3- neural network algorithms (e.g., perceptron algorithm, multilayer perceptron (MLP) algorithm, backpropagation algorithm, stochastic gradient descent algorithm, Hopfield network algorithm, radial basis function network (RBFN) algorithm), and deep learning neural network algorithms - convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), stacked autoencoder, deep Boltzmann machine (DBM). ), Deep Belief Networks (DBN), etc.), 4 - Dimensionality Reduction Algorithms (e.g., Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), etc.), 5 - Regularization Algorithms (e.g., Least Absolute Value Shrinkage and Selection Operator (LASSO), Elastic Net, etc.), 6 - Decision Tree Algorithms (e.g., Decision Strain, M5, Conditional Decision Tree, etc.), 7 - Bayesian Algorithms (e.g., Naive Bayes, Gaussian Naive Bayes, Bayesian Networks (BN), etc.), 8 - Regression Algorithms (e.g., Linear Regression, Logistic Regression, etc.), 9 - Instance-Based Algorithms (e.g., k-Nearest Neighbors (kNN), Self-Organizing Maps (SOM), Support Vector Machines (SVM), etc.), 10 - Ensemble Algorithms (e.g., Boosting, Stacked Generalization, Random Forest), etc., and many more not listed. Summary of the Invention
[0007] In an embodiment, one or more machine learning models are trained to examine audio data including at least one of i) words, ii) phonemes, and iii) non-words in speech communication. The i) words, ii) phonemes, and / or iii) non-words are annotated with one or more markup markers that guide the generation of a textual representation to cause a pronunciation that differs from the plain pronunciation that would occur in the absence of the markup markers in order to convey at least one of 1) additional intended meaning and 2) enhanced understanding of the enunciated i) phonemes, ii) words, and / or iii) non-words themselves. The one or more machine learning models are trained with training data of speakers using paralinguistic effects. A speech generation module can receive the generated text representation to guide the speech generation module so that the speech generation module is configured to produce speech with different pronunciations, with specific paralinguistic effects, in a manner that better conveys 1) additional intended meaning and / or 2) enhanced understanding of the pronounced i) phonemes, ii) words, and / or iii) non-words themselves.
[0008] In an embodiment, a training system may include at least a set of speech processing detectors, a speech recognition module, and one or more machine learning models. The speech recognition module receives training data of an audio stream of speech and produces 1) a text representation, 2) a time-aligned phonetic representation, or 3) any combination of both for individual words, individual non-words, individual phonemes, and any combination thereof in the training data of the audio stream of speech. The set of speech processing detectors analyzes the training data of the audio stream of speech from a person speaking. The set of speech processing detectors detects speech parameters indicative of one or more paralinguistic effects in addition to pronounced words, phonemes, and non-words in the audio stream. One or more machine learning models undergo supervised machine learning on their neural networks to train how to associate one or more markup markers with 1) a text representation, 2) a waveform representation, and 3) any combination of both, for each i) individual word, ii) individual non-word, iii) individual phoneme, and iv) any combination thereof pronounced with a particular paralinguistic effect, where each markup marker corresponds to its own paralinguistic effect. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram of an example conversational engagement platform embodiment that includes a dialogue management module that cooperates with one or more machine learning models that receive parameters tied to a user's speech input from other modules, where the machine learning models are trained to understand a user's communicated meaning, including potential paralinguistic effects, and to induce changes in prosody to convey the paralinguistic effects when speaking to the user. [Figure 2]FIG. 1 is a block diagram of an example training system for one or more machine learning models trained on paralinguistic effects conveyed by a person's speech and a set of markup markers that correlate to each particular paralinguistic effect. [Figure 3] 1 is a block diagram of several electronic systems and devices communicating with each other in a network environment according to an embodiment of the present design; [Figure 4] FIG. 1 is a block diagram of an embodiment of one or more computing devices that can be part of a conversation assistant for embodiments of the present design discussed herein.
[0010] While the present design is subject to various modifications, equivalents, and alternatives, specific embodiments thereof are shown by way of example in the drawings and are herein described in detail. It is to be understood that the present design is not limited to the particular embodiments disclosed, but rather encompasses all modifications, equivalents, and alternatives using the specific embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0011] This disclosure describes the inventive concepts with reference to specific examples. However, the intent is to encompass modifications, equivalents, and alternatives of the inventive concepts consistent with this disclosure. However, it will be apparent to one skilled in the art that the present techniques may be practiced without these specific details. Accordingly, the specific details described are merely exemplary and are not intended to limit the scope of the present disclosure. Features implemented in one embodiment may be implemented in other embodiments where logically possible. Furthermore, specific numerical references may be made, such as to a first user. However, the specific numerical references should not be construed as a literal order; rather, a first user should be construed as being different from a second user. Features implemented in one embodiment may be implemented in other embodiments where logically possible. It is contemplated that variations from the specific details can be made and still be within the spirit and scope of what is disclosed.
[0012] 1 shows a block diagram of an exemplary conversational engagement platform embodiment that includes a dialogue management module that cooperates with one or more machine learning models that receive parameters tied to a user's speech input from other modules, where the machine learning models are trained to understand the user's communicated meaning, including potential paralinguistic effects, and to trigger changes in prosody to convey the paralinguistic effects when speaking to the user. Conversational engagement platform 100 can use a speech generation module, such as a text-to-speech module, as an additional source of information in addition to the pronounced i) phonemes, ii) words, and / or iii) non-words to create an audio or video file with acoustic cues through prosody to trigger two or more paralinguistic effects, each of which may change the pronunciation of i) phonemes, ii) words, iii) non-words, and iv) any combination thereof, in a pronunciation that differs from the normal pronunciation of the i) phonemes, ii) words, and / or iii) non-words. The speech generation module receives 1) textual representations, 2) acoustic representations, or 3) a combination of both, annotated with one or more markup markers (e.g., diacritics) on i) phonemes, ii) words, and / or iii) non-words. These received audio representations guide the speech generation module on how to pronounce i) phonemes, ii) words, and / or iii) non-words with specific paralinguistic effects through variations in prosody to cause pronunciations that differ from the normal pronunciation of that i) phoneme, ii) word, and / or iii) non-word for a given communication in order to convey the specific paralinguistic effects. The specific paralinguistic effects are conveyed in a manner that follows actual learned examples from people who convey the specific paralinguistic effects, bringing additional information sources into play when pronouncing the i) phonemes, ii) words, and / or iii) non-words.One or more machine learning models are trained to examine i) words, ii) phonemes, and / or iii) non-words annotated with one or more markup markers and then guide the generation of 1) a textual representation, 2) a sound wave representation, or 3) a combination of both, received by a speech generation module of how to pronounce the i) phonemes, ii) words, or iii) non-words with specific paralinguistic effects through changes in prosody to cause pronunciations that differ from normal pronunciations in a way that conveys additional information to a person in addition to the pronounced i) phonemes, ii) words, and / or iii) non-words themselves. The machine learning models can assist in the annotation of the correct markup markers to the i) phoneme, ii) word, and / or iii) non-word representations and / or can directly feed the sound wave files to the speech generation module. The speech generation module is configured to pronounce the i) phonemes, ii) words, and / or iii) non-words with specific paralinguistic effects through changes in prosody through the speaker.
[0013] It is noted that phonemes can be perceptually distinct units of sound within words and / or non-words in a given language. For example, English uses approximately 40+ phonemes. An example of a phoneme is the "c" in the word "car" because it has its own unique sound. English can include 19 vowels (5 short vowels, 6 long vowels, 3 diphthongs, 2 "oo" sounds, and 3 r-controlled vowels) and 25 consonant phonemes. Additionally, any software portions of the various modules and machine learning models discussed herein are stored in one or more non-transitory storage media in a format executable by one or more processors.
[0014] In embodiments, the speech generation module can receive the generated text representation to guide the speech generation module so that the speech generation module is configured to produce speech at different pronunciations with specific paralinguistic effects in a manner that better conveys 1) additional intended meaning and / or 2) enhanced understanding of the pronounced i) phonemes, ii) words, and / or iii) non-words themselves.
[0015] A conversational engagement platform 100, such as a dialogue system, can include various modules, such as a speech activity detector, a speech recognition module, a dialogue management module that cooperates with one or more machine learning models trained on paralinguistic effects conveyed by a person's speech and a set of markup markers correlated to each particular paralinguistic effect, a user state analysis module with input from a series of speech analysis detectors, such as a user emotional state module, a user sentiment analysis module, a lexical stress module, etc., an expression marker module, a timing module, a natural language generator module, a speech generation module, an information extraction and topic interpretation module, a voice assistant pipeline, and other modules. In the conversational engagement platform 100, the system can dynamically adapt the conversational engagement platform 100 between everyday conversation and directed dialogue based on the conversational context.
[0016] The conversational capabilities of the conversational engagement platform 100 are enhanced by one or more machine learning models trained to: i) leverage multimodal inputs such as sentiment, lexical stress, and emotion; ii) build details about user and activity details to leverage them in subsequent interactions; and iii) support other near-human-like interactions.
[0017] In an embodiment, one or more machine learning models are trained to examine audio data including at least one of i) words, ii) phonemes, and iii) non-words in a speech communication, which are annotated with one or more markup markers that guide the generation of a text representation to cause a pronunciation that differs from the plain pronunciation that would occur in the absence of the markup markers in order to convey at least one of 1) additional intended meaning and 2) enhanced understanding of the pronounced i) phonemes, ii) words, and / or iii) non-words themselves, and the one or more machine learning models are trained with training data of speakers using paralinguistic effects.
[0018] Again, a variety of acoustic cues (in addition to, and often superimposed on, the words themselves) can be added to a person's speech to convey or clarify meaning. Prosody can be variations in individual speech parameters (e.g., volume, pitch, duration / length) of tone, rhythm, and intonation (e.g., the rise and fall of the voice in speech), as well as their different patterns in speech. Through variations in prosody on words, nonwords, and even phonemes, these acoustic cues serve many purposes; for example, they serve the discourse function of controlling and managing a dialogue. For example, to communicate a speaker's desire to hold the conversational floor, a speaker can use a prosodic pattern of rising pitch followed by a pitch plateau on most words. Similarly, a speaker can use the prosody of lengthening the "m" sound in "Hmmm...?" to indicate uncertainty or question certainty about what the user just said. People often convey acoustic cues as additional sources of information in addition to the actual words, nonwords, and / or phonemes uttered. Additionally, verbal stress in words, nonwords, and phonemes, as well as the position of stress within other words and nonwords being communicated, can alter their meaning or clarify their significance, depending on where acoustic stress is placed. Thus, prosodic variations and patterns can reflect various characteristics of speech, such as the speaker's emotional state, speech form (statement, question / inquiry, or command), sentiment (e.g., the presence of sarcasm or sarcasm), contrast, and emphasis within speech. These acoustic variations in prosody convey additional information beyond the definitional meaning of the word or nonword itself. Thus, acoustic cues through prosodic variations in prosody are an additional source of information, often superimposed on top of the actual words, nonwords, and / or phonemes being spoken / enunciated. However, many conventional dialogue systems ignore the prosodic components being communicated and simply attempt to understand or produce the words and / or nonwords themselves.
[0019] Thus, many current text-to-speech synthesizers are not readily able to utilize and control these paralinguistic cues to convey this additional information in addition to the pronounced words, non-words, and phonemes themselves, or similarly, to understand this additional information as the user speaks using paralinguistic effects, for example, which places significant constraints on the naturalness and understandability of the dialogue system. The present design introduces mechanisms to better understand, utilize, and convey paralinguistic effects that convey additional information as words and / or non-words are uttered.
[0020] The expressive marker module can collaborate with one or more machine learning models trained on paralinguistic effects and a set of markup markers that associate each particular paralinguistic effect with one another. The expressive marker module can collaborate with one or more machine learning models to annotate markup markers on a representation of speech that associate paralinguistic effects with one another for discourse functions ("hold the conversation floor" or "elicit feedback from the other party"). When markup markers are annotated on a representation of speech, such as the text or waveforms of words, phonemes, and non-words, the acoustic signal cues synthesized by the speech generation module can result in controllable, natural-sounding paralanguage for text-to-speech synthesis.
[0021] The markup markers annotated in the speech representation are input to a speech generation module that allows the system developer to specify and control aspects of the output related to paralinguistic features such as stress, emphasis, and discourse functions. These paralinguistic effects can be achieved at lower acoustic levels by systems that manipulate speech parameters such as duration, loudness, pitch height, pitch contour, phonation type, vocal effort, etc.
[0022] A person can convey an intended paralinguistic effect by varying these 1) speech parameters relative to the standard (e.g., differing from the standard by, say, 35%) and 2) the pattern in which these speech parameters are varied.
[0023] Example System Processing The conversational engagement platform 100 may utilize paralinguistic effects to interact and converse with one or more people. The human user's voice may be transmitted through a microphone. Thus, the human's speech, utilizing 1) i) phonemes, ii) words, iii) non-words, and iv) any combination thereof pronounced with paralinguistic effects, and 2) normal / standard pronunciation of i) phonemes, ii) words, iii) non-words, and iv) any combination thereof, is transmitted to a voice activity detector. The microphone digitizes vibrations caused by sound by measuring sound wave frequencies. The system can then filter the digitized sound to remove undesirable noise outside the frequency range of speech. The microphone converts the vibrations and frequencies of the user's voice into digitized data (e.g., a sound wave file). The microphone output provides the digitized voice to the voice activity detector.
[0024] The voice activity detector detects when speech is occurring and sends the digitized audio files and time codes to i) an automatic speech recognition module, ii) a user state analysis module, iii) a time synchronization module, and iv) a dialogue manager module. It is noted that the one or more machine learning models can stand alone as distinct components in themselves, or can be incorporated into larger modules such as a dialogue manager module or a voice generation module.
[0025] The time code helps the time synchronization module synchronize the output data from these different processes back to its relative position in the voice stream for each i) phoneme, ii) word, iii) non-word, and iv) any combination thereof in the voice stream after processing from each module.
[0026] In parallel with other processes that receive the voice activity detector output, the automatic speech recognition module can produce 1) waveforms for individual words, phonemes, and non-words, 2) text representations of individual words, phonemes, and non-words, and 3) any combination of both. The automatic speech recognition module can take the digitized voice data activity input from the voice detector and produce individual sound waveforms for words, non-words, phonemes, and any combination thereof. Similarly to the produced text, the automatic speech recognition module can take the digitized voice data input and produce individual text representations for individual words, individual non-words, individual phonemes, and / or combinations thereof.
[0027] It is noted that one form of text representation that a developer can use is conventional orthography. It is noted that orthography can be a system of writing conventions used to represent spoken English in written form, allowing a reader to make the connection from spelling to sound to meaning. Another form of textual representation can be phonemic or sub-phonemic (phonetic feature based) representation, where length, loudness, etc. are implicitly or explicitly specified for each phoneme (phonetic segment). Another form of textual representation can be the phonetic sounds of i) phonemes, ii) words, or iii) non-words.
[0028] In parallel with other processes, the user state analysis module analyzes the speech data through several different speech processing detectors for the user's emotional state, mood, stressed lexical points, and other states conveyed in the input speech data. The user state analysis module analyzes each individual word, non-word, and phoneme for these as well as individual low-level speech parameters. The user state analysis module works in conjunction with several different automatic speech analysis detectors.
[0029] The set of speech processing detectors can include user emotional state modules such as SRI International's SenSay speech platform and SRI International's J-miner platform, a user sentiment analysis module, and a lexical stress module. These detectors and their algorithms detect a series of speech parameters / phenomena and acoustic realizations of speech phenomena. The speech processing detectors detect speech parameters, which are relevant building blocks that can be combined or filtered using linguistic knowledge to produce an approximation of the desired paralinguistic effect. The user state analysis module can identify individual words, nonwords, and phonemes that were not pronounced normally and therefore most likely had paralinguistic effects that convey additional information beyond the spoken / pronounced words, nonwords, and / or phonemes themselves. For each word, nonword, and phoneme that conveys a paralinguistic effect, the user state analysis module can output the detected user state, such as emotion, sentiment, lexical stress, and other paralinguistic effects for the individual words, nonwords, and phonemes identified by the detector as nonstandardly pronounced, as well as whether the detector happened to determine other types of paralinguistic intent. The output of the individual words, non-words, and phonemes identified by the detector as being nonstandardly pronounced, as well as what the paralinguistic effects were, is sent to the expression marker module. Supervised learning can use a linguistic expert to confirm the identified paralinguistic effects suggested by the detection tools in the user state analysis module. Furthermore, through the additional paralinguistic effects input module, the linguistic expert can add additional paralinguistic effects that the expert providing supervised machine learning has detected on individual phonemes, words, and non-words, such as prosodic changes to convey the user's intent to continue holding the conversational floor, and / or other paralinguistic effects that may be detected by the user and flagged by the expert.In a feedback loop, the linguistic expert providing the supervised machine learning can provide confirmation of the machine learning model predictions of the detected paralinguistic effects and the additional information / meaning they convey when used. It is noted that the ability of a set of speech analysis detectors to detect individual low-level speech parameters can be used as components in identifying when paralinguistic effects are bound to pronounced words, non-words, and / or phonemes.
[0030] The set of detectors analyzing speech data from a person communicating one or more paralinguistic effects may further include: 1) a pitch detector configured to use fundamental frequency to find i) the highest pitch on phonemes within words and sentences, ii) rising or falling pitch trends throughout a sentence, and iii) a combination of both; 2) a lexical stress detector configured to compare the duration of how long a phoneme is said to the normal length of time for uttering that phoneme (the standard taking into account the dialect being identified) and the loudness of the word compared to other words being communicated at the same time; 3) a stress detector configured to look at a combination of pitch, volume, and duration; and 4) a mood detector configured to detect the user's mood.
[0031] The detector can use lower level audio analysis tools to extract, for example, the following audio characteristics: i) Model lengthening by performing a forced alignment, the extraction of which determines which phonemes are long (e.g., above the 90th percentile) or short (e.g., below the 10th percentile) for their class, and annotating these with an exemplary colon (:) markup marker (a markup marker that indicates the lengthened version of this phoneme). ii) A lexical stress tool is used to detect the stress of the sentence and create example markup markers corresponding to the lexical stress. iii) Detect phrase-final prosody using the pitch trajectory from the f0 tracker. It is noted that the fundamental frequency F0 is generally associated with the lowest frequency in a harmonic progression (equal to the difference between adjacent harmonics).
[0032] The expressive marker module can annotate each phoneme, word, and / or non-word and its detected paralinguistic effect with a corresponding markup marker. The markup marker can be high-level discourse-related markup. The expressive marker module can reference a table of markup markers for their corresponding paralinguistic effect, their prosodic patterns, and variations that correlate to the particular paralinguistic effect. In this way, paralinguistic effects can be annotated in an audio corpus in an automated manner.
[0033] The set of detectors, working in conjunction with the expression markers and machine learning model, provides a mechanism for the machine learning model to learn to distinguish the paralinguistic meanings of non-word lexical units that are usually written identically (e.g., "uh-huh" indicating agreement, skepticism, annoyance, or an indication of the speaker's judgment about whether the other person has finished speaking) and to use them where appropriate in a dialogue.
[0034] Again, the timing module can help synchronize the individual waveform representations, the text representations of words, non-words, and phonemes from the speech recognition module, and the identified words, non-words, and phonemes with potential paralinguistic effects from the speech detector, so that these speech representations can be annotated in a fast and efficient manner. Additionally, these inputs help the machine learning model use neural networks to piece together many different forms of the input into predicted paralinguistic effects, intended additional information, and prosodic variations and patterns that correlate with the predicted paralinguistic effects.
[0035] The output from 1) the voice activity detector, digitized data of the audio stream with its time codes; 2) sound waveform representations of individual i) phonemes, ii) words, iii) non-words, and iv) any combination thereof, annotated with markup markers; and / or 3) text representations of individual i) phonemes, ii) words, iii) non-words, and iv) any combination thereof, annotated with markup markers; and 4) the timing module timeline can be fed to one or more machine learning models trained on paralinguistic effects. In addition, a bidirectional feedback loop exists between the expression marker module and the one or more machine learning models to identify paralinguistic effects, prosodic patterns, and variations associated with various paralinguistic effects, as well as corresponding markup markers. The machine learning-based TTS system trains on and learns the complex acoustic patterns associated with these paralinguistic effects, allowing them to be utilized by a dialogue system using a speech generation module.
[0036] It is noted that FIG. 2 provides a more detailed discussion of a training system 200 for automating the pre-deployment training of one machine learning model trained on paralinguistic effects conveyed by human speech, and a set of markup markers that associate each particular paralinguistic effect with each other.
[0037] In short, the machine learning model can learn both the prosodic patterns that people use to communicate specific paralinguistic effects and the markup markers that correspond to those effects. The mapping process leverages the capabilities of a set of speech detectors that automatically detect many low-level and high-level aspects of speech, which are mapped to each corresponding paralinguistic effect and then mapped to corresponding markup markers. The mapping process with a feedback loop leverages the capabilities of the machine learning model, for example, using a neural network (e.g., a deep neural network (DNN)), to implicitly model and learn patterns. The machine learning model can also learn additional information that is typically conveyed by paralinguistic effects in addition to the pronounced word, phoneme, and / or non-word itself. The mapping process can then also map to the corresponding intended additional information conveyed by the corresponding paralinguistic effect.
[0038] In an embodiment, the table may map a set of markup markers, each different markup marker may be mapped to its corresponding specific paralinguistic effect on pronounced i) phonemes, ii) words, and / or iii) non-words. The machine learning model works with the table to generate changes in the pronunciation of i) phonemes, ii) words, and / or iii) non-words when conveying a particular paralinguistic effect. The machine learning model is trained by supervised machine learning on how to identify paralinguistic effects and corresponding markup markers, and then understand the additional information conveyed by each particular paralinguistic effect.
[0039] Machine learning models can assist both 1) in creating paralinguistic effects in text-to-speech output and 2) in more accurately identifying what additional information is conveyed when a person conveys a particular paralinguistic effect beyond the spoken words and / or nonwords themselves. Machine learning models are trained to associate and understand changes in prosody, including prosodic patterns, as well as prosodic changes on i) individual phonemes, ii) individual words, and iii) individual nonwords, with specific paralinguistic effects detected by a set of speech processing detectors within training data of audio streams of speech, as well as changes in individual speech parameters, including volume, duration, pitch, and the rising and falling intonation of these speech parameters. A machine learning model that analyzes and trains on paralinguistic effects with their corresponding markup markers can learn many functions. The machine learning model can also learn how to identify paralinguistic effects and corresponding markup markers and then understand the additional information conveyed by a speaker using the paralinguistic effect.
[0040] Similarly, for speech generated by a dialogue system, a machine learning model can learn to identify paralinguistic effects and corresponding markup markers that the system may want to annotate on the text or waveform representation of words, phonemes, or non-words to convey additional information when the text-to-speech engine pronounces speech with paralinguistic effects.
[0041] The table can be configured to map a set of markup markers (e.g., simple diacritics), each different markup marker being mapped to its own corresponding specific paralinguistic effect on pronounced i) phonemes, ii) words, and / or iii) non-words, which may have simple meanings in discourse (such as the pronunciation of an example phoneme in a word or non-word indicating an intent to "cede the conversational floor") and may correspond to complex, non-local acoustic correlations. Each paralinguistic effect corresponds to its own change in prosody: 1) individual phonetic parameters for other communication conveyed with the pronounced i) phonemes, ii) words, and / or iii) non-words, and 2) prosodic patterns conveyed with the pronounced i) phonemes, ii) words, and / or iii) non-words. The table is initially created by linguistic experts who formulate paralinguistic effects for corresponding patterns of prosodic changes and / or individual changes in prosody. However, the tables can be updated and revised as the machine learning model learns prosodic pattern changes and other paralinguistic nuances, guided by supervised machine learning that confirms or corrects results from updates to the machine learning model.
[0042] The natural language generator module may be configured to generate representations of i) phonemes, ii) words, and / or iii) non-words. The table may be referenced by the natural language generator module to mark up the textual representation, waveform representation, or any of these representations with one or more markup markers (such as diacritics) of i) phonemes, ii) words, or iii) non-words to guide the speech generation module on how to pronounce the i) phonemes, ii) words, or iii) non-words using paralinguistic effects through prosodic changes to cause pronunciations that differ from the standard / normal pronunciation of that i) phoneme, ii) word, or iii) non-word. For example, a markup marker corresponding to the paralinguistic effect of lexical stress may cause the word having the markup marker to be the most prominent word to be emphasized in a sentence compared to other words conveyed substantially simultaneously.
[0043] In embodiments, the natural language generator module can generate text representations of i) phonemes, ii) words, and / or iii) non-words. A table of markup markers can be configured so that each markup marker corresponds to its own paralinguistic effect. The table can be referenced by the natural language generator module to mark up the text representation with one or more markup markers for i) phonemes, ii) words, or iii) non-words to guide the speech generation module on how to pronounce the i) phoneme, ii) word, or iii) non-word with paralinguistic effects to cause a pronunciation that differs from the plain pronunciation of that i) phoneme, ii) word, or iii) non-word by a threshold amount in one or more speech parameters.
[0044] The dialogue manager, which has already been trained on examples from people conveying a specific set of paralinguistic effects and their corresponding markup markers, provides the natural language generator with guidance on how to mark up i) phoneme, ii) word, or iii) non-word representations with markup markers indicating the paralinguistic effects associated with a given annotated phoneme, which in turn instructs the speech generation module on how to pronounce that phoneme with a specific paralinguistic effect. The machine learning model is trained to learn how to identify each different paralinguistic effect and its corresponding markup marker, as well as how to create waveform modifications when conveying the first and second paralinguistic effects. This gives the speech generation module better control over prosody in spoken sentences or in back-channel speech generated by the speech generation module in virtual digital assistants such as Alexa or Siri.
[0045] Thus, a dialogue manager that has already trained a machine learning model on examples from people conveying a particular set of paralinguistic effects and their corresponding markup markers can also provide guidance to the natural language generator on how to express and emphasize other words, non-words, and / or phonemes in the particular word, non-word, and / or phoneme intended to be conveyed, compared to the standard for how that particular word, non-word, and / or phoneme is pronounced, and together with the set of words, non-words, and / or phonemes conveyed while the speech generation module maintains the conversational floor (these other unmarked-up words, non-words, and / or phonemes in the set will be pronounced normally, without any change in volume, pitch, or duration). The natural language generator module then generates a text representation with one or more annotated markup markings on one or more specific words, non-words, and / or phonemes intended to be conveyed while the speech generation module is holding the conversation floor, and does so by referencing a table that maps given markup markers to corresponding paralinguistic effects. It then annotates each of the text representations (e.g., word-like symbols) of the words, phonemes, or non-words stressed with appropriate diacritics to trigger the corresponding paralinguistic effects, which are then provided as input to the TTS module. Optionally, a speech generation module, such as a text-to-speech module, can reference a machine learning model for how to create sound waves to be pronounced through a speaker by text-to-speech to convey a set of text containing one or more words, non-words, or phonemes pronounced with paralinguistic effects. The trained model examines each word, phoneme, or non-word pronounced differently from its standard with appropriate diacritics to trigger the corresponding paralinguistic effect.
[0046] The machine learning model and the natural language generator can also work with a set of waveforms in a reference channel, as opposed to a table. A set of two or more channels with phoneme waveforms can be used as a reference for both natural pronunciation and pronunciation with different paralinguistic effects. Thus, for example, two or more channels with phoneme waveforms can be referenced by the natural language generator module to mark up the phoneme waveforms and compare the differences between the two input channels with markup markers corresponding to the differential paralinguistic effects.
[0047] Paralinguistic effects can often be highly nuanced and complex, with different patterns of similar lower-level speech phenomena, and their position within a word or sentence can change the meaning of the paralinguistic effect being conveyed. Machine learning helps learn these nuances so that they can be recognized and distinguished from one another. Collaboration between natural language generators, speech generation modules, and machine learning models results in a way to define simple, conversationally relevant information with paralinguistic effects beyond phonemes, to have speech generation modules that are i) more understandable, and ii) can be used in more nuanced dialogue systems.
[0048] For example, some paralinguistic effects include speakers 1) stressing and / or emphasizing certain important words in a sentence by varying their pronunciation volume or duration, 2) using a rising pitch followed by a pitch flat to convey additional information in addition to the pronounced word where the speaker still wishes to hold the conversational floor, and 3) lengthening the "m" sound in "mmm..." and flattening it in pitch, which convey additional information in addition to the pronounced word the speaker is conveying.
[0049] Some additional paralinguistic effects include a speaker 4) communicating uncertainty about something by raising the pitch of a final phoneme to convey additional information beyond the spoken word whose meaning the speaker is questioning; 5) increasing the volume and shortening the duration of a given word to convey additional information beyond the spoken word whose meaning the speaker is angry; 6) decreasing the volume and lengthening the duration of a given word to convey additional information beyond the spoken word whose meaning the speaker is sad; 7) and many other changes in prosodic patterns and / or individual prosodic parameters.
[0050] With the help of machine learning models, text-to-speech synthesis becomes more human-like and natural in its ability to create and control paralinguistic cues. Again, paralinguistic effects can be components of metacommunication that can modify meaning to impart nuanced meaning, such as conveying surprise, anger, or questions; slowing speech to increase comprehension; creating longer pauses for long lists (commands, complex information, or emotional communication); etc. An important function of paralinguistic effects is to convey information i) more clearly than in the absence of paralinguistic phenomena and / or 2) more efficiently than with additional words alone.
[0051] These combinations result in better, more paralinguistically appropriate and informative text-to-speech output from the speech generation module, and a conversational engagement platform that can handle a fuller range of speech input from its users.
[0052] A machine learning model working in conjunction with the natural language generator identifies markup markers for intended paralinguistic functions to be inserted into intended sequences of words and non-words in a system using a speech generation module, based on determinations from one or more AI models that have learned where people tend to use paralinguistic cues.
[0053] As discussed, the natural language generator marks up the waveform representation, text representation, and any of these annotated with markup markers corresponding to particular paralinguistic effects that are desired to be conveyed along with the words, phonemes, and / or non-words in the waveform representation or text representation, and supplies this to the speech generation module.
[0054] When the dialogue system desires to generate speech with paralinguistic effects, the one or more machine learning models are configured to examine words, phonemes, and / or non-words annotated with one or more markup markers and provide to the speech generation module an acoustic file of i) phonemes, ii) words, or iii) non-words with paralinguistic effects corresponding to the one or more markup markers so as to appropriately pronounce the i) phonemes, ii) words, or iii) non-words with paralinguistic effects in a manner that conveys additional information to a person in addition to the i) phonemes, ii) words, or iii) non-words themselves that are pronounced through a speaker and through prosodic variations.
[0055] Thus, the speech generation module takes as input a marked-up, e.g., orthographic representation, and uses paralinguistic effects indicated by markup markers annotated on the i) phonemes, ii) words, and / or iii) non-words to pronounce i) phonemes, ii) words, iii) non-words, and iv) any combination thereof. It is noted that how the markup and its prosodic variations and / or prosodic patterns are realized will depend on the spoken language, such as English, and may have subsequent distinctions for dialects within a language, such as the New York accent. One or more models can work together to bring this functionality to multiple spoken languages and / or different dialects within a spoken language.
[0056] Similarly, the speech generation module can reference one or more machine learning models for how to pronounce each given phoneme annotated with its markup markings for different domains, such as domains for elderly spoken, medical spoken, etc.
[0057] Machine learning models working in conjunction with text-to-speech and natural language generator modules can create many kinds of prosodic features that are used all the time in human-to-human interactions, such as to manage discourse, manage turn-taking, control the rate of information flow, and confirm or indicate lack of understanding.
[0058] Using the paralinguistic capabilities provided by waveforms and tables provided by machine learning, text-to-speech serves a variety of discourse functions and gives system developers control over the prosodic variations humans typically use. The paralinguistic layer added to text-to-speech allows for naturalness and potentially greater intelligibility in the output of text-to-speech systems, for example in the context of dialogue systems, and allows system developers to use more verbose strategies to convey information. Users of spoken-language systems can become uncertain about when the system has stopped speaking and / or what the system's state of processing or knowledge is due to the use of monotonous words alone, or worse, the arbitrary use of prosodic patterns that do not correspond to meaning. However, when the same words are explicitly communicated to the user through acoustic cues (e.g., by spelling out the state in words and without visually indicating the recipient with spelled words), the user's comprehension increases more than simply the words alone.
[0059] The potential payoff is that text-to-speech output could sound much more human and natural than today's text-to-speech systems typically do, at a much lower cost than it would cost to produce similarly natural output using today's text-to-speech systems. In addition, there could be real efficiencies to be gained in how users interact with spoken systems, since they could rely on information conveyed through the prosodic channel (simultaneously with speech) without having to explicitly detail the system's state to them in words or other means. This could lead to much greater user willingness to use the spoken system, improved user satisfaction, and a wider range of tasks for which users perceive the spoken system to perform satisfactorily.
[0060] Paralinguistic effects are a part of everyday human conversation, so users should naturally incorporate what comes from a dialogue system using this design. Users will find dialogue systems with this design easier to use and allow information to be exchanged more efficiently.
[0061] The text-to-speech output may sound much more human and natural than text-to-speech output typically sounds today. The text-to-speech output may be able to communicate more clearly and / or more efficiently.
[0062] FIG. 2 shows a block diagram of an example training system for one or more machine learning models trained on paralinguistic effects conveyed by human speech and a set of markup markers that correlate to each particular paralinguistic effect.
[0063] An exemplary training system 200 for one or more machine learning models may include various modules: a voice activity detector; a speech recognition module; one or more machine learning models trained on paralinguistic effects conveyed by human speech and a set of markup markers correlated to each particular paralinguistic effect; a user state analysis module with input from a set of speech analysis tools, such as a user emotional state module, a user sentiment analysis module, a lexical stress module, etc.; an expression marker module; a timing module; and potentially a natural language generator module; and a text-to-speech module.
[0064] Human speech utilizing both 1) i) phonemes, ii) words, iii) non-words, and iv) any combination thereof pronounced with paralinguistic effects, and 2) normal / standard pronunciation of i) phonemes, ii) words, iii) non-words, and iv) any combination thereof, is conveyed to a voice activity detector. In an embodiment, the human speech may be conveyed, for example, through a microphone. The microphone digitizes vibrations caused by sound by measuring sound wave frequencies. The system may then filter the digitized sound to remove undesirable noise outside the speech frequency range. The microphone converts the vibrations and frequencies of the user's voice into data (such as a sound wave file). In an embodiment, a memory storing the sound wave file may provide a digitized version of the speech to the voice activity detector. In an embodiment, the combination of the memory and microphone results in a digitized version of the input speech data.
[0065] The voice activity detector module, the speech recognition module, the user state analysis module with inputs from a series of speech analysis tools, such as a user emotional state module, a user sentiment analysis module, a lexical stress module, etc., and the timing module all function in a similar manner as discussed in FIG. 1.
[0066] It is noted that speech processing detectors yield low-level phonetic features and, in some cases, the proposed explicit paralinguistic effects. For example, speech processing tools such as DynaSpeak have pitch trackers and can detect lexical stress.
[0067] The speech data serves as training data for a machine learning model with a neural network. The input training data has known labels that will be associated with the data by an expression marker module. It is noted that the input training data can be a mixture of labeled and unlabeled training data. The AI model undergoes a supervised learning training process, in which the model and its neural network are required to make predictions and are corrected when those predictions are incorrect. The model must learn the paralinguistic effect structure / corresponding prosodic changes to organize the data and make predictions. The prosody of low-level speech parameters and changes in the high-level prosodic patterns of speech are mapped to each corresponding paralinguistic effect, which is mapped to corresponding markup markers, which are mapped to corresponding additional information conveyed by the corresponding paralinguistic effect. The training process continues until the model achieves a desired level of accuracy on the training data, such as at least 90% accuracy.
[0068] In an embodiment, a set of speech processing detectors can analyze training data of an audio stream of speech from a speaker. The set of speech processing detectors can detect speech parameters indicative of one or more paralinguistic effects in addition to pronounced words, phonemes, and non-words in the audio stream. A speech recognition module can receive the training data of the audio stream of speech and create text representations for individual words, individual non-words, individual phonemes, and any combination thereof in the training data of the audio stream of speech. One or more machine learning models using neural networks undergo supervised machine learning to train how to associate one or more markup markers with a text representation for each individual word, individual non-word, and / or individual phoneme pronounced with a particular paralinguistic effect. The set of speech processing detectors, the speech recognition module, and the one or more machine learning models can cooperate to automate the labeling of training data and the training of machine learning models trained for paralinguistic effects prior to deployment. Each markup marker can correspond to its own paralinguistic effect.
[0069] The machine learning model can learn how to identify paralinguistic effects and corresponding markup markers when conveying a particular paralinguistic effect, and how to work with the table to learn how to create changes in the pronunciation of i) phonemes, ii) words, and / or iii) non-words, and to create and / or update a table of mapping markup markers and corresponding paralinguistic effects to text or waveform representations of words, non-words, and / or phonemes. Supervised learning for machine learning uses linguistic experts to confirm the identified paralinguistic effects suggested by the detection tool in the user state analysis module. The detection tool itself can directly detect some paralinguistic effects. However, the detection tool itself detects many low-level aspects of speech—such as loudness and pronunciation duration—and these low-level speech aspects can be used to identify the presence of pronounced paralinguistic effects in addition to the words, non-words, and / or phonemes themselves. Additionally, through the Additional Paralinguistic Effects module, the linguistic expert can add additional paralinguistic effects that the expert detects on individual phonemes, words, and non-words, such as changes in prosody to convey the user's intent to continue holding the conversational floor, and / or other paralinguistic effects that can be detected by the user and flagged by the expert. Many explicit paralinguistic effects often correspond to combinations of lower-level detectors.
[0070] Prior to deployment during training, machine learning models, each with its neural network, are trained to look for patterns, whereby text and / or waveform representations are analyzed by a set of detectors, and when words, non-words, and / or phonemes deviate sufficiently from the standard (e.g., more than 35% from the standard pronunciation in the reference dialogue model), the output from the set of speech processing detectors conveys both 1) different speech parameters in any of duration, volume, and pitch that are increased or decreased from the standard, and 2) prosodic patterns of change in the different speech parameters, which are analyzed by one or more machine learning models and then correlated to specific paralinguistic effects and to the intended additional information to be conveyed by that specific paralinguistic effect (e.g., a rising pitch followed by a pitch plateau), which is then analyzed by one or more machine learning models. The one or more machine learning models can then correlate with additional information intended to be conveyed by paralinguistic effects, such as emphasizing the importance of the text and / or waveform expression, questioning the meaning of the text and / or waveform expression, emotions of surprise, anger or happiness, boredom, etc., all assisted by supervised machine learning that confirms or corrects the results determined by the machine learning models.
[0071] In an embodiment, a machine learning model is trained to look for patterns. A set of detectors helps the machine learning model 1) analyze and annotate the training data so that it can correlate specific paralinguistic effects with at least one of: i) their intended additional intended meaning; and ii) a more understandable manner in which the specific paralinguistic effect conveys information. The machine learning model is trained to learn how to identify each different paralinguistic effect and its corresponding markup marker, and then how to create corresponding waveforms when conveying the first and second paralinguistic effects.
[0072] As shown in FIG. 1, a set of speech processing detectors is configured to analyze training data of an audio stream of speech from a speaking person. The set of speech processing detectors detects pronounced words, phonemes, and non-words in the audio stream, as well as speech parameters indicative of one or more paralinguistic effects. The speech processing detectors can detect low-level speech phenomena naturally produced by a speaker and find a location for the phenomenon. The location can be a clearly defined segment, an indefinite segment, or something like a pitch peak that is a point in time but has branches to the pitch contour of indefinite extent surrounding the peak.
[0073] In an embodiment, a machine learning model is trained on variations in prosody, including prosodic patterns, and variations in individual speech parameters to relate and understand i) individual phoneme, ii) individual word, and iii) individual non-word variations to specific paralinguistic effects detected by a set of speech processing detectors within training data of an audio stream of speech. The speech parameters may be increased or decreased from the norm in any of the following: duration, volume, and pitch of variations in different speech parameters; and 2) prosodic patterns.
[0074] Again, in parallel with the speech processing detector that identifies low-level speech parameters, the automatic speech recognition module can create either or both waveform representations of individual words, phonemes, and non-words and text representations of individual words, phonemes, and non-words. The speech recognition module receives training data of an audio stream of speech and, for each word, each non-word, each phoneme, and any combination thereof, creates 1) a text representation, 2) a waveform representation, and 3) any combination of both in the training data of the audio stream of speech.
[0075] Again, the timing module can synchronize the identified words, non-words, and phonemes with their individual waveforms, text representations, and paralinguistic effects from the input speech. These waveform representations of individual words, phonemes, and non-words and text representations of individual words, phonemes, and non-words can be annotated with the paralinguistic effects by the expression marker module.
[0076] The expressive marker module works in conjunction with the machine learning model to provide a way to display different paralinguistic meanings on the input text and / or waveform that corresponds to the received digitized data of the input audio stream. After training, this enables the conversational engagement platform to both understand and generate paralinguistic effects to take advantage of the full range of expressive power of these words and non-words coming from and then going out to human users.
[0077] In an embodiment, the representational marker module is configured to annotate the textual representation of individual words, phonemes, and / or non-words with first markup markers corresponding to first paralinguistic effects, the first markup markers indicating particular paralinguistic effects.
[0078] The expressive marker module will annotate each phoneme, word, and / or non-word and its detected paralinguistic effects with a corresponding markup marker. The markup marker can be high-level discourse-related markup. In this way, paralinguistic effects can be annotated in the audio corpus.
[0079] The expressive marker module and / or additional paralinguistic effects input module annotate, in an automated manner, a training corpus of speech data with a set of markup markers for corresponding explicit paralinguistic effects. As a simple example, an automatic speech recognition process with forced alignment of known text can be used to discover distributions of duration, loudness, pitch height, etc. for phonemes or other units, such as syllables, and then use the tails of the distribution to yield simple annotations of extreme cases. When detection is sufficiently accurate and numerous, at runtime, the machine learning model, in conjunction with the text-to-speech module, can replicate meaningful features of the acoustic realization corresponding to the desired markup.
[0080] The expression marker module and the additional paralinguistic effect input module cooperate to identify and annotate presentations that have explicit paralinguistic effects themselves.
[0081] The expression marker module can use tables that are initially pre-filled with several different kinds of markup markers (e.g., adding an asterisk (*) diacritic to a phoneme in addition to other types of markup marker diacritics such as accents or cedillas), which gives the user more control over the prosody or expressiveness of the words, phonemes, and / or non-words that make up the utterance. For example, an asterisk (*) markup marker can signal to lengthen the pronunciation of a word from its standard, for example, by adding a colon markup marker to signal to lengthen the pronunciation of the word, non-word, or phoneme.
[0082] Part of the expression marker module can be a script that takes the output of the speech recognition module and the output of the detector of low-level speech features and any proposed detected paralinguistic effects from the detector to determine which phonemes, words, and / or non-words in the training set are, for example, long, and then tags the identified phonemes, words, and / or non-words with a corresponding markup marker, for example, a colon (:).
[0083] Thus, in the example, words may be emphasized relative to other words, non-words, and phonemes being conveyed by adding lengthening markers after those phonemes. The system uses the example length markup marker ">" after each phoneme in a word that is desired to be emphasized. In the example, complex information conveyed by a word may be better conveyed by adding long pause markers after those phonemes and by adding prosodic slowdown markers to pronounce that phoneme, word, or non-word in a slower manner than normal. In other examples, long pause markup markers may be annotated onto the expression. It is noted that there is a difference in pause duration between a regular comma in an uttered sentence and the paralinguistic effect of adding a long pause marker such as ">" to slow down the conveyance of complex information so that the user can more easily understand the conveyance of complex information by the speech generation module.
[0084] The collaboration between the set of detectors, the expressive marker module, and the machine learning model provides the ability to produce paralinguistically significant annotations that would be too difficult to achieve without automatic annotation using speech recognition and / or speech processing tools. These are explicit input markings for and control of prosody and other paralinguistic phenomena through the specification of potentially high-level, discourse-related events or dialogue states. The collaboration between the set of detectors, the expressive marker module, and the machine learning model provides the ability to utilize paralinguistic information learned automatically from a training corpus, including learning the variability of these features that occurs in human speech in a given language.
[0085] The one or more machine learning models undergo supervised machine learning on their neural networks to train at least how to associate one or more markup markers with 1) a text representation, 2) a waveform representation, and 3) any combination of both, for each individual word, individual non-word, individual phoneme, and any combination thereof pronounced with a particular paralinguistic effect, where each markup marker corresponds to its own paralinguistic effect.
[0086] The collaboration between the set of detectors, the expression marker module, and the machine learning model results in simple diacritics as markup markers that correspond to simple discourse meanings and complex acoustic realizations of paralinguistic effects. It is noted that diacritics, such as accents or cedillas, can be symbols that indicate phonetic differences from the same character when written above or below a word or in other forms, unmarked or differently marked. The collaboration between the speech recognition module, the set of detectors, the expression marker module, and the machine learning model results in automatic annotation of the training data (with the markup markers just mentioned).
[0087] The set of detectors, the expression marker module, and the machine learning model work together to produce training data and a set of markup marker annotations to produce a set of prosodic annotations that are usefully conveyed through the prosodic channel, resulting in a sound (not just robotic with default phoneme durations and pitch contours, but with expected human-like prosodic durations / contours, pitch variations, and appropriate prosodic patterns that lead to correct recognition by the end user of the communicative intent of the conversational engagement platform). The automatic labeling of paralinguistic effects on the training data in each domain by the set of detectors, the expression marker module, and the machine learning model work together to produce training data and a set of markup marker annotations that are large enough for prosodic features to be learned very well, and thereby generalizable to paralinguistic effects that are transferable to untrained and previously untrained domains with very little additional training.
[0088] Again, the supervised machine learning can, through the linguistic expert, confirm the identified paralinguistic effects suggested by any of the detectors in the user state analysis module. In addition, the supervised machine learning can, through the linguistic expert, confirm the identified paralinguistic effects annotated by the expression marker module. The role of the linguistic expert is to confirm things so that the supervised machine learning can learn. Additionally, through the additional paralinguistic effects input module, the linguistic expert can add additional paralinguistic effects that the expert has detected on individual phonemes, words, and non-words, such as prosodic changes to convey the user's intent to continue holding the conversation floor, and / or other paralinguistic effects detected by the user that may be flagged by the expert.
[0089] The output from 1) the digitized data of the audio stream of speech with its time codes from the voice activity detector, 2) the sound waveform representations of individual i) phonemes, ii) words, iii) non-words, and iv) any combination thereof annotated with markup markers, and / or 3) the text representations of individual i) phonemes, ii) words, iii) non-words, and iv) any combination thereof annotated with markup markers from the representation marker module, and 4) the timeline from the timing module may be fed to one or more machine learning models that are trained on paralinguistic effects and their corresponding markup markers. The use of a machine learning-based TTS system can learn the complex acoustic patterns associated with these paralinguistic effects and allow them to be utilized by the conversational engagement platform or even simply the training system 200 for learning paralinguistic effects in speech.
[0090] Aspects of many low-level and high-level speech parameters from a set of speech detectors can be correlated to associated changes in prosody and prosodic patterns, and then to each corresponding paralinguistic effect, and can be mapped to corresponding markup markers. The mapping process from the machine learning model, with its feedback loop through the table and expression marker module, exploits the ability of the machine learning model, e.g., using deep neural networks, to implicitly model, learn patterns, and then predict outcomes. The machine learning model can also learn additional information typically conveyed by paralinguistic effects in addition to the pronounced words, phonemes, and / or non-words themselves.
[0091] The machine learning model can learn both the prosodic patterns that humans use to communicate specific paralinguistic effects and the markup markers that correspond to those effects. After repeated supervised learning training, the neural network (e.g., DNN) eventually learns the paralinguistic effects it has been trained on and can be adapted to learn additional paralinguistic effects and translate its learning to different dialects and spoken languages. The mapping process leverages the DNN's ability to automatically detect many low-level (e.g., pitch, loudness, and phoneme duration level details) and high-level aspects of speech, as well as its ability to implicitly model patterns and arrive at predictable outcomes / decisions.
[0092] The waveform representation, the text representation, markup markers from the expression marker module about what the paralinguistic effect was, and the prosodic patterns that cause the paralinguistic effect, as well as the input from the timing module, can all be fed into the machine learning module, along with a copy of the original input speech from the audio stream. The machine learning module can be trained for paralinguistic effects on a domain-by-domain basis. For each machine learning model, the system can initially be trained within a narrow, specific domain so that it understands the normal pronunciation and prosody of words, nonwords, and phonemes within that domain and can mimic the prosodic patterns used by humans within that domain when producing text-to-speech output from the speech generation module. It is noted that after initially training within a specific domain, the trained model can generalize to other similar domains and tasks not previously trained on, as well as to additional prosodic patterns and variations found within the training data.
[0093] Again, the supervised machine learning, through the linguistic expert, can confirm the identified paralinguistic effects suggested by any of the detection tools in the user state analysis module. In addition, the supervised machine learning, through the linguistic expert, can confirm the identified paralinguistic effects annotated by the expression marker module. In addition, through the additional paralinguistic effects input module, the linguistic expert can add additional paralinguistic effects that the expert detects on individual phonemes, words, and non-words, such as prosodic changes to convey the user's intent to continue holding the conversational floor, and / or other paralinguistic effects detected by the user and that may be flagged by the expert.
[0094] Again, a feedback loop may be implemented between the expression marker module and the machine learning model to potentially fine-tune, train, and improve the automatic annotation of markup markers to text and / or acoustic expression.
[0095] For example, when a sufficient amount of training data is used via supervised machine learning that is sufficiently accurate and consistent to achieve the goal of identifying a set of paralinguistic effects with 95% or better accuracy, the neural network learns both how to recognize these effects in human speech input and how to accurately annotate expressions. Additionally, after a sufficient amount of training data to identify a set of paralinguistic effects with 95% or better accuracy (note that 95% is an example number, and accuracy can be 91% to 99.9%), the machine learning model can also assist the output of the speech generation module in creating these paralinguistic effect phenomena at runtime. The machine learning model is configured to create a paralinguistic effect waveform modification table that maps each annotated markup marker on the text generated by the natural language generator to how the speech generation module should modify the pronunciation of i) phonemes, ii) words, and / or iii) non-words with specific paralinguistic effects through changes in prosody.
[0096] Prior to runtime, the machine learning model can learn changes to speech parameters corresponding to paralinguistic effects and markup markers corresponding to particular paralinguistic effects. The machine learning model is trained by supervised machine learning on how to identify paralinguistic effects and corresponding markup markers and then understand the additional information conveyed by each particular paralinguistic effect.
[0097] Thus, prior to deployment in the field, the machine learning model is further configured to work with a set of detectors that analyze examples from people speaking a particular paralinguistic effect, such that based on receiving 1) textual representations, 2) acoustic representations, or 3) a combination of both annotated with markup markers, the speech generation module undergoes supervised machine learning to train how to associate (pair) one or more markup markers with textual and / or acoustic representations of i) phonemes, ii) words, or iii) non-words to speak the particular paralinguistic effect.
[0098] The machine learning module can also learn through supervised machine learning to output waveforms with paralinguistic effects to be sent to the text-to-speech module and to annotate the text representations, waveform representations, or a combination of both generated by the natural language generator with a table of expressive markup markers and corresponding paralinguistic effects.
[0099] The machine learning module can update and / or verify a table of expressive markup markers and the corresponding paralinguistic effects that correspond to the markup markers. The lookup table becomes updated according to the machine learning model's understanding of paralinguistic effects and their associated prosodic patterns and changes in prosody, and the additional information they convey. Supervised learning can confirm or correct the updates.
[0100] At runtime, the machine learning module can work directly with the speech generation module to output waveforms with paralinguistic effects and a table of expressive markup markers and their corresponding paralinguistic effects. The machine learning model can examine and train on patterns and prosodic changes, including volume, duration, and pitch in rising and falling intonation, to associate and understand individual phoneme, word, and non-word variations and prosody with paralinguistic effects. The machine learning model that analyzes and trains on paralinguistic effects with corresponding markup markers can learn many functions.
[0101] As discussed, the machine learning model can also use supervised machine learning for initial training on how to create waveforms to guide the speech generation module on how to pronounce words, non-words, and / or phonemes differently from their standard pronunciation using given paralinguistic effects.
[0102] The set of detectors, the expression marker module, and the machine learning model work together to reduce the cost and time to develop a system with powerful, discourse-appropriate acoustic characteristics.
[0103] network 3 shows a block diagram of several electronic systems and devices communicating with each other in a network environment according to an embodiment of the present design. Components within the conversation engagement platform may communicate within a network environment.
[0104] The network environment has a communications network 320 connecting server computing systems 304A-304B and at least one or more client computing systems 302A-302G. As shown, there may be many server computing systems 304A-304B and many client computing systems 302A-302G connected to each other through network 320, which may be, for example, the Internet. It is noted that network 320 may alternatively be or include one or more of an optical network, a cellular network, the Internet, a local area network (LAN), a wide area network (WAN), a satellite link, a fiber network, a cable network, or any combination thereof. Each server computing system 304A-304B may have circuitry and software for communicating with the other server computing systems 304A-304B and client computing systems 302A-302G over network 320. Each server computing system 304A-304B may be associated with one or more databases 306A-306B. Each server 304A-304B may have one or more instances of virtual servers running on its physical server, and multiple virtual instances may be implemented by the present design. To protect data integrity on client computing system 302D, a firewall may be established between the client computing system, e.g., 302D, and network 320.
[0105] Cloud provider services allow application software to be installed and run within the cloud, and users can access the software services from client devices. Cloud users with sites within the cloud do not have sole management of the cloud infrastructure and platform on which their applications run. Thus, servers and databases can be shared hardware, with users being given dedicated use of a certain amount of these resources. A user's cloud-based site is given a virtual amount of dedicated space and bandwidth within the cloud. Cloud applications can differ from other applications in their scalability, which can be achieved by cloning tasks to multiple virtual machines at runtime to adapt to changing workloads. A load balancer distributes work across a set of virtual machines. This process is transparent to cloud users, who only see a single point of access.
[0106] The cloud-based remote access is coded to utilize protocols such as Hypertext Transfer Protocol (HTTP) to engage in request and response cycles with both mobile device applications resident on the client devices 302A-302G and web browser applications resident on the client devices 302A-302G. In some circumstances, the cloud-based remote access to the wearable electronic device 302C may be accessed through a mobile, desktop, or tablet device cooperating with the wearable electronic device 302C. The cloud-based remote access between the client devices 302A-302G and the cloud-based provider site 304A is coded to engage in one or more of: 1) a request and response cycle from any web browser-based application; 2) an SMS / Twitter-based request and response message exchange; 3) a request and response cycle from a dedicated online server; 4) a direct request and response cycle between the native mobile application resident on the client device and the cloud-based remote access to the wearable electronic device; and 5) combinations thereof.
[0107] In an embodiment, the server computing system 304A may include a server engine, a web page management component or an online service or application component, a content management component, and a database management component. The server engine performs basic processing and operating system level tasks. The web page management component, online services, or online application component handles the creation and display and / or routing of web pages or screens associated with receiving and serving digital content and digital advertisements. Users may access the server computing device using its associated URL. The content management component handles most of the functionality in the embodiments described herein. The database management component includes storage and retrieval tasks with respect to databases, queries against databases, and storage of data.
[0108] Computing Devices 4 shows a block diagram of an embodiment of one or more computing devices that may be part of a conversational engagement platform for embodiments of the present design discussed herein. Components of the conversational engagement platform may implement aspects of the computing devices as follows. For example, a detector in a set of detectors may have such an architecture.
[0109] The computing device may include one or more processors or processing units 420 for executing instructions, one or more memories 430-432 for storing information, one or more data input components 460-463 for receiving data input from a user of the computing device 400, one or more modules including a management module, a network interface communication circuitry 470 for establishing communication links for communicating with other computing devices external to the computing device, one or more sensors whose output is used to detect certain triggering conditions and then generate one or more pre-programmed actions in response thereto, a display screen 491 for displaying at least some of the information stored in the one or more memories 430-432, and other components. It is noted that portions of the present design embodied in software 444, 445, 446 can be stored in one or more memories 430-432 and executed by one or more processors 420. The processing unit 420 may have one or more processing cores coupled to a system bus 421, which couples various system components including the system memory 430. The system bus 421 may be any of several types of bus structures selected from a memory bus, an interconnect fabric, a peripheral bus, and a local bus using any of a variety of bus architectures.
[0110] The computing device 402 typically includes a variety of computer-readable media. Machine-readable media can be any available medium that can be accessed by the computing device 402, including both volatile and nonvolatile media, and removable and non-removable media. By way of example, the use of computer-readable media includes, but is not limited to, the storage of information such as computer-readable instructions, data structures, other executable software, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other tangible medium that can be used to store desired information and that can be accessed by the computing device 402. Transient media, such as wireless channels, are not included as machine-readable media. Machine-readable media typically embodies computer-readable instructions, data structures, and other executable software.
[0111] In the example, volatile memory drive 441 is shown for storing operating system 444 , application programs 445 , other executable software 446 , and program data 447 .
[0112] A user may enter commands and information into the computing device 402 through input devices such as a keyboard, touch screen, or software or hardware input buttons 462, a microphone 463, a pointing device and / or scrolling input component such as a mouse, trackball, or touchpad 461. The microphone 463 may work with voice recognition software. These and other input devices are often connected to the processing unit 420 through a user input interface 460 coupled to the system bus 421, but may also be connected by other interface and bus structures, such as a light port, a game port, or a universal serial bus (USB). A display monitor 491 or other type of display screen device is also connected to the system bus 421 through an interface such as a display interface 490. In addition to the monitor 491, computing devices may also include other peripheral output devices, such as speakers 497, a vibrating device 499, and other output devices, which may be connected through an output peripheral interface 495.
[0113] The computing device 402 can operate in a networked environment using logical connections to one or more remote computers / client devices, such as a remote computing system 480. The remote computing system 480 can be a personal computer, a mobile computing device, a server, a router, a network PC, a peer device or other common network node, and typically includes many or all of the elements described above with respect to the computing device 402. The logical connections can include a personal area network (PAN) 472 (e.g., Bluetooth), a local area network (LAN) 471 (e.g., Wi-Fi), and a wide area network (WAN) 473 (e.g., a cellular network). Such networking environments are commonplace in offices, enterprise computer networks, intranets, and the Internet. A browser application and / or one or more local applications can be resident on the computing device and stored in memory.
[0114] When used in a LAN networking environment, computing device 402 is connected to LAN 471 through network interface 470, which may be, for example, a Bluetooth® or Wi-Fi adapter. When used in a WAN networking environment (e.g., the Internet), computing device 402 typically includes some means for establishing communications over WAN 473. For mobile communications technologies, for example, a wireless interface, which may be internal or external, may be connected to system bus 421 through network interface 470 or other appropriate mechanism. In a networked environment, other software depicted relative to computing device 402, or portions thereof, may be stored in remote memory storage devices. By way of example, and not limitation, remote application programs 485 may reside on remote computing device 480. It will be understood that the network connections shown are examples, and that other means of establishing a communications link between computing devices may be used.
[0115] It should be noted that the design may be performed on a computing device such as that described in connection with the figures, but the design may also be performed on a server, a computing device dedicated to message handling, or on a distributed system in which different parts of the design are performed on different parts of the distributed computing system.
[0116] It is noted that applications discussed herein include, but are not limited to, software applications, mobile applications, and programs that are part of operating system applications. Some portions of this description are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. These algorithms may be written in a number of different software programming languages, such as C, C++, HTTP, Java, or other similar languages. Also, an algorithm may be implemented by lines of code in software, configured logic gates in hardware, or a combination of both. In embodiments, logic consists of electronic circuits that follow the rules of Boolean logic, software containing patterns of instructions, or any combination of both. Modules may be implemented in hardware electronic components, software components, or a combination of both.
[0117] Generally, applications include programs, routines, objects, widgets, plug-ins, and other similar structures that perform particular tasks or implement particular abstract data types. Those skilled in the art can implement the descriptions and / or figures herein as computer-executable instructions, which may be embodied on any form of computer-readable medium discussed herein.
[0118] Many functions performed by electronic hardware components can be replicated through software emulation. Thus, software programs written to achieve those same functions can emulate the functions of the hardware components in input-output circuits.
[0119] While the above designs and embodiments have been described in considerable detail, it is not the applicant's intention that the designs and embodiments described herein be limiting. Additional adaptations and / or modifications are possible, and the broader aspects encompass these adaptations and / or modifications. Accordingly, departures may be made from the above designs and embodiments without departing from the scope imposed by the appended claims, which, when properly interpreted, are limited only by the claims.
Claims
1. one or more machine learning models trained to: examine audio data including at least one of i) words, ii) phonemes, and iii) non-words in a spoken communication annotated with one or more markup markers; and guide the generation of a text representation to cause pronunciations of the pronounced i) phonemes, ii) words, and / or iii) non-words themselves that differ from plain pronunciations that would occur in the absence of the markup markers in order to convey at least one of 1) additional intended meaning and 2) enhanced understanding, wherein the one or more machine learning models have been trained with training data of speakers using paralinguistic effects; a speech generation module configured to receive the generated text representation to guide the speech generation module, thereby producing speech with the different pronunciation and the specific paralinguistic effect in a manner that better conveys 1) the additional intended meaning and / or 2) the enhanced understanding of the pronounced i) phonemes, ii) words, and / or iii) non-words themselves; The apparatus, wherein the speech generation module and any software portions of the one or more machine learning models are stored on one or more non-transitory storage media in a state executable by one or more processors.
2. a table configured to map a set of markup markers, each different markup marker being mapped to its corresponding specific paralinguistic effect on pronounced i) phonemes, ii) words, and / or iii) non-words; the machine learning model is configured to cooperate with the table to generate variations in the pronunciation of the i) phonemes, ii) words, and / or iii) non-words when conveying a particular paralinguistic effect; 10. The apparatus of claim 1.
3. a natural language generator module configured to generate textual representations of the i) phonemes, ii) words, and / or iii) non-words; a table of markup markers, each markup marker corresponding to its own paralinguistic effect; marking up the text representation with one or more markup markers of the i) phoneme, ii) word, or iii) non-word, such that the table is referenced by the natural language generator module to guide the speech generation module on how to pronounce the i) phoneme, ii) word, or iii) non-word using the paralinguistic effect to cause the pronunciation to differ from the plain pronunciation of that i) phoneme, ii) word, or iii) non-word; 10. The apparatus of claim 1.
4. The apparatus of claim 3 , wherein the different pronunciation differs in one or more phonetic parameters from the plain pronunciation of that i) phoneme, ii) word, or iii) non-word by a threshold amount.
5. 10. The apparatus of claim 1, wherein the machine learning model is trained by supervised machine learning on how to identify the paralinguistic effects and corresponding markup markers and understand the additional information conveyed by each particular paralinguistic effect.
6. a set of speech processing detectors configured to analyze training data of an audio stream of speech from a speaker, the set of speech processing detectors detecting pronounced words, phonemes, and non-words in the audio stream as well as speech parameters indicative of one or more paralinguistic effects; a speech recognition module configured to receive the training data of the audio stream of speech and to produce a text representation for each word, each non-word, each phoneme, and any combination thereof, in the training data of the audio stream of speech; one or more machine learning models using neural networks configured to undergo supervised machine learning to train how to associate one or more markup markers with the text representation of each individual word, individual non-word, and / or individual phoneme pronounced with a particular paralinguistic effect; the set of speech processing detectors, the speech recognition module, and the one or more machine learning models are configured to cooperate to automate labeling of the training data and pre-deployment training of the machine learning models trained for paralinguistic effects, wherein a first markup marker corresponds to a first paralinguistic effect and a second markup marker corresponds to a second paralinguistic effect.
7. 7. The apparatus of claim 6, wherein the machine learning model is trained on variations in prosody, including prosodic patterns, and variations in individual speech parameters to associate and understand i) individual phoneme, ii) individual word, and iii) individual non-word variations to specific paralinguistic effects detected by the set of speech processing detectors in the training data of the audio stream of the speech.
8. the machine learning model is trained to look for patterns; 7. The apparatus of claim 6, wherein the set of detectors is configured to analyze the training data to assist in annotating the training data such that the machine learning model can associate at least 1) a first particular paralinguistic effect with at least one of: i) an intended additional intended meaning, and ii) a more understandable way of conveying information with the first particular paralinguistic effect.
9. 7. The apparatus of claim 6, wherein the machine learning model is trained to learn how to identify each different paralinguistic effect and its corresponding markup marker and how to generate corresponding waveforms when conveying the first paralinguistic effect and the second paralinguistic effect.
10. a representation marker module configured to annotate the text representation of the individual words, phonemes, and / or non-words with the first markup markers corresponding to the first paralinguistic effects, the first markup markers indicating particular paralinguistic effects; The apparatus of claim 6 further comprising:
11. 1. A method for a conversational engagement platform, comprising: configuring one or more machine learning models trained to: examine audio data including at least one of i) words, ii) phonemes, and iii) non-words in a spoken communication annotated with one or more markup markers; and guide the generation of a text representation to cause pronunciations that differ from the plain pronunciation that would occur in the absence of the markup markers in order to convey at least one of 1) additional intended meaning and 2) enhanced understanding of the pronounced i) phonemes, ii) words, and / or iii) non-words themselves, as compared to plain pronunciations; wherein the one or more machine learning models have been trained with training data of speakers using paralinguistic effects; configuring a speech generation module to receive the generated text representation to guide the speech generation module to the different pronunciation states with the specific paralinguistic effect in a manner that better conveys 1) the additional intended meaning and / or 2) the enhanced understanding of the pronounced i) phonemes, ii) words, and / or iii) non-words themselves; The method, wherein any software portions of the speech generation module and the one or more machine learning models are stored on one or more non-transitory storage media in a state executable by one or more processors.
12. constructing a table to map a set of markup markers, each different markup marker being mapped to its corresponding specific paralinguistic effect on pronounced i) phonemes, ii) words, and / or iii) non-words; configuring the machine learning model to cooperate with the table to generate changes in the pronunciation of i) phonemes, ii) words, and / or iii) non-words when conveying a particular paralinguistic effect; 12. The method for a conversational engagement platform of claim 11, further comprising:
13. configuring a natural language generator module to generate textual representations of i) phonemes, ii) words, and / or iii) non-words; constructing a table of markup markers, each markup marker corresponding to its own paralinguistic effect; constructing a table to be referenced by the natural language generator module to mark up the text representation with one or more markup markers of the i) phonemes, ii) words, or iii) non-words so as to guide the speech generation module on how to pronounce the i) phonemes, ii) words, or iii) non-words using the paralinguistic effects to cause the pronunciation to differ from the plain pronunciation of the i) phonemes, ii) words, or iii) non-words; 12. The method for a conversational engagement platform of claim 11, further comprising:
14. 14. The method for a conversational engagement platform of claim 13, wherein the different pronunciation differs in one or more phonetic parameters from the plain pronunciation of that i) phoneme, ii) word, or iii) non-word by a threshold amount.
15. 12. The method for a conversational engagement platform of claim 11, wherein the machine learning model is trained by supervised machine learning on how to identify the paralinguistic effects and corresponding markup markers and understand the additional information conveyed by each particular paralinguistic effect.
16. 1. A method for a conversational engagement platform, comprising: configuring a set of speech processing detectors to analyze training data of an audio stream of speech from a speaking person, the set of speech processing detectors detecting pronounced words, phonemes, and non-words in the audio stream as well as speech parameters indicative of one or more paralinguistic effects; configuring a speech recognition module to receive the training data of the audio stream of speech and to create a text representation for each word, each non-word, each phoneme, and any combination thereof, in the training data of the audio stream of speech; training one or more machine learning models using a supervised machine learning neural network to train how to associate one or more markup markers with the text representation of each individual word, non-word, and / or phoneme pronounced with a particular paralinguistic effect; configuring the set of speech processing detectors, wherein the speech recognition module and the one or more machine learning models cooperate to automate labeling of the training data and pre-deployment training of the machine learning models trained for paralinguistic effects, wherein a first markup marker corresponds to a first paralinguistic effect and a second markup marker corresponds to a second paralinguistic effect; A method comprising:
17. 17. The method for a conversational engagement platform of claim 16, wherein the machine learning model is trained on a set of individual speech parameters and variations in prosody, including prosodic patterns, to associate and understand i) individual phoneme, ii) individual word, and iii) individual non-word variations with specific paralinguistic effects detected by the set of speech processing detectors in the training data of the audio stream of the speech.
18. configuring the machine learning model to be trained to search for patterns; analyzing the training data with the set of detectors to annotate the training data such that the machine learning model can associate at least 1) a first particular paralinguistic effect with at least one of: i) an intended additional intended meaning, and ii) a more understandable way of conveying information with the first particular paralinguistic effect; 20. The method for a conversational engagement platform of claim 16, further comprising:
19. configuring the machine learning model to train how to learn to identify each different paralinguistic effect and its corresponding markup marker, and how to generate corresponding waveforms when conveying the first paralinguistic effect and the second paralinguistic effect; 17. The method of claim 16, further comprising:
20. configuring a representational marker module to annotate the text representation of the individual words, phonemes, and / or non-words with the first markup markers corresponding to the first paralinguistic effects that indicate particular paralinguistic effects; 17. The method of claim 16, further comprising: