Detecting and using non-textual information in human speech

A deep learning-based system using a Transformer model like Whisper analyzes speech to identify intonation units and non-verbal cues, addressing the limitations of traditional methods in capturing prosodic information for comprehensive speech understanding.

WO2025141559A1PCT designated stage expired Publication Date: 2025-07-03YEDA RES & DEV CO LTD

Patent Information

Application Number
PCT/IL2024/051195
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2024-12-17
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Current approaches to computerized analysis of spoken text struggle with robust contextualization and heavily rely on narrow semantic searches, leading to partial interpretation of language, failing to effectively capture non-verbal linguistic information conveyed through prosody.

Method used

A deep learning-based system using a weakly-supervised encoder-decoder Transformer model, such as Whisper, is trained to recognize and transcribe speech, identifying intonation units and associating non-verbal labels like prosodic unit prototypes, discourse functions, emotions, and attitudes.

Benefits of technology

The system provides accurate and comprehensive analysis of non-verbal cues in speech, enhancing the understanding of human communication by effectively capturing prosodic information often missed by traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2024051195_03072025_PF_FP_ABST
    Figure IL2024051195_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Automatic recognition of non-verbal messages in speech, and in particular to detection or analysis of prosodic multilayered analysis of intonation units such as prosodic unit prototypes and their multi-labeled variations, may form a hierarchical classification for the analysis of non¬ verbal information or cues in speech. A speech captured by a microphone is fed to a weakly- supervised deep learning acoustic model for speech recognition and transcription, that may be based on encoder-decoder Transformer architecture, such as Whisper by OpenAI. The model is trained to output multiple words form the text in the captured speech, to identify Intonation Units (IUs) that include one or more words, and associate non-verbal labels to each of the IUs. The labels may indicate a prototype, a discourse function (such as a conversation action), an emotion, an emphasis, or an attitude, as well as a genre of a part of, or whole of, the entire captured speech.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DETECTING AND USING NON-VERBAL INFORMATION IN HUMAN SPEECH

[0002] RELATED APPLICATION

[0003] This patent application claims priority from U.S. Provisional Application Ser. No. 63 / 614,588, which was filed on December 24, 2023 and entitled: “Method and System for Detection of Non-Verbal Messages in Speech”, and from U.S. Provisional Application Ser. No. 63 / 555,115, which was filed on February 19, 2024 and entitled: “NON-VERBAL INFORMATION IN SPONTANEOUS SPEECET\ both of which are hereby incorporated herein by reference in their entirety.

[0004] TECHNICAL FIELD

[0005] This disclosure generally relates to an apparatus and method for automatic recognition of non-verbal messages in speech, and in particular to detection or analysis of prosodic multilayered analysis of intonation units such as prosodic unit prototypes and their multi-labeled variations, forming a hierarchical classification for the analysis of non-verbal information in speech, allowing for the detection of non-verbal cues in speech.

[0006] BACKGROUND

[0007] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art by inclusion in this section.

[0008] A human voice is much more than words; the voice has intonations, intentions, and patterns and makes us truly unique in our communication. A major part of human communication is conveyed through intonation or prosody - the music of speech. It is the least conscious, most instinctive activity of speech, conveying all non-verbal linguistic information including sentiment, emphasis, conversation action, and a large variety of signals: whether we are happy, irritated, or even truthful can be detected through how words are uttered. Non-verbal linguistic signals carry crucial information, encoded by the ‘music of speech’, termed ‘Prosody’. Prosodic messages range from conversation action (e.g., request, command), via discourse functions (e.g., narration, parentheticals), saliency of information (de / emphasis), and attitude (e.g., sarcasm), all the way to uninhibited emotion. Current approaches to computerized analysis of spoken text are generally heuristic, struggle for robust means of contextualization, and further rely heavily on narrow semantic searches, resulting in a very partial interpretation of language and text. Sound. In general, a sound is a mechanical wave that is an oscillation of pressure transmitted through some medium (like air or water), composed of frequencies within the range of hearing. In physics, sound is a vibration that propagates as an acoustic wave through a transmission medium such as a gas, liquid, or solid. As used herein, a sound is the reception of such waves and their perception by the brain, and typically comprises only acoustic waves that commonly have frequency components between about 20 Hz and 20 kHz, the audio frequency range, elicit an auditory percept in humans. In air at atmospheric pressure, these represent sound waves with wavelengths of 17 meters (56 ft) to 1.7 centimeters (0.67 in). Sound waves above 20 kHz are known as ultrasound and are not audible to humans. Sound waves below 20 Hz are known as infrasound. Different animal species have varying hearing ranges.

[0009] Voice. Voice consists of sounds made by a human being using the vocal folds for talking, singing, laughing, crying, screaming, etc. The human voice is specifically that part of human sound production in which the vocal folds (vocal cords) are the primary sound source. The human voice consists of sounds made by a human being using the vocal tract, including talking, singing, laughing, crying, screaming, shouting, humming, or yelling. The human voice frequency is specifically a part of human sound production in which the vocal folds (vocal cords) are the primary sound source. Other sound production mechanisms produced from the same general area of the body involve the production of unvoiced consonants, clicks, whistling, and whispering.

[0010] Speech. Speech is a human vocal communication using language. Each language uses phonetic combinations of vowel and consonant sounds that form the sound or phonology of its words (that is, all English words sound different from all French words, even if they are the same word, e.g., "role" or "hotel"), and using those words in their semantic value or character as words in the lexicon of a language according to the syntactic constraints that govern lexical words' function in a sentence. In speaking, speakers perform many different intentional speech acts, e.g., informing, declaring, asking, persuading, directing, and can use enunciation, intonation, degrees of loudness, tempo, and other non-representational or paralinguistic aspects of vocalization to convey meaning. In their speech, speakers also unintentionally communicate many aspects of their social position such as power relations in conversation, social rank, gender (sex), age, place of origin (through accent or dialect), physical states (alertness and sleepiness, vigor or weakness, health or illness), psychological states (emotions or moods), physio- psychological states (sobriety or drunkenness, normal consciousness and trance states), education or experience, and the like. Speech production is an unconscious multi-step process by which thoughts are generated into utterances. Production involves the unconscious mind selecting appropriate words and the appropriate form of those words from the lexicon and morphology, and the organization of those words through the syntax. Then, the phonetic properties of the words are retrieved and the sentence is articulated through the articulations associated with those phonetic properties.

[0011] Phonetics is the study of how the tongue, lips, jaw, vocal cords, and other speech organs are used to make sounds. Speech sounds are categorized by manner of articulation and place of articulation. Place of articulation refers to where in the neck or mouth the airstream is constricted. Manner of articulation refers to the manner in which the speech organs interact, such as how closely the air is restricted, what form of an airstream is used (e.g., pulmonic, implosive, ejectives, and clicks), whether or not the vocal cords are vibrating, and whether the nasal cavity is opened to the airstream. The concept is primarily used for the production of consonants but can be used for vowels in qualities such as voicing and nasalization. For any place of articulation, there may be several manners of articulation, and therefore several homorganic consonants. Normal human speech is pulmonic, produced with pressure from the lungs, which creates phonation in the glottis in the larynx, which is then modified by the vocal tract and mouth into different vowels and consonants. However, humans can pronounce words without the use of the lungs and glottis in laryngeal speech, of which there are three types: esophageal speech, pharyngeal speech, and buccal speech.

[0012] MP3. MP3 (formally MPEG-1 Audio Layer III or MPEG-2 Audio Layer III) is a coding format for digital audio developed largely by the Fraunhofer Society in Germany under the lead of Karlheinz Brandenburg. MP3 (or mp3) as a file format commonly designates files containing an elementary stream of MPEG-1 Audio or MPEG-2 Audio encoded data, without other complexities of the MP3 standard. MP3 uses lossy data compression to encode data using inexact approximations and the partial discarding of data. This allows a large reduction in file sizes when compared to uncompressed audio. The combination of small size and acceptable fidelity led to a boom in the distribution of music over the Internet in the mid-to-late 1990s, with MP3 serving as an enabling technology at a time when bandwidth and storage were still at a premium. The MP3 format soon became associated with controversies surrounding copyright infringement, music piracy, and the file ripping / sharing services MP3.com and Napster, among others. With the advent of portable media players, a product category also including smartphones, MP3 support remains near-universal.

[0013] MP3 compression works by reducing (or approximating) the accuracy of certain components of sound that are considered (by psychoacoustic analysis) to be beyond the hearing capabilities of most humans. This method is commonly referred to as perceptual coding or as psychoacoustic modelling. The remaining audio information is then recorded in a space-efficient manner, using MDCT and FFT algorithms. The Moving Picture Experts Group (MPEG) designed MP3 as part of its MPEG-1, and later MPEG-2, standards. MPEG-1 Audio (MPEG-1 Part 3), which included MPEG-1 Audio Layer I, II, and III, was published in 1993 as ISO / IEC 11172-3:1993. An MPEG-2 Audio (MPEG-2 Part 3) extension with lower sample- and bit-rates was published in 1995 as ISO / IEC 13818-3:1995, and it requires only minimal modifications to existing MPEG-1 decoders (recognition of the MPEG-2 bit in the header and addition of the new lower sample and bit rates).

[0014] WMA. Windows Media Audio (WMA) is a series of audio codecs and their corresponding audio coding formats developed by Microsoft. WMA consists of four distinct codecs. The original WMA codec, known simply as WMA, was conceived as a competitor to the popular MP3 and RealAudio codecs. WMA Pro, a newer and more advanced codec, supports multichannel and high-resolution audio. A lossless codec, WMA Lossless, compresses audio data without loss of audio fidelity (the regular WMA format is lossy). WMA Voice, targeted at voice content, applies compression using a range of low bit rates. A WMA file is in most circumstances contained in the Advanced Systems Format (ASF), a proprietary Microsoft container format for digital audio or digital video. The ASF container format specifies how metadata about the file is to be encoded, similar to the ID3 tags used by MP3 files. Metadata may include song name, track number, artist name, and audio normalization values. This container can optionally support Digital Rights Management (DRM) using a combination of elliptic curve cryptography key exchange, DES block cipher, a custom block cipher, RC4 stream cipher, and the SHA-1 hashing function.

[0015] Windows Media Audio (WMA) is the most common codec of the four WMA codecs. The colloquial usage of the term WMA, especially in marketing materials and device specifications, usually refers to this codec only. WMA is a lossy audio codec based on the study of psychoacoustics. Audio signals that are deemed to be imperceptible to the human ear are encoded with reduced resolution during the compression process. WMA can encode audio signals sampled at up to 48 kHz with up to two discrete channels (stereo). WMA 9 introduced variable bit rate (VBR) and average bit rate (ABR) coding techniques into the MS encoder although both were technically supported by the original format. WMA 9.1 also added support for low-delay audio, which reduces latency for encoding and decoding.

[0016] WAV. Waveform Audio File Format (WAVE, or WAV due to its filename extension; is an audio file format standard, developed by IBM and Microsoft, for storing an audio bitstream on personal computers. It is the main format used on Microsoft Windows systems for uncompressed audio. The usual bitstream encoding is the linear pulse-code modulation (LPCM) format. WAV is an application of the Resource Interchange File Format (RIFF) bitstream format method for storing data in chunks and thus is similar to the 8SVX and the Audio Interchange File Format (AIFF) format used on Amiga and Macintosh computers, respectively.

[0017] The WAV file is an instance of a Resource Interchange File Format (RIFF) defined by IBM and Microsoft. The RIFF format acts as a wrapper for various audio coding formats. Though a WAV file can contain compressed audio, the most common WAV audio format is uncompressed audio in the linear pulse-code modulation (LPCM) format. LPCM is also the standard audio coding format for audio CDs, which store two-channel LPCM audio sampled at 44.1 kHz with 16 bits per sample. Since LPCM is uncompressed and retains all of the samples of an audio track, professional users or audio experts may use the WAV format with LPCM audio for maximum audio quality. WAV files can also be edited and manipulated with relative ease using software. Since the sampling rate of a WAV file can vary from 1 Hz to 4.3 GHz, and the number of channels can be as high as 65,535, ‘.wav’ files have also been used for non-audio data.

[0018] Sounder. As used herein, a sounder is a component or device that converts electrical energy to sound waves transmitted through the air, an elastic solid material, or a liquid, usually utilizing a vibrating or moving ribbon or diaphragm. The sound may be audio or audible, having frequencies in the approximate range of 20 to 20,000 hertz, capable of being detected by human organs of hearing. Alternatively or in addition, the sounder may be used to emit inaudible frequencies, such as ultrasonic (a.k.a. ultrasound) acoustic frequencies that are above the range audible to the human ear, or above approximately 20,000 Hz. A sounder may be omnidirectional, unidirectional, bidirectional, or provide other directionality or polar patterns.

[0019] A loudspeaker (a.k.a. speaker) is a sounder that produces sound in response to an electrical audio signal input, typically audible sound. The most common form of loudspeaker is the electromagnetic (or dynamic) type, which uses a paper cone supporting a moving voice coil electromagnet acting on a permanent magnet. Where accurate reproduction of sound is required, multiple loudspeakers may be used, each reproducing a part of the audible frequency range. A loudspeaker is commonly optimized for middle frequencies; tweeters for high frequencies; and sometimes a super-tweeter is used which is optimized for the highest audible frequencies.

[0020] A loudspeaker may be a piezo (or piezoelectric) speaker that contains a piezoelectric crystal coupled to a mechanical diaphragm and is based on the piezoelectric effect. An audio signal is applied to the crystal, which responds by flexing in proportion to the voltage applied across the crystal surfaces, thus converting electrical energy into mechanical. Piezoelectric speakers are frequently used as beepers in watches and other electronic devices and are sometimes used as tweeters in less-expensive speaker systems, such as computer speakers and portable radios. A loudspeaker may be a magneto-strictive transducer, based on magnetostriction, that has been predominantly used as sonar ultrasonic sound wave radiators, but their usage has spread also to audio speaker systems.

[0021] A loudspeaker may be an electrostatic loudspeaker (ESL), in which sound is generated by the force exerted on a membrane suspended in an electrostatic field. Such speakers use a thin flat diaphragm usually consisting of a plastic sheet coated with a conductive material such as graphite sandwiched between two electrically conductive grids, with a small air gap between the diaphragm and grids. The diaphragm is usually made from a polyester film (thickness 2-20 pm) with exceptional mechanical properties, such as PET film. Utilizing the conductive coating and an external high-voltage supply, the diaphragm is held at a DC potential of several kilovolts with respect to the grids. The grids are driven by the audio signal, and the front and rear grids are driven in antiphase. As a result, a uniform electrostatic field proportional to the audio signal is produced between both grids. This causes a force to be exerted on the charged diaphragm, and its resulting movement drives the air on either side of it.

[0022] A loudspeaker may be a magnetic loudspeaker, that may be a ribbon or planar type, based on a magnetic field. A ribbon speaker consists of a thin metal-film ribbon suspended in a magnetic field. The electrical signal is applied to the ribbon, which moves with it to create the sound. Planar magnetic speakers are speakers with roughly rectangular flat surfaces that radiate in a bipolar (i.e., front and back) manner, and may have printed or embedded conductors on a flat diaphragm. Planar magnetic speakers consist of a flexible membrane with a voice coil printed or mounted on it. The current flowing through the coil interacts with the magnetic field of carefully placed magnets on either side of the diaphragm, causing the membrane to vibrate more uniformly and without much bending or wrinkling. A loudspeaker may be a bending wave loudspeaker, which uses an intentionally flexible diaphragm.

[0023] A sounder may be an electromechanical type, such as an electric bell, which may be based on an electromagnet, causing a metal ball to clap on a cup or half-sphere bell. A sounder may be a buzzer (or beeper), a chime, a whistle, or a ringer. Buzzers may be either electromechanical or ceramic -based piezoelectric sounders, which make a high-pitch noise and may be used for alerting. The sounder may emit a single or multiple tones and can be in continuous or intermittent operation. Speech-to-text (STT). Speech-to-text (STT), also known as speech recognition, Automatic Speech Recognition (ASR), or computer speech recognition, is a computerized, algorithmic process that transcribes a human’s spoken input into digital text in a written format. STT focuses on the translation of speech from a verbal format to a text one whereas voice recognition just seeks to identify an individual user’s voice. STT is typically implemented by a speech recognition software that enables the recognition and translation of spoken language into text through computational linguistics. STT is software that works by listening to audio and delivering an editable, verbatim transcript on a given device. A computer program draws on linguistic algorithms to sort auditory signals from spoken words and transfer those signals into text using characters called Unicode.

[0024] Some modem general-purpose speech recognition systems are based on Hidden Markov Models (HMMs). These are statistical models that output a sequence of symbols or quantities. HMMs are based on that a speech signal may be viewed as a piecewise stationary signal or a short-time stationary signal. In a short time-scale (e.g., 10 milliseconds), speech can be approximated as a stationary process. Speech can be thought of as a Markov model for many stochastic purposes. Further, HMMs may be trained automatically and are simple and computationally feasible. In speech recognition, the hidden Markov model would output a sequence of n-dimensional real-valued vectors (with n being a small integer, such as 10), outputting one every 10 milliseconds. The vectors would consist of cepstral coefficients, which are obtained by taking a Fourier transform of a short time window of speech and decorrelating the spectrum using a cosine transform, then taking the first (most significant) coefficients. The hidden Markov model will tend to have in each state a statistical distribution that is a mixture of diagonal covariance Gaussians, which will give a likelihood for each observed vector. Each word, or (for more general speech recognition systems), each phoneme, will have a different output distribution; a hidden Markov model for a sequence of words or phonemes is made by concatenating the individual trained hidden Markov models for the separate words and phonemes.

[0025] Dynamic Time Warping (DTW) is an approach that was historically used for speech recognition but has now largely been displaced by the more successful HMM-based approach. Dynamic time warping is an algorithm for measuring similarity between two sequences that may vary in time or speed. For instance, similarities in walking patterns would be detected, even if in one video the person was walking slowly and if in another he or she were walking more quickly, or even if there were accelerations and deceleration during the course of one observation. DTW has been applied to video, audio, and graphics - indeed, any data that can be turned into a linear representation can be analyzed with DTW.

[0026] Neural networks emerged as an attractive acoustic modeling approach in ASR in the late 1980s. Since then, neural networks have been used in many aspects of speech recognition such as phoneme classification, phoneme classification through multi-objective evolutionary algorithms, isolated word recognition, audiovisual speech recognition, audiovisual speaker recognition, and speaker adaptation. Neural networks make fewer explicit assumptions about feature statistical properties than HMMs and have several qualities making them attractive recognition models for speech recognition. When used to estimate the probabilities of a speech feature segment, neural networks allow discriminative training in a natural and efficient manner. However, in spite of their effectiveness in classifying short-time units such as individual phonemes and isolated words, early neural networks were rarely successful for continuous recognition tasks because of their limited ability to model temporal dependencies.

[0027] Deep Neural Networks and Denoising Autoencoders are also under investigation. A deep feedforward neural network (DNN) is an artificial neural network with multiple hidden layers of units between the input and output layers. Similar to shallow neural networks, DNNs can model complex non-linear relationships. DNN architectures generate compositional models, where extra layers enable composition of features from lower layers, giving a huge learning capacity and thus the potential of modeling criminative complex patterns of speech data.

[0028] Intonation. Intonation is the variation in pitch used to indicate the speaker's attitudes and emotions, to highlight or focus an expression, to signal the illocutionary act performed by a sentence, or to regulate the flow of discourse. Intonation is considered as the ‘melody’ of speech and generally refers to how the pitch of the voice rises and falls, and how speakers use this pitch variation to convey linguistic and pragmatic meaning. For example, the English question "Does Maria speak Spanish or French?” is interpreted as a yes-or-no question when it is uttered with a single rising intonation contour, but is interpreted as an alternative question when uttered with a rising contour on "Spanish" and a falling contour on "French". Although intonation is primarily a matter of pitch variation, its effects almost always work hand-in-hand with other prosodic features, and functions attributed to intonation such as the expression of attitudes and emotions, or highlighting aspects of grammatical structure, almost always involve concomitant variation in other prosodic features Intonation is distinct from tone, the phenomenon where pitch is used to distinguish words or to mark grammatical features.

[0029] Prosody, or intonation, is a critically important component of spoken communication. The automatic extraction of prosodic information is necessary for machines to process speech with human levels of proficiency. A thesis entitled: “Automatic Detection and Classification of Prosodic Events” by Andrew Rosenberg, submitted 2009 in partial fulfillment of the Requirements for the degree of Doctor of Philosophy in the Graduate School of Arts and Sciences at COLUMBIA UNIVERSITY, which is incorporated in its entirety for all purposes as if fully set forth herein, provides understanding of prosodic events and the use of prosody in spoken language processing towards the goal of human-like processing of speech by machines. The thesis describes work on the automatic detection and classification of prosodic events - specifically, pitch accents and prosodic phrase boundaries. The thesis presents novel techniques, feature representations, and state-of-the-art performance in each of these tasks, as well as three proof-of-concept applications - speech summarization, story segmentation, and non-native speech assessment - showing that access to hypothesized prosodic event information can be used to improve the performance of downstream spoken language processing tasks.

[0030] Intonation Unit (IU). An “Intonation Unit” (IU) is a piece of utterance, a continuous stream of meaningful sounds, bounded by a fairly perceptible decay of pitch, intensity and / or speech rate, pause, or any combination thereof. Typically, an IU is a unit of speech bounded by pauses that have movement, of music and rhythm, and spontaneous speech may be considered as produced in chunks called Intonation Units (IUS). Voice decaying or pausing in some sense is a way of packaging the information such that the lexical items put together in an intonation unit form certain psychological and lexicogrammatically realities or entities. Typical examples would be the inclusion of subordinate clauses and prepositional phrases in intonation units. Any feature of intonation may be analyzed and discussed against a background of tonic stress placement, choke of tones, and features or keys that are applicable to almost all intonation units. Closely related to the notion of pausing is that a change of meaning may be brought about; certain pauses in a stream of speech can have significant meaning variations in the message to be conveyed. Linguistic theory suggests that IUs pace the flow of information and serve as a window to the dynamic focus of attention in speech processing.

[0031] The Intonation Unit (IU) (also known as ‘Prosody Unit’, ‘Tone Group’, or (Intermediate) Intonational Phrase) is central in models of prosodic hierarchy and is hypothesized to be a universal characteristic of speech. Spontaneous speech is produced in chunks called Intonation Units (IUs), that are defined by a set of prosodic cues and occur in all human languages. In conversation, an IU is on average one second and (in English) three words long. Cross-linguistic evidence indicates that IUs elicit a neural response and form a consistent low-frequency rhythm (~lHz), leading to the hypothesis that they are utilized by the neural system in tracking speech. Qualitative analyses have shown that a single IU typically contains only one new "idea" (such as an entity, event, or state). Thus, it has been proposed that speech is produced as a concatenation of IUS to help regulate and manage the flow of information. In addition, the effective turn-taking mechanism underlying conversations has been found to rely on the discemibility of IU boundaries: the prosodic marking of completeness versus incompleteness.

[0032] In a typical speech, on average an IU includes 3-4 words and is 0.5 - 1.8 seconds long. However, an IU may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 words. Alternatively or in addition, an IU may include no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 words. Further, an IU may be at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long. Alternatively or in addition, an IU may last less than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, 2.5, or 3.0 seconds.

[0033] Identifying a neural response unique to the boundary defined by the IU is described in an article by Maya Inbar, Shir Genzer, Anat Perry, Eitan Grossman, and Ayelet N. Landau, posted on January 26, 2023 and entitled: “ Intonation Units in spontaneous speech evoke a neural response” [DOI: https: / / doi.org / 10.1101 / 2023.01.26.525707], which is incorporated in its entirety for all purposes as if fully set forth herein. In this article, a neural response unique to the boundary defined by the IU is identified. The EEG of participants who listened to different speakers recounting an emotional life event is measured, the speech stimuli linguistically analyzed, and modeled the EEG response at word offset using a GLM approach. The EEG response to lU-final words is found to differ from the response to lU-nonfinal words when acoustic boundary strength is held constant. It is demonstrated in spontaneous speech under naturalistic listening conditions, and under a theoretical framework that connects the prosodic chunking of speech, on the one hand, with the flow of information during communication, on the other. Finally, the findings of the body of research on rhythmic brain mechanism in speech processing are related by comparing the topographical distributions of neural speech tracking in model-predicted and empirical EEG. This qualitative comparison suggests that lU-related neural activity contributes to the previously characterized delta-band neural speech tracking.

[0034] Self-Supervised Learning (SSL) from speech data has produced models that have achieved remarkable performance in many tasks, and that are known to implicitly represent many aspects of information latently present in speech signals. However, relatively little is known about the suitability of such models for prosody-related tasks or the extent to which they encode prosodic information. A new evaluation framework, “SUPERB- prosody,” consisting of three prosody-related downstream tasks and two pseudo tasks, is presented in an article by Guan-Ting Lin, Chi-Luen Feng, Wei-Ping Huang, Yuan Tseng, Tzu-Han Lin, Chen- An Li, Hung-yi Lee, and Nigel G. Ward, published in 2022 IEEE Spoken Language Technology Workshop (SLT) [979-8-3503-9690-4 / 23 / $31.00 I DOI: 10.1109 / SLT54892.2023.10023234] entitled: “ON THE UTILITY OF SELF-SUPERVISED MODELS FOR PROSODY-RELATED TASKS”, which is incorporated in its entirety for all purposes as if fully set forth herein. The article finds that 13 of the 15 SSL models outperformed the baseline on all the prosody-related tasks. The article also shows good performance on two pseudo tasks: prosody reconstruction and future prosody prediction. The article further analyzes the layer-wise contributions of the SSL models. Overall the article concludes that SSL speech models are highly effective for prosody- related tasks.

[0035] The task of automatically detecting syllable stress is a key module in computer-assisted language learning systems. There are numerous studies proposed in the literature for automatic syllable stress detection by using different knowledge-based prosodic features. Also, different statistical machine learning and deep learning models are explored for this task using knowledge-based features. However, the acoustic parameters considered to compute knowledgebased features might not always represent the stress phenomena, hence the knowledge-based features are not always suitable for generalization and scalability. Recently, the rapidly emerging self-supervised learning-based representations are outperforming the existing state-of-the-art knowledge-based features in all speech applications. Also, these representations allow the models to be built in an end-to-end fashion. The use of self- supervised representations (Wav2Vec-2.0), for syllable stress detection and compare the performance with state-of-the-art knowledge-based features, is explored in an article by Jhansi Mallela, Sai Harshitha Aluru, and Chiranjeevi Yarra published in 2024 National Conference on Communications (NCC) by IEEE [979-8-3503-5922-0 / 24 / $31.00 I DOI: 10.1109 / NCC60321.2024.10486028] entitled: “Exploring the use of self-supervised representations for automatic syllable stress detection”, which is incorporated in its entirety for all purposes as if fully set forth herein. Further, the article uses recently proposed explicit representation learning framework, modeled by jointly optimizing variational autoencoder (VAE) and DNN for stress detection. The article analyzes the performance of representation learning framework with two different state-of-the-art classifiers, support vector machines (SVM) and simple deep neural network (DNN). The article describes experiments conducted on two non-native English speakers’ datasets from ISLE corpus i.e., German (GER), and Italian (ETA). From the analysis study, it is observed that the classification accuracy for syllable stress detection using self-supervised representations significantly improved by 3.2% and 2.7% over knowledge-based features in GER and ITA, respectively. From the t-SNE plots, it is observed that the representations learned by explicit representation learning framework with VAE show better discrimination among stressed and unstressed syllables compared with representations learned implicitly with simple DNN.

[0036] Automatic phrase boundary detection could be useful in applications, including computer-assisted pronunciation tutoring, spoken language understanding, and automatic speech recognition. The problem of phrase boundary detection on English utterances spoken by native American speakers is considered in an article by Pavan Kumar, Chiranjeevi Yarra, and Prasanta Kumar Ghosh, published IEEE 2021 in National Conference on Communications (NCC) [978- l-6654-4177-3 / 21 / $31.00; DOI: 10.1109 / NCC52529.2021.9530147] entitled: “DNN based phrase boundary detection using knowledge-based features and feature representations from CNN”, which is incorporated in its entirety for all purposes as if fully set forth herein. Most of the existing works on boundary detection use either knowledge-based features or representations learnt from a convolutional neural network (CNN) based architecture, considering word segments. However, the article hypothesizes that combining knowledge-based features and learned representations could improve the boundary detection task’s performance. For this, the article considers a fusion-based model considering deep neural network (DNN) and CNN, where CNNs are used for learning representations and DNN is used to combine knowledge-based features and learned representations. Further, unlike existing data-driven methods, the article considers two CNNs for learning representation, one for word segments and another for wordfinal syllable segments. Experiments on Boston University radio news and Switchboard corpora show the benefit of the proposed fusion-based approach compared to a baseline using knowledge-based features only and another baseline using feature representations from CNN only.

[0037] Automatic speech recognition (ASR) and natural language processing (NLP) are expected to benefit from an effective, simple, and reliable method to automatically parse conversational speech. The ability to parse conversational speech depends crucially on the ability to identify boundaries between prosodic phrases. This is done naturally by the human ear, yet it has proved surprisingly difficult to achieve reliably and simply in an automatic manner. Efforts to date have focused on detecting phrase boundaries using a variety of linguistic and acoustic cues. A method that does not require model training and utilizes two prosodic cues that are based on ASR output is described in an article by Tirza Biron, Daniel Baum, Dominik Freche, Nadav Matalon, Netanel Ehrmann, Eyal Weinreb, David Biron, and Elisha Moses published May 3, 2021, and entitled: “Automatic detection of prosodic boundaries in spontaneous speech”, which is incorporated in its entirety for all purposes as if fully set forth herein. Boundaries are identified using discontinuities in speech rate (pre-boundary lengthening and phrase-initial acceleration) and silent pauses. The resulting phrases preserve syntactic validity, exhibit pitch reset, and compare well with manual tagging of prosodic boundaries. Collectively, the article supports the notion of prosodic phrases that represent coherent patterns across textual and acoustic parameters.

[0038] A prosodic speech recognition engine configured to identify prosodic features and patterns in a speech continuum is disclosed in U.S. Patent No. 11,600,264 to Moses et al. entitled: “ Extracting content from speech prosody”, which is incorporated in its entirety for all purposes as if fully set forth herein. The prosodic speech recognition engine is for the extraction of linguistic content including para-syntactic content, discourse function, information structure, meaning, and speaker sentiment.

[0039] Examples of peak values in pitch and intensity-pitch are shown in a graph view 10 in FIG. 1, illustrating curves of normalized unit over time for a spoken text of “Paul will have some beer”, which includes 5 words and 15 phones. The graphs view 10 comprises a graph of interpolated pitch curve 12, a graph of intensity curve 11, and a pitch-intensity curve 13 that is estimated by multiplying the scaled pitch and intensity. Curves peak values are marked, and measure the distance between the maxima and the nearby minima.

[0040] Expressive reading, considered the defining attribute of oral reading fluency, comprises the prosodic realization of phrasing and prominence. In the context of evaluating oral reading, it helps to establish the speaker’s comprehension of the text. We consider a labeled dataset of children’s reading recordings for the speaker-independent detection of prominent words using acoustic-prosodic and lexico-syntactic features. A previous well-tuned random forest ensemble predictor is replaced by an RNN sequence classifier to exploit potential context dependency across the longer utterance. Further, deep learning is applied to obtain word-level features from low-level acoustic contours of fundamental frequency, intensity and spectral shape in an end-to- end fashion. Performance comparisons are presented across the different feature types and across different feature learning architectures for prominent word prediction to draw insights wherever possible.

[0041] Mel-Frequency Cepstrum (MFC). The Mel-Frequency Cepstrum (MFC) is a representation of the short-term power spectrum of a sound, based on a linear cosine transform of a log power spectrum on a nonlinear Mel scale of frequency. Mel-Frequency Cepstral Coefficients (MFCCs) are coefficients that collectively make up an MFC. The difference between the Cepstrum and the Mel-frequency Cepstrum is that in the MFC, the frequency bands are equally spaced on the Mel scale, which approximates the human auditory system's response more closely than the linearly-spaced frequency bands used in the normal Cepstrum. This frequency warping can allow for better representation of sound, for example, in audio compression.

[0042] MFCCs are commonly derived by the steps of: taking the Fourier transform of (a windowed excerpt of) a signal, mapping the powers of the spectrum obtained above onto the Mel scale, using triangular overlapping windows; taking the logs of the powers at each of the Mel frequencies; taking the discrete cosine transform of the list of Mel log powers, as if it were a signal, and the MFCCs are the amplitudes of the resulting spectrum. There can be variations in this process, for example: differences in the shape or spacing of the windows used to map the scale, or the addition of dynamics features such as "delta" and "delta-delta" (first- and second- order frame-to-frame difference) coefficients. Calculating and using MFCC is further described in European Telecommunications Standards Institute (ETSI) 2003 Standard ETSI ES 201 108 vl.1.3 (2003-09) entitled: “ Speech Processing, Transmission and Quality Aspects (STQ); Distributed speech recognition; Front-end feature extraction algorithm; Compression algorithms", in an article in J. Computer Science & Technology, 16(6):582-589, Sept. 2001 by Fang Zheng, Guoliang Zhang, and Zhanjiang Song entitled: “Comparison of Different Implementations of MFCC”, and in RWTH Aachen, University of Technology, Aachen Germany publication by Sirko Molau, Michael Pitz, Ralf Schluter, and Hermann Ney, entitled: “Computing MEL-Frequency Cepstral Coefficients on the Power Spectrum”, which are all incorporated in their entirety for all purposes as if fully set forth herein.

[0043] TTS. Speech synthesis is the artificial production of human speech. A computer system used for this purpose is called a speech synthesizer and can be implemented in software or hardware products. A text- to- speech (TTS) system converts normal language text into speech; other systems render symbolic linguistic representations like phonetic transcriptions into speech. Synthesized speech can be created by concatenating pieces of recorded speech that are stored in a database. Systems differ in the size of the stored speech units; a system that stores phones or diphones provides the largest output range but may lack clarity. For specific usage domains, the storage of entire words or sentences allows for high- quality output. Alternatively, a synthesizer can incorporate a model of the vocal tract and other human voice characteristics to create a completely "synthetic" voice output. The quality of a speech synthesizer is judged by its similarity to the human voice and by its ability to be understood clearly. An intelligible text- to- speech program allows people with visual impairments or reading disabilities to listen to written words on a home computer. Practical, engineering aspects of text-to-speech and speech synthesis are described in a book by Paul Taylor published 2009 by Cambridge University Press [978-0-521-89927-7] entitled: “Text-to- Speech Synthesis”, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0044] A text-to-speech system (or "engine") is composed of two parts: a front-end and a back- end. The front-end has two major tasks. First, it converts raw text containing symbols like numbers and abbreviations into the equivalent of written-out words. This process is often called text normalization, pre-processing, or tokenization. The front-end then assigns phonetic transcriptions to each word and divides and marks the text into prosodic units, like phrases, clauses, and sentences. The process of assigning phonetic transcriptions to words is called text- to-phoneme or grapheme-to-phoneme conversion. Phonetic transcriptions and prosody information together make up the symbolic linguistic representation that is output by the frontend. The back-end — often referred to as the synthesizer — then converts the symbolic linguistic representation into sound. In certain systems, this part includes the computation of the target prosody (pitch contour, phoneme durations), which is then imposed on the output speech.

[0045] The two primary technologies generating synthetic speech waveforms are concatenative synthesis and formant synthesis. Concatenative synthesis is based on the concatenation (stringing together) of segments of recorded speech. Generally, concatenative synthesis produces the most natural- sounding synthesized speech. However, differences between natural variations in speech and the nature of the automated techniques for segmenting the waveforms sometimes result in audible glitches in the output. There are three main sub-types of concatenative synthesis. Formant synthesis does not use human speech samples at runtime. Instead, the synthesized speech output is created using additive synthesis and an acoustic model (physical modelling synthesis). Parameters such as fundamental frequency, voicing, and noise levels are varied over time to create a waveform of artificial speech. This method is sometimes called rules-based synthesis; however, many concatenative systems also have rules- based components. Many systems based on formant synthesis technology generate artificial, robotic- sounding speech that would never be mistaken for human speech. However, maximum naturalness is not always the goal of a speech synthesis system, and formant synthesis systems have advantages over concatenative systems. Formant-synthesized speech can be reliably intelligible, even at very high speeds, avoiding the acoustic glitches that commonly plague concatenative systems.

[0046] A short but comprehensive overview of text-to-speech synthesis by highlighting its natural language processing (NLP) and digital signal processing (DSP) components is provided in a paper by M. Z. Rashad; Nikos Mastorakis; Hazem M. El-Bakry; and Islam R. Isma'il published July 2010 LATEST TRENDS on COMMUNICATIONS and INFORMATION TECHNOLOGY [ISSN: 1792-4316; ISBN: 978-960-474-207-3] entitled: “An Overview of Text-To-Speech Synthesis Techniques” , which is incorporated in its entirety for all purposes as if fully set forth herein. First, the front-end or the NLP component comprised of text analysis, phonetic analysis, and prosodic analysis is introduced then two rule-based synthesis techniques (formant synthesis and articulatory synthesis) are explained. After that concatenative synthesis is explored. Compared to rule-based synthesis, concatenative synthesis is simpler since there is no need to determine speech production rules. However, concatenative synthesis introduces the challenges of prosodic modification to speech units and resolving discontinuities at unit boundaries. Prosodic modification results in artifacts in the speech that make the speech sound unnatural. Unit selection synthesis, which is a kind of concatenative synthesis, solves this problem by storing numerous instances of each unit with varying prosodies. The unit that best matches the target prosody is selected and concatenated. Finally, hidden Markov model (HMM) synthesis is introduced.

[0047] Text- to- speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a research topic in speech, language, and machine learning communities and has broad applications in the industry. With the development of deep learning and artificial intelligence, neural network-based TTS has significantly improved the quality of synthesized speech in recent years. A comprehensive survey on neural TTS, aiming to provide a good understanding of current research and future trends, is described in a paper by Xu Tan, Tao Qin, Frank Soong, and Tie-Yan Liu published 23 Jul 2021 [https: / / doi.org / 10.48550 / arXiv.2106.15561] entitled: “A Survey on Neural Speech Synthesis”, which is incorporated in its entirety for all purposes as if fully set forth herein. The key components in neural TTS, including text analysis, acoustic models, and vocoders, and several advanced topics, including fast TTS, low-resource TTS, robust TTS, expressive TTS, and adaptive TTS, etc., are disclosed. Resources related to TTS (e.g., datasets, and open-source implementations) are further summarized.

[0048] Prosody. The term “Prosody” refers to the sound of syllables, words, phrases, and sentences produced by pitch variation, rhythm variation, stress, and intensity variation, as well as timbre (or voice quality) variation of the human voice. Pitch is a perceptual property that allows the ordering of sounds on a frequency-related scale, and is a major auditory attribute of musical tones, along with duration, loudness, and timbre. Prosody may reflect various features of the speaker or the utterance: the emotional state of the speaker; the modality of the utterance (statement, question, or command); the presence of irony or sarcasm; emphasis, contrast, and focus; or other elements of language that may not be encoded by the morpho-syntax or choice of vocabulary. In sign languages, prosody involves the rhythm, length, and tension of gestures, along with mouthing and facial expressions. Prosody is largely absent from writing, which can occasionally result in reader misunderstanding. Orthographic conventions to mark or substitute for prosody include punctuation (commas, exclamation marks, question marks, scare quotes, and ellipses), and typographic styling for emphasis (italic, bold, and underlined text).

[0049] Prosodic features involve the magnitude, duration, and changing over time characteristics of acoustic parameters of the spoken voice, such as Tempo (fast or slow), timbre or harmonics (few or many), pitch level, and in particular pitch variations (high or low), envelope (sharp or round), pitch contour (up or down), amplitude and amplitude variations (small or large), tonality mode (major or minor), and rhythmic or non-rhythmic behavior.

[0050] Prosody comprises aspects of speech that communicate information beyond written words related to syntax, sentiment, intent, discourse, and comprehension. Decades of research have confirmed the importance of prosody in human speech perception and production, yet spoken language technology has made limited use of prosodic information. This limitation is due to several reasons. Words (written or transcribed) are often treated as discrete units while speech signals are continuous, which makes it challenging to combine these two modalities appropriately in spoken language systems. In addition, as variable as text can often be, text has fewer sources of variation than speech. Different meanings of a written or transcribed sentence can be communicated through punctuation, but a sentence can be spoken in many more ways, where prosody is often essential in conveying information not reflected in the word sequence. Moreover, given the highly variable nature of speech, most successful systems require a lot of data that covers these different aspects, which in turn requires powerful computing technology that was not available until recently.

[0051] A thesis submitted 2020 by Trang Tran entitled: “Neural Models for Integrating Prosody in Spoken Language Understanding” , which is incorporated in its entirety for all purposes as if fully set forth herein, describes mechanisms for integrating prosody in spoken language systems, using spontaneous and expressive speech. This thesis focuses on two language understanding tasks: (a) constituency parsing (identifying the syntactic structure of a sentence), motivated by the fact that prosodic boundaries align with constituent boundaries, and (b) dialog act recognition (identifying the segmentation and intents of utterances in discourse), motivated by the fact that prosodic boundaries signal dialog act boundaries, and intonational cues help disambiguate intents. Both parsing and dialog act recognition are important components of spoken language systems. From the modeling perspective, the thesis proposes a method for integrating prosody effectively in spoken language understanding systems, which is shown empirically to advance the state of the art in parsing and dialog act recognition tasks. Further, the described methods can be extended to other spoken language processing tasks. Speech understanding has a broad impact on many areas, as it facilitates accessibility and allows for more natural human-computer interactions in education, health care, elder care, and Al-assisted domains in general.

[0052] Transformer. As used herein, a ‘transformer’ is a deep learning model that adopts the mechanism of self-attention, differentially weighting the significance of each part of the input data. It is used primarily in the fields of Natural Language Processing (NLP) and Computer Vision (CV). Like recurrent neural networks (RNNs), transformers are designed to process sequential input data, such as natural language, with applications towards tasks such as translation and text summarization. However, unlike RNNs, transformers process the entire input all at once. The attention mechanism provides context for any position in the input sequence. For example, if the input data is a natural language sentence, the transformer does not have to process one word at a time. This allows for more parallelization than RNNs and therefore reduces training times.

[0053] A new simple network architecture, a Transformer, that is based solely on attention mechanisms, dispensing with recurrence and convolutions entirely, is described in an article by Ashish Vaswani; Noam Shazeer; Niki Parmar; Jakob Uszkoreit; Llion Jones; Aidan N. Gomez; Lukasz Kaiser; and Ulia Polosukhin dated 12 Jun 2017 (arXiv: 1706.03762 [cs.CL]; https: / / doi.org / 10.48550 / arXiv.1706.0376) and entitled: “Atention Is All You Need”, which is incorporated in its entirety for all purposes as if fully set forth herein. The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best-performing models also connect the encoder and decoder through an attention mechanism. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. The article shows that the transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.

[0054] Like earlier seq2seq models, the original Transformer model architecture uses an encoder-decoder architecture. First, the input text is parsed into tokens by a byte pair encoding tokenizer, and each token is converted via a word embedding into a vector. Then, the positional information of the token is added to the word embedding. The encoder consists of encoding layers that process the input iteratively one layer after another, while the decoder consists of decoding layers that do the same thing to the encoder's output. The function of each encoder layer is to generate encodings that contain information about which parts of the inputs are relevant to each other. It passes its encodings to the next encoder layer as inputs. Each decoder layer does the opposite, taking all the encodings and using their incorporated contextual information to generate an output sequence. To achieve this, each encoder and decoder layer makes use of an attention mechanism. For each input, attention weighs the relevance of every other input and draws from them to produce the output. Each decoder layer has an additional attention mechanism that draws information from the outputs of previous decoders before the decoder layer draws information from the encodings. Both the encoder and decoder layers have a feed-forward neural network for additional processing of the outputs and contain residual connections and layer normalization steps.

[0055] The transformer building blocks are scaled dot-product attention units. When a sentence is passed into a transformer model, attention weights are calculated between every token simultaneously. The attention unit produces embeddings for every token in the context that contains information about the token itself along with a weighted combination of other relevant tokens each weighted by its attention weight. One set of matrices is called an attention head, and each layer in a transformer model has multiple attention heads. While each attention head attends to the tokens that are relevant to each token, with multiple attention heads the model can do this for different definitions of "relevance". In addition, the influence field representing relevance can become progressively dilated in successive layers. Many transformer attention heads encode relevant relations that are meaningful to humans. For example, attention heads can attend mostly to the next word, while others mainly attend from verbs to their direct objects. The computations for each attention head can be performed in parallel, which allows for fast processing. The outputs for the attention layer are concatenated to pass into the feed-forward neural network layers.

[0056] Each encoder consists of two major components: a self- attention mechanism and a feedforward neural network. The self-attention mechanism accepts input encodings from the previous encoder and weighs their relevance to each other to generate output encodings. The feed-forward neural network further processes each output encoding individually. These output encodings are then passed to the next encoder as its input, as well as to the decoders. The first encoder takes positional information and embeddings of the input sequence as its input, rather than encodings. The positional information is necessary for the transformer to make use of the order of the sequence because no other part of the transformer makes use of this. The encoder is bidirectional. Attention can be placed on tokens that are before and after the current token. Each decoder consists of three major components: a self-attention mechanism, an attention mechanism over the encodings, and a feed-forward neural network. The decoder functions in a similar fashion to the encoder, but an additional attention mechanism is inserted which instead draws relevant information from the encodings generated by the encoders. This mechanism can also be called the encoder-decoder attention. Like the first encoder, the first decoder takes positional information and embeddings of the output sequence as its input, rather than encodings. The transformer must not use the current or future output to predict an output, so the output sequence must be partially masked to prevent this reverse information flow. This allows for autoregressive text generation. For all attention heads, attention can't be placed on following tokens. The last decoder is followed by a final linear transformation and softmax layer, to produce the output probabilities over the vocabulary.

[0057] The remarkable success of transformers in the field of natural language processing has sparked the interest of the speech-processing community, leading to an exploration of their potential for modeling long-range dependencies within speech sequences. Recently, transformers have gained prominence across various speech-related domains, including automatic speech recognition, speech synthesis, speech translation, speech para-linguistics, speech enhancement, spoken dialogue systems, and numerous multimodal applications. A comprehensive survey that aims to bridge research studies from diverse subfields within speech technology is presented in a paper published 21 Mar 2023 [https: / / doi.org / 10.48550 / arXiv.2303.11607] and authored by Siddique Latif, Aun Zaidi, Heriberto Cuayahuitl, Fahad Shamshad, Moazzam Shoukat, and Junaid Qadir, entitled: “Transformers in Speech Processing: A Survey”, which is incorporated in its entirety for all purposes as if fully set forth herein. Consolidating findings from across the speech technology landscape provides a valuable resource for researchers interested in harnessing the power of transformers to advance the field. The paper identifies the challenges encountered by transformers in speech processing while also offering insights into potential solutions to address these issues.

[0058] Whisper. ‘Whisper’ is a machine learning model for speech recognition and transcription, created by OpenAI (headquartered in San Francisco, California, U.S.A. 94110) and first released as open-source software in September 2022. The Whisper model is capable of transcribing speech in English and several other languages and is also capable of translating several non-English languages into English. The combination of different training data used in its development has led to improved recognition of accents, background noise, and jargon compared to previous approaches. Whisper is a weakly- supervised deep learning acoustic model, made using an encoder-decoder transformer architecture, and has been trained using semi- supervised learning on 680,000 hours of multilingual and multitask data, of which about one-fifth (117,000 hours) were non-English audio data. Whisper has a differing error rate with respect to transcribing different languages, with a higher word error rate in languages not well- represented in the training data. The model has been used as the base for a unified model for speech recognition and more general sound recognition. The Whisper architecture is based on an encoder-decoder transformer. Input audio is split into 30-second chunks converted into a Mel- frequency Cepstrum, which is passed to an encoder. A decoder is trained to predict later text captions. Special tokens are used to perform several tasks such as phrase-level timestamps

[0059] ‘Whisper’ is described, for example, in a study published Sep 21, 2022 by OpenAI, San Francisco, CA 94110, USA, authored by Alec Radford Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Hya Sutskever, entitled: “Robust Speech Recognition via Large-Scale Weak Supervision” , which is attached to this document, and is incorporated in its entirety for all purposes as if fully set forth herein. The study describes the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any finetuning. When compared to humans, the models approach their accuracy and robustness. The study further describes models and inference code to serve as a foundation for further work on robust speech processing.

[0060] A schematic architecture 20 of the Whisper model is shown in FIG. 2, which is based on a simple end-to-end approach, implemented as an encoder-decoder Transformer. It will take the audio recording, split it into 30-second chunks, and process them one by one. For each 30- second recording, it will encode the audio using the encoder section, and save the position of each word said, and leverage this encoded information to find what was said using the decoder. Input audio is split into 30-second chunks, converted into a Mel-frequency Cepstrum shown as a log-Mel spectrogram 21, which is then processed by a “2X ConvlD + GELU” block 21a. After sinusoidal positional encoding 22, the data is then passed into a set of encoders 23a, 23b, 23c, and 23d. A corresponding set of decoders 48a, 48b, 48c, and 48d is trained to predict the corresponding text caption, intermixed by learned positional encoding 37a with special tokens in multitask training format 47a, that direct the single model to perform tasks such as language identification, phrase-level timestamps, multilingual speech transcription, and to-English speech translation. The decoders set 48 will predict next-token prediction 47b from all this information, which is basically each word being said. Then, it will repeat this process for the next word using all the same information as well as the predicted previous words, helping it guess the next one that would make more sense.

[0061] Self-attention mechanisms have enabled transformers to achieve superhuman-level performance on many Speech-To-Text (STT) tasks, yet the challenge of automatic prosodic segmentation has remained unsolved. A paper by Nathan Roll, Calbert Graham, and Simon Todd dated February 2023 [DOI: 10.48550 / arXiv .2302.01984] entitled: “PSST! Prosodic Speech Segmentation with Transformers” , which is incorporated in its entirety for all purposes as if fully set forth herein, discloses a finetune Whisper, a pretrained STT model, to annotate intonation unit (IU) boundaries by repurposing low-frequency tokens. The approach in the paper achieves an accuracy of 95.8%, outperforming previous methods without the need for large-scale labeled data or enterprise-grade compute resources. The paper describes diminishing input signals by applying a series of filters, finding that low pass filters at a 3.2 kHz level improve segmentation performance in out of sample and out of distribution contexts. The model may be used as both a transcription tool and a baseline for further improvements in prosodic segmentation. In one example, PSST applies a method that is based fine-tuning of the WHISPER model on a transcription, which is enriched with lU-boundary tags, trained on the data of Santa Barbara Corpus (SBC).

[0062] SDK. As used herein, the term Software Development Kit (SDK) refers to a specific software package, software framework, hardware platform, or a set of development tools and the like at the time of the establishment of the operating system software. Typically, an SDK includes a programming package that enables a programmer to develop applications for a specific platform, and may include one or more APIs, programming tools, and documentation. It may be as simple as the implementation of one or more Application Programming Interfaces (APIs) in the form of some libraries to interface to a particular programming language or to include sophisticated hardware that can communicate with a particular embedded system. Common tools include debugging facilities and other utilities, often presented in an Integrated Development Environment (IDE). The SDKs also frequently include sample code and supporting technical notes or other supporting documentation to help clarify points made by the primary reference material. Some SDKs may have attached licenses that make them unsuitable for building software intended to be developed under an incompatible license. For example, a proprietary SDK will probably be incompatible with free software development, while a GPL- licensed SDK could be incompatible with proprietary software development. LGPL SDKs are typically safe for proprietary development. A software engineer typically receives the SDK from the target system developer. Often the SDK can be downloaded directly via the Internet or via SDKs marketplaces. Many SDKs are provided for free to encourage developers to use the system or language. Sometimes this is used as a marketing tool. Freely offered SDKs may still be able to monetize, based on user data taken from the apps, which may serve the interests of big players in the ecosystem, for example, the operating system. An SDK for an operating system add-on (for instance, QuickTime for classic Mac OS) may include the add-on software itself to be used for development purposes, albeit not necessarily for redistribution together with the developed product. Any part of, or the whole of, any of the methods described herein may be provided as part of, or used as, an SDK.

[0063] API. An Application Programming Interface (API) refers to an intermediary software serving as the interface allowing the interaction and data sharing between an application software and the application platform, across which few or all services are provided, and commonly used to expose or use a specific software functionality, while protecting the rest of the application. The API may be based on, or according to, the Portable Operating System Interface (POSIX) standard, defining the API along with command line shells and utility interfaces for software compatibility with variants of Unix and other operating systems, such as POSIX.1-2008 that is simultaneously IEEE STD. 1003.1™ - 2008 entitled: “Standard for Information Technology - Portable Operating System Interface (POSIX(R)) Description” , and The Open Group Technical Standard Base Specifications, Issue ?, IEEE STD. 1003.1™, 2013 Edition. Any part of, or the whole of, any of the methods described herein may be provided as part of, or used as, an Application Programming Interface (API).

[0064] Feature engineering. Feature engineering (or feature extraction or feature discovery) is the process of extracting features (characteristics, properties, attributes) from raw data. Feature engineering refers to the process of using domain knowledge to select and transform the most relevant variables from raw data when creating a predictive model using machine learning or statistical modeling. The goal of feature engineering and selection is to improve the performance of Machine Learning (ML) algorithms. The feature engineering pipeline is the preprocessing steps that transform raw data into features that can be used in machine learning algorithms, such as predictive models. Predictive models consist of an outcome variable and predictor variables, and it is during the feature engineering process that the most useful predictor variables are created and selected for the predictive model. Feature engineering in ML consists of four main steps: Feature Creation, Transformations, Feature Extraction, and Feature Selection.

[0065] Feature engineering consists of the creation, transformation, extraction, and selection of features, also known as variables, that are most conducive to creating an accurate ML algorithm. Feature Creation refers to creating features that involves identifying the variables that will be most useful in the predictive model. This is a subjective process that requires human intervention and creativity. Existing features are mixed via addition, subtraction, multiplication, and ratio to create new derived features that have greater predictive power. Transformation involves manipulating the predictor variables to improve model performance; e.g., ensuring the model is flexible in the variety of data it can ingest; ensuring variables are on the same scale, making the model easier to understand; improving accuracy; and avoiding computational errors by ensuring all features are within an acceptable range for the model. Feature extraction is the automatic creation of new variables by extracting them from raw data. The purpose of this step is to automatically reduce the volume of data into a more manageable set for modeling. Some feature extraction methods include cluster analysis, text analytics, edge detection algorithms, and principal components analysis. Feature selection algorithms essentially analyze, judge, and rank various features to determine which features are irrelevant and should be removed, which features are redundant and should be removed, and which features are most useful for the model and should be prioritized.

[0066] Feature engineering may be used to get better results when training a model, and the feature engineering process can help improve results by modifying the data’s features to better capture the nature of the problem. Ideas on how to improve the performance of a machine learning solution are provided in a guide by PABEO DUBOUE published 2020 by Cambridge University Press [DOI: 10.1017 / 9781108671682; ISBN: 9781108671682] entitled: “The Art of Feature Engineering Essentials for Machine Learning”, which is incorporated in its entirety for all purposes as if fully set forth herein. Beginning with the basic concepts and techniques of feature engineering, this guide builds up to a unique cross-domain approach that spans data on graphs, texts, time series, and images, with fully worked-out case studies. Key topics include binning, out-of-fold estimation, feature selection, dimensionality reduction, and encoding variable-length data.

[0067] Feature engineering can improve the performance of machine learning models by creating relevant and informative features from raw data. By engineering features, ME models can make more accurate predictions, handle complex and distributed data, reduce overfitting, and extract valuable insights from categorical and numerical data. Common techniques used in feature engineering include one-hot encoding, feature scaling, handling missing values (e.g., imputation), creating interaction features (e.g., polynomial features), dimensionality reduction (e.g., PCA), feature selection (e.g., using statistical tests or feature importance), and transforming variables (e.g., logarithmic or power transformations). When handling missing data during feature engineering, instances may be removed with missing values, fill in missing values with mean / median / mode, or use more advanced techniques like imputation methods (e.g., K- nearest neighbors or regression imputation). For outliers, you can consider removing them if they are shown to be incorrect or transform them using techniques like winsorization or capping.

[0068] Feature engineering techniques for machine learning may include Imputation, referring to handling missing values in data. While deleting records missing specific values is one way of dealing with this issue, it could also mean losing out on valuable data. Imputation may be classified into two types: Categorical Imputation, where missing categorical variables are generally replaced by the most commonly occurring value in other record, and Numerical Imputation, where missing numerical values are generally replaced by the mean of the corresponding value in other records. Discretization involves logically grouping sets of data values together into bins (or buckets). Binning can apply to numerical values as well as to categorical data values, and may help prevent data from overfitting but comes at the cost of loss of granularity of data. The grouping of data may involve grouping of equal intervals; grouping based on equal frequencies (of observations in the bin); or grouping based on decision tree sorting (to establish a relationship with the target).

[0069] Categorical encoding refers to encoding categorical features into numerical values, which are usually simpler for an algorithm to understand, such as a ‘One Hot Encoding’ (OHE) technique of categorical encoding. Categorical values may be converted into simple numerical 1’s and 0’s without losing information. OHE may substantially increase the number of features and result in highly correlated features. Splitting features into parts may improve the value of the features toward the target to be learned. Outliers refer to unusually high or low values in the dataset, which are unlikely to occur in normal scenarios. Since these outliers may adversely affect the prediction, they are handled appropriately, such as by removal, where the records containing outliers are removed from the distribution; replacing values, where the outliers may be treated as missing values and replaced by using appropriate imputation; and capping the maximum and minimum values and replacing them with an arbitrary value or a value from a variable distribution.

[0070] Variable transformation techniques help with normalizing skewed data. In one example, such a transformation is based on logarithmic transformation, that operates to compress the larger numbers and relatively expand the smaller numbers. This, in turn, results in less skewed values, especially in the case of heavy-tailed distributions. Other variable transformations used include square root and Box-Cox transformations, which generalize the former two. Feature scaling, sometimes referred to as feature normalization, is used due to the sensitivity of some machine learning algorithms to the scale of the input values, and may include Min-Max Scaling, which involves rescaling all values in a feature from 0 to 1; and Standardization / Variance Scaling, where all the data points are subtracted by their mean, and the result is divided by the distribution's variance to arrive at a distribution with a 0 mean and variance of 1. Feature creation involves deriving new features from existing ones, such as by performing simple mathematical operations such as aggregations to obtain the mean, median, mode, sum, or difference and even product of two values. Although derived directly from the given input data, these features can impact the performance when carefully chosen to relate to the target.

[0071] Microphone. A microphone comprises an electroacoustic sensor that responds to sound waves (which are essentially vibrations transmitted through an elastic solid or a liquid or gas), and converts sound into electrical energy, usually by means of a ribbon or diaphragm set into motion by the sound waves. The sound may be audio or audible, having frequencies in the approximate range of 20 to 20,000 hertz, capable of being detected by human organs of hearing. Alternatively or in addition, the microphone may be used to sense inaudible frequencies, such as ultrasonic (a.k.a. ultrasound) acoustic frequencies that are above the range audible to the human ear, or above approximately 20,000 Hz. A microphone may be a condenser microphone (a.k.a. capacitor or electrostatic microphone) where the diaphragm acts as one plate of a two plates capacitor, and the vibrations changes the distance between plates, hence changing the capacitance. An electret microphone is a capacitor microphone based on a permanent charge of an electret or a polarized ferroelectric material. A dynamic microphone is based on electromagnetic induction, using a diaphragm attached to a small movable induction coil that is positioned in a magnetic field of a permanent magnet. The incident sound waves cause the diaphragm to vibrate, and the coil to move in the magnetic field, producing a current. Similarly, a ribbon microphone uses a thin, usually corrugated metal ribbon suspended in a magnetic field, and its vibration within the magnetic field generates the electrical signal. A loudspeaker is commonly constructed similar to a dynamic microphone, and thus may be used as a microphone as well. In a carbon microphone, the diaphragm vibrations apply varying pressure to a carbon, thus changing its electrical resistance. A piezoelectric microphone (a.k.a. crystal or piezo microphone) is based on the phenomenon of piezoelectricity in piezoelectric crystals such as potassium sodium tartrate. A microphone may be omnidirectional, unidirectional, bidirectional, or provide other directionality or polar patterns. Actuator. As used herein, an actuator is any component, device, or functionality that affects, creates, or changes a physical phenomenon associated with an object, and the object may be gas, air, liquid, or solid. The actuator may be controlled by a digital input, and may be electrical actuator powered by an electrical energy. The actuator may be operative to affect timedependent characteristic such as a time-integrated, an average, an RMS (Root Mean Square) value, a frequency, a period, a duty-cycle, a time-integrated, or a time-derivative, of the affected or produced phenomenon. Further, the actuator may be operative to affect or change spacedependent characteristics of the phenomenon, such as a pattern, a linear density, a surface density, a volume density, a flux density, a current, a direction, a rate of change in a direction, or a flow, of the sensed phenomenon. Examples of actuators are described in U.S. Patent Application Publication No. 2013 / 0201316 to Binder el al., entitled: “System and Method for Server Based Control”, which is incorporated in its entirety for all purposes as if fully set forth herein, that discloses various home and building automation systems in a building or vehicle for an actuator operation in response to a sensor according to a control logic. Any actuator herein may include one or more actuators, each affecting or generating a physical phenomenon in response to an electrical command, which can be an electrical signal (such as voltage or current), or by changing a characteristic (such as resistance or impedance) of a device. The actuators may be identical, similar or different from each other, and may affect or generate the same or different phenomena. Two or more actuators may be connected in series or in parallel. The actuator command signal may be conditioned by a signal conditioning circuit. In the case of analog actuator, a digital to analog (D / A) converter may be used to convert the digital command data to analog signals for controlling the actuators.

[0072] Any actuator herein may be a light source used to emit light by converting electrical energy into light, and where the luminous intensity may be fixed or may be controlled, commonly for illumination or indication purposes. The actuator may be used to activate or control the light emitted by a light source, being based on converting electrical energy or another energy to a light. The light emitted may be a visible light, or invisible light such as infrared, ultraviolet, X-ray or gamma rays. A shade, reflector, enclosing globe, housing, lens, and other accessories may be used, typically as part of a light fixture, in order to control the illumination intensity, shape or direction. Electrical sources of illumination commonly use a gas, a plasma (such as in arc and fluorescent lamps), an electrical filament, or Solid-State Lighting (SSL), where semiconductors are used. An SSL may be a Light-Emitting Diode (LED), an Organic LED (OLED), Polymer LED (PLED), or a laser diode. A light source may consist of, or may comprise, a lamp which may be an arc lamp, a fluorescent lamp, a gas-discharge lamp (such as a fluorescent lamp), or an incandescent light (such as a halogen lamp). An arc lamp is the general term for a class of lamps that produce light by an electric arc voltaic arc. Such a lamp consists of two electrodes, first made from carbon but typically made today of tungsten, which are separated by a noble gas.

[0073] Any actuator herein may comprise, or may consist of, a motion actuator that may be a rotary actuator that produces a rotary motion or torque, commonly to a shaft or axle. The motion produced by a rotary motion actuator may be either continuous rotation, such as in common electric motors, or movement to a fixed angular position as for servos and stepper motors. A motion actuator may be a linear actuator that creates motion in a straight line. A linear actuator may be based on an intrinsically rotary actuator, by converting from a rotary motion created by a rotary actuator, using a screw, a wheel and axle, or a cam. A screw actuator may be a leadscrew, a screw jack, a ball screw or roller screw. A wheel-and-axle actuator operates on the principle of the wheel and axle, and may be hoist, winch, rack and pinion, chain drive, belt drive, rigid chain, or rigid belt actuator. Similarly, a rotary actuator may be based on an intrinsically linear actuator, by converting from a linear motion to a rotary motion, using the above or other mechanisms. Motion actuators may include a wide variety of mechanical elements and / or prime movers to change the nature of the motion such as provided by the actuating / transducing elements, such as levers, ramps, screws, cams, crankshafts, gears, pulleys, constant-velocity joints, or ratchets. A motion actuator may be part of a servomotor system.

[0074] A motion actuator may be a pneumatic actuator that converts compressed air into rotary or linear motion, and may comprises a piston, a cylinder, valves, or ports. Motion actuators are commonly controlled by an input pressure to a control valve, and may be based on moving a piston in a cylinder. A motion actuator may be a hydraulic actuator using a pressure of the liquid in a hydraulic cylinder to provide force or motion. A hydraulic actuator may be a hydraulic pump, such as a vane pump, a gear pump, or a piston pump. A motion actuator may be an electric actuator where electrical energy may be converted into motion, such as an electric motor. A motion actuator may be a vacuum actuator producing a motion based on vacuum pressure.

[0075] An electric motor may be a DC motor, which may be a brushed, brushless, or uncommutated type. An electric motor may be a stepper motor, and may be a Permanent Magnet (PM) motor, a Variable reluctance (VR) motor, or a hybrid synchronous stepper. An electric motor may be an AC motor, which may be an induction motor, a synchronous motor, or an eddy current motors. An AC motor may be a two-phase AC servo motor, a three-phase AC synchronous motor, or a single-phase AC induction motor, such as a split-phase motor, a capacitor start motor, or a Permanent-Split Capacitor (PSC) motor. Alternatively or in addition, an electric motor may be an electrostatic motor, and may be MEMS based. A rotary actuator may be a fluid power actuator, and a linear actuator may be a linear hydraulic actuator or a pneumatic actuator. A linear actuator may be a piezoelectric actuator, based on the piezoelectric effect, may be a wax motor, or may be a linear electrical motor, which may be a DC brush, a DC brushless, a stepper, or an induction motor type. A linear actuator may be a telescoping linear actuator. A linear actuator may be a linear electric motor, such as a linear induction motor (LIM), or a Linear Synchronous Motor (LSM).

[0076] A motion actuator may be a linear or rotary piezoelectric motor based on acoustic or ultrasonic vibrations. A piezoelectric motor may use piezoelectric ceramics such as Inchworm or PiezoWalk motors, may use Surface Acoustic Waves (SAW) to generate the linear or the rotary motion, or may be a Squiggle motor. Alternatively or in addition, an electric motor may be an ultrasonic motor. A linear actuator may be a micro- or nanometer comb-drive capacitive actuator. Alternatively or in addition, a motion actuator may be a Dielectric or Ionic based Electroactive Polymers (EAPs) actuator. A motion actuator may also be a solenoid, thermal bimorph, or a piezoelectric unimorph actuator.

[0077] Any actuator herein may be a pump, typically used to move (or compress) fluids or liquids, gasses, or slurries, commonly by pressure or suction actions, and the activating mechanism is often reciprocating or rotary. A pump may be a direct lift, impulse, displacement, valveless, velocity, centrifugal, vacuum pump, or gravity pump. A pump may be a positive displacement pump, such as a rotary-type positive displacement type such as internal gear, screw, shuttle block, flexible vane or sliding vane, circumferential piston, helical twisted roots or liquid ring vacuum pumps, a reciprocating-type positive displacement type, such as piston or diaphragm pumps, and a linear-type positive displacement type, such as rope pumps and chain pumps, a rotary lobe pump, a progressive cavity pump, a rotary gear pump, a piston pump, a diaphragm pump, a screw pump, a gear pump, a hydraulic pump, and a vane pump. A rotary positive displacement pumps may be a gear pump, a screw pump, or a rotary vane pumps. Reciprocating positive displacement pumps may be plunger pumps type, diaphragm pumps type, diaphragm valves type, or radial piston pumps type. A pump may be an impulse pump such as hydraulic ram pumps type, pulser pumps type, or airlift pumps type. A pump may be a rotodynamic pump such as a velocity pump or a centrifugal pump. A centrifugal pump may be a radial flow pump type, an axial flow pump type, or a mixed flow pump. Any actuator herein may be an electrochemical or chemical actuator, used to produce, change, or otherwise affect a matter structure, properties, composition, process, or reactions, such as oxidation / reduction or an electrolysis process.

[0078] Any actuator herein may be a sounder which converts electrical energy to sound waves transmitted through the air, an elastic solid material, or a liquid, usually by means of a vibrating or moving ribbon or diaphragm. The sound may be audible or inaudible (or both), and may be omnidirectional, unidirectional, bidirectional, or provide other directionality or polar patterns. A sounder may be an electromagnetic loudspeaker, a piezoelectric speaker, an electrostatic loudspeaker (ESL), a ribbon or planar magnetic loudspeaker, or a bending wave loudspeaker. A sounder may be an electromechanical type, such as an electric bell, a buzzer (or beeper), a chime, a whistle or a ringer and may be either electromechanical or ceramic-based piezoelectric sounders. The sounder may emit a single or multiple tones, and can be in continuous or intermittent operation.

[0079] The sounder may be used to play digital audio content, either stored in, or received by, the sounder, the actuator unit, the router, the control server, or any combination thereof. The audio content stored may be either pre-recorded or using a synthesizer. Few digital audio files may be stored, selected by a control logic. Alternatively or in addition, the source of the digital audio may be a microphone serving as a sensor. In another example, the system uses the sounder for simulating the voice of a human being or generates music. The music produced, can emulate the sounds of a conventional acoustical music instrument, such as a piano, tuba, harp, violin, flute, guitar and so forth. A talking human voice may be played by the sounder, either prerecorded or using human voice synthesizer, and the sound may be a syllable, a word, a phrase, a sentence, a short story or a long story, and can be based on speech synthesis or pre-recorded, using male or female voice. A human speech may be produced using a hardware, software (or both) speech synthesizer, which may be Text-To-Speech (TTS) based. The speech synthesizer may be a concatenative type, using unit selection, diphone synthesis, or domain- specific synthesis. Alternatively or in addition, the speech synthesizer may be a formant type, and may be based on articulatory synthesis or hidden Markov models (HMM) based.

[0080] Any actuator herein may be used to generate an electric or magnetic field, and may be an electromagnetic coil or an electromagnet.

[0081] Any actuator herein may be a display for presentation of visual data or information, commonly on a screen, and may consist of an array (e.g., matrix) of light emitters or light reflectors, and may present text, graphics, image or video. A display may be a monochrome, gray-scale, or color type, and may be a video display. The display may be a projector (commonly by using multiple reflectors), or alternatively (or in addition) have the screen integrated. A projector may be based on an Eidophor, Liquid Crystal on Silicon (LCoS or LCOS), or LCD, or may use Digital Light Processing (DLP™) technology, and may be MEMS based or be a virtual retinal display. A video display may support Standard-Definition (SD) or High-Definition (HD) standards, and may support 3D. The display may present the information as scrolling, static, bold or flashing. The display may be an analog display, such as having NTSC, PAL or SECAM formats. Similarly, analog RGB, VGA (Video Graphics Array), SVGA (Super Video Graphics Array), SCART or S-video interface, or may be a digital display, such as having IEEE 1394 interface (a.k.a. FireWire™), may be used. Other digital interfaces that can be used are USB, SDI (Serial Digital Interface), HDMI (High-Definition Multimedia Interface), DVI (Digital Visual Interface), UDI (Unified Display Interface), DisplayPort, Digital Component Video or DVB (Digital Video Broadcast) interface. Various user controls may include an on / off switch, a reset button and others. Other exemplary controls involve display associated settings such as contrast, brightness and zoom.

[0082] A display may be a Cathode-Ray Tube (CRT) display, or a Liquid Crystal Display (LCD) display. The LCD display may be passive (such as CSTN or DSTN based) or active matrix, and may be Thin Film Transistor (TFT) or LED-backlit LCD display. A display may be a Field Emission Display (FED), Electroluminescent Display (ELD), Vacuum Fluorescent Display (VFD), or may be an Organic Light-Emitting Diode (OLED) display, based on passivematrix (PMOLED) or active-matrix OLEDs (AMOLED). A display may be based on an Electronic Paper Display (EPD), and be based on Gyricon technology, Electro-Wetting Display (EWD), or Electrofluidic display technology. A display may be a laser video display or a laser video projector, and may be based on a Vertical-Extemal-Cavity Surface-Emitting-Laser (VECSEL) or a Vertical-Cavity Surface-Emitting Laser (VCSEL).

[0083] Any actuator herein may be a thermoelectric actuator such as a cooler or a heater for changing the temperature of a solid, liquid or gas object, and may use conduction, convection, thermal radiation, or by the transfer of energy by phase changes. A heater may be a radiator using radiative heating, a convector using convection, or a forced convection heater. A thermoelectric actuator may be a heating or cooling heat pump, and may be electrically powered, compression-based cooler using an electric motor to drive a refrigeration cycle. A thermoelectric actuator may be an electric heater, converting electrical energy into heat, using resistance, or a dielectric heater. A thermoelectric actuator may be a solid-state active heat pump device based on the Peltier effect. A thermoelectric actuator may be an air cooler, using a compressor-based refrigeration cycle of a heat pump. An electric heater may be an induction heater.

[0084] Any actuator herein may include a signal generator serving as an actuator for providing an electrical signal (such as a voltage or current), or may be coupled between the processor and the actuator for controlling the actuator. A signal generator may be an analog or digital signal generator, and may be based on software (or firmware) or may be a separated circuit or component. A signal may generate repeating or non-repeating electronic signals, and may include a digital to analog converter (DAC) to produce an analog output. Common waveforms are a sine wave, a saw-tooth, a step (pulse), a square, and a triangular waveforms. The generator may include some sort of modulation functionality such as Amplitude Modulation (AM), Frequency Modulation (FM), or Phase Modulation (PM). A signal generator may be an Arbitrary Waveform Generators (AWGs) or a logic signal generator.

[0085] Any actuator herein may be a light source that emits visible or non-visible light (infrared, ultraviolet, X-rays, or gamma rays) such as for illumination or indication. The actuator may comprise a shade, a reflector, an enclosing globe, or a lens, for manipulating the emitted light. The light source may be an electric light source for converting electrical energy into light, and may consist of, or comprise, a lamp, such as an incandescent, a fluorescent, or a gas discharge lamp. The electric light source may be based on Solid-State Lighting (SSL) such as a Light Emitting Diode (LED) which may be Organic LED (OLED), a polymer LED (PLED), or a laser diode. The actuator may be a chemical or electrochemical actuator, and may be operative for producing, changing, or affecting a matter structure, properties, composition, process, or reactions, such as producing, changing, or affecting an oxidation / reduction or an electrolysis reaction.

[0086] Any actuator herein may be a motion actuator and may cause linear or rotary motion or may comprise a conversion mechanism (may be based on a screw, a wheel and axle, or a cam) for converting to rotary or linear motion. The conversion mechanism may be based on a screw, and the system may include a leadscrew, a screw jack, a ball screw or a roller screw, or may be based on a wheel and axle, and the system may include a hoist, a winch, a rack and pinion, a chain drive, a belt drive, a rigid chain, or a rigid belt. The motion actuator may comprise a lever, a ramp, a screw, a cam, a crankshaft, a gear, a pulley, a constant-velocity joint, or a ratchet, for affecting the produced motion. The motion actuator may be a pneumatic actuator, a hydraulic actuator, or an electrical actuator. The motion actuator may be an electrical motor such as brushed, a brushless, or an uncommutated DC motor, or a Permanent Magnet (PM) motor, a Variable reluctance (VR) motor, or a hybrid synchronous stepper DC motor. The electrical motor may be an induction motor, a synchronous motor, or an eddy current AC motor. The AC motor may be a single-phase AC induction motor, a two-phase AC servo motor, or a three-phase AC synchronous motor, and may be a split-phase motor, a capacitor-start motor, or a Permanent-Split Capacitor (PSC) motor. The electrical motor may be an electrostatic motor, a piezoelectric actuator, or a MEMS-based motor.

[0087] The motion actuator may be a linear hydraulic actuator, a linear pneumatic actuator, or a linear electric motor such as linear induction motor (LIM) or a Linear Synchronous Motor (LSM). The motion actuator may be a piezoelectric motor, a Surface Acoustic Wave (SAW) motor, a Squiggle motor, an ultrasonic motor, or a micro- or nanometer comb-drive capacitive actuator, a Dielectric or Ionic based Electroactive Polymers (EAPs) actuator, a solenoid, a thermal bimorph, or a piezoelectric unimorph actuator.

[0088] Any actuator herein may be operative to move, force, or compress a liquid, a gas or a slurry, and may be a compressor or a pump. The pump may be a direct lift, an impulse, a displacement, a valveless, a velocity, a centrifugal, a vacuum, or a gravity pump. The pump may be a positive displacement pump such as a rotary lobe, a progressive cavity, a rotary gear, a piston, a diaphragm, a screw, a gear, a hydraulic, or a vane pump. The positive displacement pump may be a rotary-type positive displacement pump such as an internal gear, a screw, a shuttle block, a flexible vane, a sliding vane, a rotary vane, a circumferential piston, a helical twisted roots, or a liquid ring vacuum pump. The positive displacement pump may be a reciprocating-type positive displacement type such as a piston, a diaphragm, a plunger, a diaphragm valve, or a radial piston pump. The positive displacement pump may be a linear-type positive displacement type such as rope-and-chain pump. The pump may be an impulse pump such as a hydraulic ram, a pulser, or an airlift pump. The pump may be a rotodynamic pump, such as a velocity pump or a centrifugal pump, that may be a radial flow, an axial flow, or a mixed flow pump.

[0089] Any actuator herein may be a sounder for converting an electrical energy to emitted audible or inaudible sound waves, emitted as omnidirectional, unidirectional, or bidirectional pattern. The sound may be audible, and the sounder may be an electromagnetic loudspeaker, a piezoelectric speaker, an electrostatic loudspeaker (ESL), a ribbon or planar magnetic loudspeaker, or a bending wave loudspeaker. The sounder may be electromechanical or ceramic based, and may be operative to emit a single or multiple tones, and may be operative to continuous or intermittent operation. The sounder may be an electric bell, a buzzer (or beeper), a chime, a whistle or a ringer. The sounder may be a loudspeaker, and the system may be operative to play one or more digital audio content files (which may include a pre-recorded audio) stored entirely or in part in the second device, the router, or the control server. The system may comprise a synthesizer for producing the digital audio content. The sensor may be a microphone for capturing the digital audio content to play by the sounder. The control logic or the system may be operative to select one of the digital audio content files, and may be operative for playing the selected file by the sounder. The digital audio content may be music, and may include the sound of an acoustical musical instrument such as a piano, a tuba, a harp, a violin, a flute, or a guitar. The digital audio content may be a male or female human voice saying a syllable, a word, a phrase, a sentence, a short story or a long story. The system may comprise a speech synthesizer (such as a Text-To-Speech (TTS) based) for producing a human speech, being part of the second device, the router, the control server, or any combination thereof. The speech synthesizer may be a concatenative type, and may use unit selection, diphone synthesis, or domain-specific synthesis. Alternatively or in addition, the speech synthesizer may be a formant type, articulatory synthesis based, or hidden Markov models (HMM) based.

[0090] Any actuator herein may be a thermoelectric actuator (such as an electric thermoelectric actuator) and may be a heater or a cooler, and may be operative for affecting or changing the temperature of a solid, a liquid, or a gas object. The thermoelectric actuator may be coupled to the object by conduction, convection, force convention, thermal radiation, or by the transfer of energy by phase changes. The thermoelectric actuator may include a heat pump, or may be a cooler based on an electric motor based compressor for driving a refrigeration cycle. The thermoelectric actuator may be an induction heater, may be an electric heater such as a resistance heater or a dielectric heater, or may be solid-state based such as an active heat pump device based on the Peltier effect. The actuator may be an electromagnetic coil or an electromagnet and may be operative for generating a magnetic or electric field.

[0091] Any actuator herein may directly or indirectly create, change or otherwise affect the rate of change of the physical quantity (gradient) versus the direction around a particular location, or between different locations. For example, a temperature gradient may describe the differences in the temperature between different locations. Further, an actuator may affect time-dependent or time-manipulated values of the phenomenon, such as time-integrated, average or Root Mean Square (RMS or rms), relating to the square root of the mean of the squares of a series of discrete values (or the equivalent square root of the integral in a continuously varying value). Further, a parameter relating to the time dependency of a repeating phenomenon may be affected, such as the duty-cycle, the frequency (commonly measured in Hertz - Hz) or the period. An actuator may be based on the Micro Electro-Mechanical Systems - MEMS (a.k.a. Micro-mechanical electrical systems) technology. An actuator may affect environmental conditions such as temperature, humidity, noise, vibration, fumes, odors, toxic conditions, dust, and ventilation.

[0092] Any actuator herein may change, increase, reduce, or otherwise affect the amount of a property or of a physical quantity or the magnitude relating to a physical phenomenon, body or substance. Alternatively or in addition, any actuator herein may be used to affect the time derivative thereof, such as the rate of change of the amount, the quantity or the magnitude. In the case of space related quantity or magnitude, an actuator may affect the linear density, relating to the amount of property per length, an actuator may affect the surface density, relating to the amount of property per area, or an actuator may affect the volume density, relating to the amount of property per volume. In the case of a scalar field, an actuator may further affect the quantity gradient, relating to the rate of change of property with respect to position. Alternatively or in addition, an actuator may affect the flux (or flow) of a property through a cross-section or surface boundary. Alternatively or in addition, an actuator may affect the flux density, relating to the flow of property through a cross-section per unit of the cross-section, or through a surface boundary per unit of the surface area. Alternatively or in addition, an actuator may affect the current, relating to the rate of flow of property through a cross-section or a surface boundary, or the current density, relating to the rate of flow of property per unit through a cross-section or a surface boundary. An actuator may include or consists of a transducer, defined herein as a device for converting energy from one form to another for the purpose of measurement of a physical quantity or for information transfer. Further, a single actuator may be used to affect two or more phenomena. For example, two characteristics of the same element may be affected, each characteristic corresponding to a different phenomenon. An actuator may have multiple states, where the actuator state is depending upon the control signal input. An actuator may have a two- state operation such as 'on' (active) and 'off (non-active), based on a binary input such as 'O' or '1', or 'true' and 'false'. In such a case, it can be activated by controlling an electrical power supplied or switched to it, such as by an electric switch,

[0093] Any actuator herein may be a light source used to emit light by converting electrical energy into light, and where the luminous intensity is fixed or may be controlled, commonly for illumination or indicating purposes. Further, an actuator may be used to activate or control the light emitted by a light source, being based on converting electrical energy or other energy to a light. The light emitted may be a visible light, or invisible light such as infrared, ultraviolet, X- ray or gamma rays. A shade, reflector, enclosing globe, housing, lens, and other accessories may be used, typically as part of a light fixture, in order to control the illumination intensity, shape or direction. The illumination (or the indication) may be steady, blinking or flashing. Further, the illumination can be directed for lighting a surface, such as a surface including an image or a picture. Further, a single-state visual indicator may be used to provide multiple indications, for example by using different colors (of the same visual indicator), different intensity levels, variable duty-cycle and so forth.

[0094] Electrical sources of illumination commonly use a gas, a plasma (such as in an arc and fluorescent lamps), an electrical filament, or Solid-State Lighting (SSL), where semiconductors are used. An SSL may be a Light-Emitting Diode (LED), an Organic LED (OLED), or Polymer LED (PLED). Further, an SSL may be a laser diode, which is a laser whose active medium is a semiconductor, commonly based on a diode formed from a p-n junction and powered by the injected electric current.

[0095] A light source may consist of, or comprise, a lamp, which is typically replaceable and is commonly radiating a visible light. A lamp, sometimes referred to as 'bulb', may be an arc lamp, a Fluorescent lamp, a gas-discharge lamp, or an incandescent light. An arc lamp (a.k.a. arc light) is the general term for a class of lamps that produce light by an electric arc (also called a voltaic arc). Such a lamp consists of two electrodes, first made from carbon but typically made today of tungsten, which are separated by a gas. The type of lamp is often named by the gas contained in the bulb; including Neon, Argon, Xenon, Krypton, Sodium, metal Halide, and Mercury, or by the type of electrode as in carbon-arc lamps. The common fluorescent lamp may be regarded as a low-pressure mercury arc lamp.

[0096] Any actuator herein may be a thermoelectric actuator such as a cooler or a heater for changing the temperature of an object, that may be solid, liquid or gas (such as the air temperature), using conduction, convection, thermal radiation, or by the transfer of energy by phase changes. Radiative heaters contain a heating element that reaches a high temperature. The element is usually packaged inside a glass envelope resembling a light bulb and with a reflector to direct the energy output away from the body of the heater. The element emits infrared radiation that travels through air or space until it hits an absorbing surface, where it is partially converted to heat and partially reflected. In a convection heater, the heating element heats the air next to it by convection. Hot air is less dense than cool air, so it rises due to buoyancy, allowing more cool air to flow in to take its place. This sets up a constant current of hot air that leaves the appliance through vent holes and heats up the surrounding space. These are generally filled with oil, in an oil heater, due to oil functioning as an effective heat reservoir. They are ideally suited for heating a closed space. They operate silently and have a lower risk of ignition hazard in the event that they make unintended contact with furnishings compared to radiant electric heaters. This is a good choice for long periods of time, or if left unattended. A fan heater, also called a forced convection heater, is a variety of convection heater that includes an electric fan to speed up the airflow. This reduces the thermal resistance between the heating element and the surroundings faster than passive convection, allowing heat to be transferred more quickly.

[0097] A thermoelectric actuator may be a heat pump, which is a machine or device that transfers thermal energy from one location, called the "source," which is at a lower temperature, to another location called the "sink" or "heat sink", which is at a higher temperature. Heat pumps may be used for cooling or for heating. Thus, heat pumps move thermal energy opposite to the direction that it normally flows, and may be electrically driven such as compressor-driven air conditioners and freezers. A heat pump may use an electric motor to drive a refrigeration cycle, drawing energy from a source such as the ground or outside air and directing it into the space to be warmed. Some systems can be reversed so that the interior space is cooled and the warm air is discharged outside or into the ground.

[0098] A thermoelectric actuator may be an electric heater, converting electrical energy into heat, such as for space heating, cooking, water heating, and industrial processes. Commonly, the heating element inside every electric heater is simply an electrical resistor, and works on the principle of Joule heating: an electric current through a resistor converts electrical energy into heat energy. In a dielectric heater, high-frequency alternating electric field, or radio wave or microwave electromagnetic radiation heats a dielectric material, and is based on heating caused by molecular dipole rotation within the dielectric. Microwave heaters, as distinct from RF heating, is a sub-category of dielectric heating at frequencies above 100 MHz, where an electromagnetic wave can be launched from a small dimension emitter and conveyed through space to the target. Modem microwave ovens make use of electromagnetic waves (microwaves) with electric fields of much higher frequency and shorter wavelength than RF heaters. Typical domestic microwave ovens operate at 2.45 GHz, but 0.915 GHz ovens also exist, thus the wavelengths employed in microwave heating are 12 or 33 cm, providing for highly efficient, but less penetrative, dielectric heating.

[0099] Any actuator herein may use pneumatics, involving the application of pressurized gas to affect mechanical motion. A motion actuator may be a pneumatic actuator that converts energy (typically in the form of compressed air) into rotary or linear motion. In some arrangements, a motion actuator may be used to provide force or torque. Similarly, force or torque actuators may be used as motion actuators. A pneumatic actuator mainly consists of a piston, a cylinder, and valves or ports. The piston is covered by a diaphragm, or seal, which keeps the air in the upper portion of the cylinder, allowing air pressure to force the diaphragm downward, moving the piston underneath, which in turn moves the valve stem, which is linked to the internal parts of the actuator. Pneumatic actuators may only have one spot for a signal input, top or bottom, depending on the action required. Valves input pressure is the "control signal", where each different pressure is a different set point for a valve. Valves typically require little pressure to operate and usually double or triple the input force. The larger the size of the piston, the larger the output pressure can be. Having a larger piston can also be good if air supply is low, allowing the same forces with less input.

[0100] Any actuator herein may use hydraulics, involving the application of a fluid to affect mechanical motion. Common hydraulics systems are based on Pascal's famous theory, which states that the pressure of the liquid produced in an enclosed structure has the capacity of releasing a force up to ten times the pressure that was produced earlier. A hydraulic actuator may be a hydraulic cylinder, where pressure is applied to the fluids (oil), to get the desired force. The force acquired is used to power the hydraulic machine. These cylinders typically include the pistons of different sizes, used to push down the fluids in the other cylinder, which in turn exerts the pressure and pushes it back again. A hydraulic actuator may be a hydraulic pump, is responsible for supplying the fluids to the other essential parts of the hydraulic system. The power generated by a hydraulic pump is about ten times more than the capacity of an electrical motor. There are different types of hydraulic pumps such as the vane pumps, gear pumps, piston pumps, etc. Among them, the piston pumps are relatively more costly, but they have a guaranteed long life and are even able to pump thick, difficult fluids. Further, a hydraulic actuator may be a hydraulic motor, where the power is achieved with the help of exerting pressure on the hydraulic fluids, which is normally oil. The benefit of using hydraulic motors is that when the power source is mechanical, the motor develops a tendency to rotate in the opposite direction, thus acting like a hydraulic pump.

[0101] Any actuator herein may be used to generate an electric or magnetic field. An electromagnetic coil (sometimes referred to simply as a "coil") is formed when a conductor (usually an insulated solid copper wire) is wound around a core or form, to create an inductor or electromagnet. One loop of wire is usually referred to as a turn, and a coil consists of one or more turns. Coils are often coated with varnish or wrapped with insulating tape to provide additional insulation and secure them in place. A completed coil assembly with taps is often called a winding. An electromagnet is a type of magnet in which the magnetic field is produced by the flow of electric current, and disappears when the current is turned off. A simple electromagnet consisting of a coil of insulated wire wrapped around an iron core. The strength of the magnetic field generated is proportional to the amount of current. Any actuator herein may produce a physical, chemical, or biological action, stimulation or phenomenon, such as a changing or generating temperature, humidity, pressure, audio, vibration, light, motion, sound, proximity, flow rate, electrical voltage, and electrical current, in response to the electrical input (current or voltage). For example, an actuator may provide visual or audible signaling, or physical movement. An actuator may include motors, winches, fans, reciprocating elements, extending or retracting, and energy conversion elements, as well as a heater or a cooler

[0102] Any actuator herein may be or may include a visual or audible signaling device, or any other device that indicates a status to the person. In one example, the device illuminates a visible light, such as a Light-Emitting-Diode (LED). However, any type of visible electric light emitter such as a flashlight, an incandescent lamp and compact fluorescent lamps can be used. Multiple light emitters may be used, and the illumination may be steady, blinking or flashing. Further, the illumination can be directed for lighting a surface, such as a surface including an image or a picture. Further, a single single-state visual indicator may be used to provide multiple indications, for example, by using different colors (of the same visual indicator), different intensity levels, variable duty-cycle and so forth.

[0103] Any actuator herein may include a solenoid, which is typically a coil wound into a packed helix, and used to convert electrical energy into a magnetic field. Commonly, an electromechanical solenoid is used to convert energy into linear motion. Such electromagnetic solenoid commonly consists of an electromagnetically inductive coil, wound around a movable steel or iron slug (the armature), and shaped such that the armature can be moved along the coil center. Any actuator herein may include a solenoid valve, used to actuate a pneumatic valve, where the air is routed to a pneumatic device, or a hydraulic valve, used to control the flow of a hydraulic fluid. In another example, the electromechanical solenoid is used to operate an electrical switch. Similarly, a rotary solenoid may be used, where the solenoid is used to rotate a ratcheting mechanism when power is applied.

[0104] Any actuator herein may be used for effecting or changing magnetic or electrical quantities such as voltage, current, resistance, conductance, reactance, magnetic flux, electrical charge, magnetic field, electric field, electric power, S -matrix, power spectrum, inductance, capacitance, impedance, phase, noise (amplitude or phase), trans-conductance, trans-impedance, and frequency.

[0105] Chatbot. A chatbot is a computer program that processes natural-language input from a user and generates smart and relative responses that are then sent back to the user. Currently, chatbots are powered by rules-driven engines or Artificial Intelligent (Al) engines that interact with users via a text-based interface primarily. These are independent computer programs that can be plugged into any of the multiple messaging platforms that have opened to developers via APIs such as Facebook Messenger, Slack, Skype, Microsoft Teams, and so on. With the advancement of voice technology in recent years, companies such as Google, Apple, and Amazon have debuted artificial intelligent agents for voice. Apple launched Siri, which comes on the iPhone, iPad, and macOS. Google launched Google Home, and Amazon launched Alexa, which are both physically devices for your home or office that can help you with tasks such as ordering a hired car, switching on / off your lights, playing your favorite tunes from Spotify, managing your calendars, and so on. The technology behind chatbots is based on similar technology to voice-based assistants. All voice-based systems have the added complexity of converting the speech to text for any computer application to work with. The processing of the text from a chatbot or a voice-based system is done in the same way, and you will look at the underlying workflow and implement your own system in this book. A chatbot is typically designed to have textual or spoken conversations. Modem chatbots are typically online and use generative artificial intelligence systems that are capable of maintaining a conversation with a user in natural language and simulating the way a human would behave as a conversational partner. Such chatbots often use deep learning and natural language processing, but simpler chatbots have existed for decades.

[0106] Modem chatbots like ChatGPT are often based on large language models called Generative Pre-trained Transformers (GPT). They are based on a deep learning architecture called the transformer, which contains artificial neural networks. They learn how to generate text by being trained on a large text corpus, which provides a solid foundation for the model to perform well on downstream tasks with limited amounts of task- specific data.

[0107] Chatbots have great potential to serve as an alternate source for customer service. Many high-tech banking organizations are looking to integrate automated Al-based solutions such as chatbots into their customer service in order to provide faster and cheaper assistance to their clients who are becoming increasingly comfortable with technology. In particular, chatbots can efficiently conduct a dialogue, usually replacing other communication tools such as email, phone, or SMS. In banking, their major application is related to quick customer service answering common requests, as well as transactional support. Deep learning techniques can be incorporated into chatbot applications to allow them to map conversations between users and customer service agents, especially in social media. Research has shown that methods incorporating deep learning can learn writing styles from a brand and transfer them to another, promoting the brand's image on social media platforms. The history of chatbots, including when they were invented and how they became popular, and how to deploy applications with a chat interface on platforms such as Facebook Messenger, Skype, and so on, which automatically respond to user queries without any human intervention, are described in a book by Rashid Khan and Anik Das, published 2018 [https: / / doi.org / 10.1007 / 978-l-4842-3111-l_l; ISBN-13 (pbk): 978-1-4842-3110-4], entitled: “Build Beter Chatbots - A Complete Guide to Getting Started with Chatbots”, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0108] PESInet. PESInet is an Automatic Prosody Recognition system aiming at classifying Information Units as Statement, Question or Exclamation, by predicting the punctuation mark at the end of each sentence (question, period, exclamation mark), based on the assumption that speech segmentation is best performed using textual features. PESInet scheme is described in a paper by Sonia Cenceschi, Roberto Tedesco, Licia Sbattella, Davide Losio, and Mauro Luchetti, presented October 2019 in the “CLiC-it 2019 Italian Conference on Computational Linguistics” conference [Downloaded from https: / / ceur-ws.org / Vol-2481 / paperl6.pdf], entitled: “PESInet: Automatic Recognition of Italian Statements, Questions, and Exclamations With Neural Networks”, which is incorporated in its entirety for all purposes as if fully set forth herein. PESInet adopts a modular architecture, with a master NN evaluating the results of two independent BLSTM NNs that work on audio and its transcription. PESInet has been trained with our own three-class, balanced corpus composed of about 1.5 million text phrases and 60 000 utterances of recited and spontaneous speech. PESInet reached an accuracy of 80% on three classes, and 91% on two classes (Question vs Non-question). Finally PESInet, compared against human listeners on a two-class test based on a different corpus, reached a better Accuracy (89% for PESInet, against 80% for human listeners).

[0109] CLAP. Contrastive Language- Audio Pretraining (CLAP) is a modelling method that uses an audio encoder trained together with a text decoder to transform both modalities into the same embedded space, in which the audio is close to its caption. This method may be used to model prosody by either finding a text that classifies the audio according to a description of the prosodic labels, or by fine-tuning. Mainstream Audio Analytics models are trained to learn under the paradigm of one class label to many recordings focusing on one task. Learning under such restricted supervision limits the flexibility of models because they require labeled audio for training and can only predict the predefined categories.

[0110] The Contrastive Language-Audio Pretraining (CLAP) approach is proposed to learn audio concepts from natural language supervision, which learns to connect language and audio by using two encoders and a contrastive learning to bring audio and text descriptions into a joint multimodal space, as described in a paper by Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang published 9 June 2022 [arXiv:2206.04769vl [cs.SD] 9 Jun 2022], entitled: “CLAP: LEARNING AUDIO CONCEPTS FROM NATURAL LANGUAGE SUPERVISION’, which is incorporated in its entirety for all purposes as if fully set forth herein. CLAP was trained with 128k audio and text pairs and evaluated it on 16 downstream tasks across 8 domains, such as Sound Event Classification, Music tasks, and Speech-related tasks. Although CLAP was trained with significantly less pairs than similar computer vision models, it establishes SoTA for Zero-Shot performance. Additionally, CLAP was evaluated in a supervised learning setup and achieve SoTA in 5 tasks. Hence, CLAP’s Zero-Shot capability removes the need of training with class labels, enables flexible class prediction at inference time, and generalizes to multiple downstream tasks.

[0111] PENGI. In the domain of audio processing, Transfer Learning has facilitated the rise of Self-Supervised Learning and Zero-Shot Learning techniques. These approaches have led to the development of versatile models capable of tackling a wide array of tasks, while delivering state-of-the-art performance. However, current models inherently lack the capacity to produce the requisite language for open-ended tasks, such as Audio Captioning or Audio Question Answering. Pengi, a novel Audio Language Model that leverages Transfer Learning by framing all audio tasks as text-generation tasks, was introduced in 37th Conference on Neural Information Processing Systems (NeurlPS 2023) in a paper by Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang, entitled: “Pengi: An Audio Language Model for Audio Tasks”, which is incorporated in its entirety for all purposes as if fully set forth herein. PENGI takes as input, an audio recording, and text, and generates free-form text as output. The input audio is represented as a sequence of continuous embeddings by an audio encoder. A text encoder does the same for the corresponding text input. Both sequences are combined as a prefix to prompt a pre-trained frozen language model. The unified architecture of Pengi enables open- ended tasks and close-ended tasks without any additional fine-tuning or task-specific extensions. When evaluated on 21 downstream tasks, the Pengi approach yields state-of-the-art performance in several of them, and the results show that connecting language models with audio models is a major step towards general-purpose audio understanding

[0112] MLLM. Multimodal Large Language Models (MLLMs), that may use new models such as chatGPT4 and Gemini, are able of receive audio as input, a detection of prosody may be achieved by prompt engineering. An example of such an attempt uses Whisper in an In-Context- Leaming (ICL) setting. A Speech-based In-Context Learning (SICL) approach that is proposed for test-time adaptation, is presented in a paper by Siyin Wang, Chao-Han Yang, Ji Wu, and Chao Zhang, dated 20 Mar 2024 [arXiv:2309.07081v2 [eess.AS] 20 Mar 2024; https: / / doi.org / 10.48550 / arXiv.2309.07081], entitled: “CAN WHISPER PERFORM SPEECHBASED IN-CONTEXT LEARNING?” , that investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI, which is incorporated in its entirety for all purposes as if fully set forth herein. The SICL can reduce the word error rates (WERs) with only a small number of labelled speech samples without gradient descent. Language-level adaptation experiments using Chinese dialects showed that when applying SICL to isolated word ASR, consistent and considerable relative WER reductions can be achieved using Whisper models of any size on two dialects, which is on average 32.3%. A k- nearest-neighbours-based in-context example selection technique can be applied to further improve the efficiency of SICL, which can increase the average relative WER reduction to 36.4%. The findings are verified using speaker adaptation or continuous speech recognition tasks, and both achieved considerable relative WER reductions. Detailed quantitative analyses are also provided to shed light on SICL’s adaptability to phonological variances and dialectspecific lexical nuances. SICL using Whisper is expected to do worse than fine-tuning, since LLMs only benefit from ICL for models of 50B parameters and more, whereas whisper relies on only ~2B parameters for audio encoding and text decoding combined. In practice, this method indeed results in inferior performance for the segmentation task.

[0113] Techniques for training and applying prosody models for speech synthesis are provided in U.S. Patent No. 9,070,365 to Stephens, Jr. entitled: “Training and applying prosody models”, which is incorporated in its entirety for all purposes as if fully set forth herein. A speech recognition engine processes audible speech to produce text annotated with prosody information. A prosody model is trained with this annotated text. After initial training, the model is applied during speech synthesis to generate speech with non-standard prosody from input text. Multiple prosody models can be used to represent different prosody styles.

[0114] A technology by which synthesized speech generated from text is evaluated against a prosody model (trained offline) to determine whether the speech will sound unnatural, is disclosed in U.S. Patent No. 8,583,438 to Zhao et al. entitled: “Unnatural prosody detection in speech synthesis”, which is incorporated in its entirety for all purposes as if fully set forth herein. The speech is regenerated with modified data. The evaluation and regeneration may be iterative until deemed natural sounding. For example, the text is built into a lattice that is then (e.g., Viterbi) searched to find the best path. The sections (e.g., units) of data on the path are evaluated via a prosody model. If the evaluation deems a section to correspond to unnatural prosody, that section is replaced, e.g., by modifying / pruning the lattice and re-performing the search. Replacement may be iterative until all sections pass the evaluation. Unnatural prosody detection may be biased such that during evaluation, unnatural prosody is falsely detected at a higher rate relative to a rate at which unnatural prosody is missed.

[0115] A voice processing apparatus for recognizing an input voice on the basis of a prosody characteristic of said voice is disclosed in U.S. Patent No. 7,979,270 to Yamada entitled: “Speech recognition apparatus and method”, which is incorporated in its entirety for all purposes as if fully set forth herein. Said voice processing apparatus includes: voice acquisition means for acquiring said input voice; acoustic analysis means for finding a relative pitch change on the basis of a frequency-direction difference between a first frequency characteristic seen at each frame time of said input voice acquired by said voice acquisition means and a second frequency characteristic determined in advance; and prosody recognition means for carrying out a prosody recognition process on the basis of said relative pitch change found by said acoustic analysis means in order to produce a result of said prosody recognition process.

[0116] A method and system for monitoring a conversation between a pair of speakers for detecting an emotion of at least one of the speakers are provided in U.S. Patent No. 6,151,571 to Pertrushin entitled: “System, method and article of manufacture for detecting emotion in voice signal through analysis of a plurality of voice signal parameters”, which is incorporated in its entirety for all purposes as if fully set forth herein. First, a voice signal is received after which a particular feature is extracted from the voice signal. Next, an emotion associated with the voice signal is determined based on the extracted feature. The emotion is screened and feedback is provided only if the emotion is determined to be a negative emotion selected from the group of negative emotions consisting of anger, sadness, and fear. Such determined negative emotion is then outputted to a third party during the conversation.

[0117] The classification of speech according to emotional content that employs acoustic measures in addition to pitch as classification input is disclosed in U.S. Patent No. 6,173,260 to Slaney entitled: “System and method for automatic classification of speech based upon affective content”, which is incorporated in its entirety for all purposes as if fully set forth herein. In one embodiment, two different kinds of features in a speech signal are analyzed for classification purposes. One set of features is based on pitch information that is obtained from a speech signal, and the other set of features is based on changes in the spectral shape of the speech signal over time. This latter feature is used to distinguish long, smoothly varying sounds from quickly changing sounds, which may indicate the emotional state of the speaker. These changes are determined by means of a low-dimensional representation of the speech signal, such as MFCC or LPC. Additional features of the speech signal, such as energy, can also be employed for classification purposes. Different variations of pitch and spectral shape features can be measured and analyzed, to assist in the classification of individual utterances. In one implementation, the features are measured individually for each of the first, middle, and last thirds of an utterance, as well as for the utterance as a whole, to generate multiple sets of data for each utterance.

[0118] A method and an apparatus for extracting a prosodic feature of a speech signal are disclosed in U.S. Patent No. 8,566,092 to Liu et al. entitled: “Method and apparatus for extracting prosodic feature of speech signal”, which is incorporated in its entirety for all purposes as if fully set forth herein. The method includes: dividing the speech signal into speech frames; transforming the speech frames from time domain to frequency domain; and extracting respective prosodic features for different frequency ranges. According to the above technical solution of the present invention, it is possible to effectively extract the prosodic feature which can combine with a traditional acoustics feature without any obstacle.

[0119] A system for carrying out voice pattern recognition and a method for achieving the same is disclosed in U.S. Patent No. 9,754,580 to WEISSBERG et al. entitled: “System and method for extracting and using prosody features” , which is incorporated in its entirety for all purposes as if fully set forth herein. The system includes an arrangement for acquiring an input voice, a signal processing library for extracting acoustic and prosodic features of the acquired voice, a database for storing a recognition dictionary, at least one instance of a prosody detector for carrying out a prosody detection process on extracted respective prosodic features, communicating with an end-user application for applying control thereto.

[0120] Systems and methods for scoring speech are described in U.S. Patent No. 9,087,519 to Zechner et al. entitled: “Computer -implemented systems and methods for evaluating prosodic features of speech”, which is incorporated in its entirety for all purposes as if fully set forth herein. A speech sample is received, where the speech sample is associated with a script. The speech sample is aligned with the script. An event recognition metric of the speech sample is extracted, and locations of prosodic events are detected in the speech sample based on the event recognition metric. The locations of the detected prosodic events are compared with the locations of model prosodic events, where the locations of model prosodic events identify the expected locations of prosodic events of a fluent, native speaker speaking the script. A prosodic event metric is calculated based on the comparison, and the speech sample is scored using a scoring model based upon the prosodic event metric.

[0121] While speech recognition Word Error Rate (WER) has reached human parity for English, continuous speech recognition scenarios such as voice typing and meeting transcriptions still suffer from segmentation and punctuation problems, resulting from irregular pausing patterns or slow speakers. Transformer sequence tagging models are effective at capturing long bi-directional context, which is crucial for automatic punctuation. Automatic Speech Recognition (ASR) production systems, however, are constrained by real-time requirements, making it hard to incorporate the right context when making punctuation decisions. The context within the segments produced by ASR decoders can be helpful but limiting in overall punctuation performance for a continuous speech session. A paper by Piyush Behre, Sharman Tan, Padma Varadharajan, and Shuangyu Chang (of Microsoft Corporation), published International Journal on Natural Language Computing (UNLC) Vol.11, No.6, December 2022 [DOI: 10.5121 / ijnlc.2022.11601], entitled: "STREAMING PUNCTUATION: A NOVEL PUNCTUATION TECHNIQUE LEVERAGING BIDIRECTIONAL CONTEXT FOR CONTINUOUS SPEECH RECOGNITION' , is incorporated in its entirety for all purposes as if fully set forth herein. The paper proposes a streaming approach for punctuation or repunctuation of ASR output using dynamic decoding windows and measures its impact on punctuation and segmentation accuracy across scenarios. The new system tackles oversegmentation issues, improving segmentation F0.5-score by 13.9%. Streaming punctuation achieves an average BLEU score improvement of 0.66 for the downstream task of Machine Translation (MT).

[0122] Table 1 below demonstrates the acoustic parameters of tone sequences significantly contributing to the variance of attributions of emotional states. It was further observed that regarding timbre, tone sequences with fewer upper harmonics were judged as significantly more pleasant, happy, and bored, while tone sequences with more upper harmonics were judged as significantly more active, potent, angry, disgusted, and fearful. Regarding tonality, it was observed that the major mode is indicative of pleasantness and happiness, while the minor mode suggests disgust and anger. Similarly, rhythmic indicates more active, fearful, and surprised, while nonrhythmic suggests boredom. The intonation pattern may correspond to emotions as described in the table, including pleasantness, activity, potency, anger, boredom, disgust, fear, happiness, sadness, and surprise.

[0123] Table 1

[0124] | Scale ^effects) listed in order of predictive strength i i Fast tempo, few harmonics, large pitch variation, sharp envelope, low pitch level, pitch contour ii i Pleasantness down, small amplitude variation (salient configuration: large pitch variation plus pitch contour ii llp'

[0125] L . . Tast tempo, high pitch level, many harmonics, large pitch variation, sharp envelope, small $ iyi amplitude variation ii Many harmonics, fast tempo, high pitch level, round envelope, pitch contour up (salient Potency configurations: large amplitude variation plus high pitch level, high pitch level plus many I harmonics)

[0126] Many harmonics, fast tempo, high pitch level, small pitch variation, pitch contours up (salient ii Anger Iconfiguration: small pitch variation plus pitch contour up) ii

[0127] I Slow tempo, low pitch level, few harmonics, pitch contour down, round envelope, small pitch $ Boredom I variation ii

[0128] I Many harmonics, small pitch variation, round envelope, slow tempo (salient configuration: small i Disgust

[0129] I pitch variation plus pitch contour up) $ i Pitch contour up, fast sequence, many harmonics, high pitch level, round envelope, small pitch ii

[0130] Fear variation (salient configurations: small pitch variation plus pitch contour up, fast tempo plus ii i many harmonics) ii Fast tempo, large pitch variation, sharp envelope, few harmonics, moderate amplitude variation i; Happiness (salient configurations: large pitch variation plus pitch contour up, fast tempo plus few i; I harmonics) $

[0131] The term 'utterance' herein refers to the smallest unit of speech, typically it is a continuous piece of speech that begins and ends with a clear pause. In general, it could be anything from "Ugh!" to a full sentence.

[0132] In one example, part of, or all of, the steps, methods, or flow charts described herein are executed (independently or in cooperation) by a client device, or any device such as the device 35 shown in FIG. 3. Alternatively or in addition, part of, or all of, the steps, methods, or flow charts described herein are executed (independently or in cooperation) by a server device, such as server 24 shown as part on the arrangement 30 shown in FIG. 3. In one example, a client device (such as the device 35) and a server (such as the server 24) cooperatively perform part of, or all of, the steps, methods, or flow charts described herein. For example, lower computing power processor 26 may be used in the device 35, since the heavy or resourceful computations are performed at a remote server. Such a scheme may obviate the need for expensive and resourceful devices. In another example, memory resources may be saved at the client device by using data stored on a server. A storage 31c in the device 35 may store the Instructions 37a and the Operating System 37b. The device 35 may mainly be used for interfacing with the user 36, while the major storing and processing resources and activities are provided by the server 24.

[0133] The output component 34 may include a color display for displaying screen elements or for organizing on-screen items and controls for data entry. Further, the device may support the display of split-screen views. The input component 38 may include dedicated hard controls for frequently used / accessed functions (e.g., repeat system message). Many systems use re- configurable keys / buttons whose functions change depending on the application. For example, a switch may be used to activate the voice recognition system and it may increase system reliability. The input component 38 and the output component 34 may further cooperate to provide both auditory and visual feedback to confirm driver inputs and availability of the speech command. Further, a strategy to alert drivers through auditory tones / beeps in advance of the presentation of information, and / or changes in display status, may be used. This may limit the need for drivers to continuously monitor the system, or repeat system messages. In one example, the input component 38 may comprise a microphone 33.

[0134] Any microphone herein may be an electro-acoustic sensor for measuring, sensing, or detecting sound, such as a microphone. Typically, microphones are based on converting audible or inaudible (or both) incident sound to an electrical signal by measuring the vibration of a diaphragm or a ribbon. The microphone may be a condenser microphone, an electret microphone, a dynamic microphone, a ribbon microphone, a carbon microphone, or a piezoelectric microphone.

[0135] The device 35 may serve as a client device and may access data, such as retrieving data from, or sending data to, the server 24 over the Internet 25, such as via an ISP. The communication with the server 24 may be via a wireless network 39, by using the antenna 29 and the wireless transceiver 28 in the device 35.

[0136] A diagrammatic representation of a machine in the example forms of the computing device 35 within which a set of instructions, for causing the machine to perform any one or more of the methods discussed herein, may be executed. An example of the device 35 that may be used with any of the steps, methods, or flow-charts herein is schematically described as part of an arrangement 30 shown in FIG. 3. The components in the device 35 communicate over a bus 32. The computing device 35 may include a mobile phone, a smartphone, a netbook computer, a rackmount server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, etc., within which a set of instructions, for causing the machine to perform any one or more of the methods discussed herein, may be executed. In alternative embodiments, the machine may be connected (e.g., networked) to other machines in a LAN, an intranet, an extranet, or the Internet. The machine may operate in the capacity of a server machine in client-server network environment. The machine may be a Personal Computer (PC), a Set-Top Box (STB), a server, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while only a single machine is illustrated, the term “machine” may also include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein.

[0137] The device 35 may also include an interface bus for facilitating communication from various interface devices (for example, one or more output components 34, one or more peripheral interfaces, and one or more communication components such as the wireless transceiver 28) to the basic configuration via the bus / interface controller that controls the bus 32. Some of the examples output components include a graphics processing unit and an audio processing unit, which may be configured to communicate to various external devices such as a display or speakers via one or more A / V ports. One or more example peripheral interfaces may include a serial interface controller or a parallel interface controller, which may be configured to communicate with external devices such as input components (for example, keyboard, mouse, pen, voice input device, touch input device, etc.) or other peripheral output devices (for example, printer, scanner, etc.) via one or more I / O ports.

[0138] The device 35 may be part of, may include, or may be integrated with, a general-purpose computing device, arranged in accordance with at least some embodiments described herein. In an example basic configuration, the device 35 may include one or more processors 26 and one or more memories or any other computer-readable media. A dedicated memory bus may be used to communicate between the processor 26 and the device memories, such as the ROM 31b, the main memory 31a, and the storage 31c. Depending on the desired configuration, the processor 26 may be of any type, including but not limited to a microprocessor (pP), a microcontroller (pC), a digital signal processor (DSP), or any combination thereof. The processor 26 may include one or more levels of caching, such as a cache memory, a processor core, and registers. The example processor core may include an arithmetic logic unit (ALU), a floating-point unit (FPU), a digital signal processing core (DSP Core), or any combination thereof. An example memory controller may also be used with the processor 26, or in some implementations, the memory controller may be an internal part of the processor 26.

[0139] Depending on the desired configuration, the device memories may be of any type including but not limited to volatile memory (such as RAM), non-volatile memory (such as ROM, flash memory, etc.), or any combination thereof. The storage 31c may correspond to the storage device 31a, and may be part of, may comprise, or may be integrated with the ROM 31b and the main memory 31a. The storage 31c may include an operating system 37b, instruction set 37a that may include steps or part of, or whole of, the flow-charts described herein. The storage 31c may further include a control module, and program data, which may include path data. Any of the memories or storages of the device 35 may include read-only memory (ROM), such as ROM 31b, flash memory, Dynamic Random Access Memory (DRAM) such as synchronous DRAM (SDRAM), a static memory (e.g., flash memory, static random-access-memory (SRAM)) and a data storage device, which communicate with each other via the bus 32.

[0140] The device 35 may have additional features or functionality, and additional interfaces to facilitate communications between the basic configuration shown in FIG. 3 and any desired devices and interfaces. For example, a bus / interface controller may be used to facilitate communications between the basic configuration and one or more data storage devices via a storage interface bus. The data storage devices may be one or more removable storage devices, one or more non-removable storage devices, or a combination thereof. Examples of removable storage and the non-removable storage devices include magnetic disk devices such as flexible disk drives and hard-disk drives (HDDs), optical disk drives such as compact disk (CD) drives or Digital- Versatile-Disk (DVD) drives, Solid-State Drives (SSDs), and tape drives to name a few. For example, computer storage media may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer readable instructions, data structures, program modules, or other data.

[0141] The device 35 may receive inputs from a user 36 via an input component 38. In one example, the input component 38 may be used for receiving instructions from the user 36. The device 35 notifies or outputs information to the user 36 using an output component 34. In one example, the output component 34 may be used for displaying guidance to the user 36.

[0142] The interface with the user 36 may be based on the input component 38 and the output component 34. For example, receiving input (visually or acoustically) from the user 36 via the input component 38. Similarly, outputting data (visually or acoustically) to the user 36 via the output component 34. The input component 38 may be a piece of computer hardware equipment used to provide data and control signals to an information processing system such as a computer or information appliance. Such input component 38 may be an integrated or a peripheral input device (e.g., hard / soft keyboard, mouse, resistive or capacitive touch display, etc.). Examples of input components include keyboards, mouse, scanners, digital cameras, and joysticks. Input components 38 can be categorized based on the modality of input (e.g., mechanical motion, audio, visual, etc.), whether the input is discrete (e.g., pressing of a key) or continuous (e.g., a mouse's position, though digitized into a discrete quantity, is fast enough to be considered continuous), the number of degrees of freedom involved (e.g., two-dimensional traditional mice, or three-dimensional navigators designed for CAD applications). Pointing devices (such as ‘computer mouse’), which are input components used to specify a position in space, can further be classified according to whether the input is direct or indirect. With direct input, the input space coincides with the display space, i.e., pointing is done in the space where visual feedback or the pointer appears. Touchscreens and light pens involve direct input. Examples involving indirect input include the mouse and trackball, and whether the positional information is absolute (e.g., on a touch screen) or relative (e.g., with a mouse that can be lifted and repositioned). Direct input is almost necessarily absolute, but indirect input may be either absolute or relative. For example, digitizing graphics tablets that do not have an embedded screen involve indirect input and sense absolute positions and are often run in an absolute input mode, but they may also be set up to simulate a relative input mode like that of a touchpad, where the stylus or puck can be lifted and repositioned.

[0143] In the case of wireless networking, the wireless network 39 may use any type of modulation, such as Amplitude Modulation (AM), Frequency Modulation (FM), or Phase Modulation (PM). Further, the wireless network 39 may be a control network (such as ZigBee or Z-Wave), a home network, a WPAN (Wireless Personal Area Network), a WEAN (Wireless Focal Area Network), a WWAN (Wireless Wide Area Network), or a cellular network. An example of a Bluetooth-based wireless controller that may be included in a wireless transceiver is SPBT2632C1A Bluetooth module available from STMicroelectronics NV and described in the data sheet DoclD022930 Rev. 6 dated April 2015 entitled: “SPBT2632C1A - Bluetooth® technology class-1 module”, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0144] Some embodiments may be used in conjunction with one or more types of wireless communication signals and / or systems, for example, Radio Frequency (RF), Infra-Red (IR), Frequency-Division Multiplexing (FDM), Orthogonal FDM (OFDM), Time-Division Multiplexing (TDM), Time-Division Multiple Access (TDMA), Extended TDMA (E-TDMA), General Packet Radio Service (GPRS), extended GPRS, Code-Division Multiple Access (CDMA), Wideband CDMA (WCDMA), CDMA 2000, single-carrier CDMA, multi-carrier CDMA, Multi-Carrier Modulation (MDM), Discrete Multi-Tone (DMT), Bluetooth (RTM), Global Positioning System (GPS), Wi-Fi, Wi-Max, ZigBee (TM), Ultra-Wideband (UWB), Global System for Mobile communication (GSM), 2G, 2.5G, 3G, 3.5G, Enhanced Data rates for GSM Evolution (EDGE), or the like. Further, wireless communication may be based on, or may be compatible with, wireless technologies that are described in Chapter 20: "Wireless Technologies" of the publication number 1-587005-001-3 by Cisco Systems, Inc. (7 / 99) entitled: "Internetworking Technologies Handbook" , which is incorporated in its entirety for all purposes as if fully set forth herein. Alternatively or in addition, the networking or the communication of the wireless- capable device 35 with the server 24 over the wireless network 39 may be using, may be according to, may be compatible with, or may be based on, Near Field Communication (NFC) using passive or active communication mode, and may use the 13.56 MHz frequency band, and data rate may be 106Kb / s, 212Kb / s, or 424 Kb / s, and the modulation may be Amplitude-Shift- Keying (ASK), and may be according to, may be compatible with, or based on, ISO / IEC 18092, ECMA-340, ISO / IEC 21481, or ECMA-352. In such a case, the wireless transceiver 28 may be an NFC transceiver and the respective antenna 29 may be an NFC antenna.

[0145] Alternatively or in addition, the networking or the communication with the wireless- capable device 35 with the server 24 over the wireless network 39 may be using, may be according to, may be compatible with, or may be based on, a Wireless Personal Area Network (WPAN) that may be according to, may be compatible with, or based on, Bluetooth™ or IEEE 802.15.1-2005 standards, and the wireless transceiver 28 may be a WPAN modem, and the respective antenna 29 may be a WPAN antenna. The WPAN may be a wireless control network according to, may be compatible with, or based on, ZigBee™ or Z-Wave™ standards, such as IEEE 802.15.4-2003.

[0146] Alternatively or in addition, the networking or the communication of the wireless- capable device 35 with the server 24 over the wireless network 39 may be using, may be according to, may be compatible with, or may be based on, a Wireless Local Area Network (WLAN) that may be according to, may be compatible with, or based on, IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.1 In, or IEEE 802.1 lac standards, and the wireless transceiver 28 may be a WLAN modem, and the respective antenna 29 may be a WLAN antenna.

[0147] Alternatively or in addition, the networking or the communication of the wireless- capable device 35 with the server 24 over the wireless network 39 may be using, may be according to, may be compatible with, or may be based on, a wireless broadband network or a Wireless Wide Area Network (WWAN), and the wireless transceiver 28 may be a WWAN modem, and the respective antenna 29 may be a WWAN antenna. The WWAN may be a WiMAX network such as according to, may be compatible with, or based on, IEEE 802.16- 2009, and the wireless transceiver 28 may be a WiMAX modem, and the respective antenna 29 may be a WiMAX antenna. Alternatively or in addition, the WWAN may be a cellular telephone network, the wireless transceiver 28 may be a cellular modem, and the respective antenna 29 may be a cellular antenna. The WWAN may be a Third Generation (3G) network and may use UMTS W-CDMA, UMTS HSPA, UMTS TDD, CDMA2000 IxRTT, CDMA2000 EV-DO, or GSM EDGE-Evolution. The cellular telephone network may be a Fourth Generation (4G) network and may use HSPA+, Mobile WiMAX, LTE, LTE-Advanced, MBWA, or may be based on, or may be compatible with, IEEE 802.20-2008. Alternatively or in addition, the WWAN may be a satellite network, the wireless transceiver 28 may be a satellite modem, and the respective antenna 29 may be a satellite antenna.

[0148] Alternatively or in addition, the networking or the communication of the wireless- capable device 35 with the server 24 over the wireless network 39 may be using, may be according to, may be compatible with, or may be based on, a licensed or an unlicensed radio frequency band, such as the Industrial, Scientific, and Medical (ISM) radio band. For example, an unlicensed radio frequency band may be used that may be about 60 GHz, may be based on beamforming, and may support a data rate of above 7Gb / s, such as according to, may be compatible with, or based on, WiGig™, IEEE 802. Had, WirelessHD™ or IEEE 802.15.3c- 2009, and may be operative to carry uncompressed video data, and may be according to, may be compatible with, or based on, WHDI™. Alternatively or in addition, the wireless network may use a white space spectrum that may be an analog television channel consisting of a 6MHz, 7MHz, or 8MHz frequency band, and allocated in the 54-806 MHz band. The wireless network may be operative for channel bonding, and may use two or more analog television channels, and may be based on Wireless Regional Area Network (WRAN) standard using OFDMA modulation. Further, the wireless communication may be based on geographically-based cognitive radio, and may be according to, may be compatible with, or may be based on, IEEE 802.22 or IEEE 802.11af standards.

[0149] Display. The Output Component 34 may include a display for a presentation of visual data or information, commonly on a screen. A display typically consists of an array of light emitters (typically in a matrix form) and commonly provides a visual depiction of a single, integrated, or organized set of information, such as text, graphics, image or video. A display may be a monochrome (a.k.a. black-and-white) type, which typically displays two colors, one for the background and one for the foreground. A display may be a gray-scale type, which is capable of displaying different shades of gray, or may be a color type, capable of displaying multiple colors, anywhere from 16 to over many millions different colors, and may be based on Red, Green, and Blue (RGB) separate signals. A video display is designed for presenting video content. The screen is the actual location where the information is actually optically visualized by humans. The screen may be an integral part of the display. Alternatively or in addition, the display may be an image or video projector, that projects an image (or a video consisting of moving images) onto a screen surface, which is a separate component and is not mechanically enclosed with the display housing. Most projectors create an image by shining a light through a small transparent image, but some newer types of projectors can project the image directly, by using lasers. A projector may be based on an Eidophor, Liquid Crystal on Silicon (LCoS or LCOS), or LCD, or may use Digital Light Processing (DLP™) technology, and may further be MEMS based. A virtual retinal display, or retinal projector, is a projector that projects an image directly on the retina instead of using an external projection screen. Common display resolutions used today include SVGA (800x600 pixels), XGA (1024x768 pixels), 720p (1280x720 pixels), and 1080p (1920x1080 pixels). Standard-Definition (SD) standards, such as those used in SD Television (SDTV), are referred to as 576i, derived from the European-developed PAL and SEC AM systems with 576 interlaced lines of resolution; and 480i, based on the American National Television System Committee (ANTSC) NTSC system. High-Definition (HD) video refers to any video system of higher resolution than standard-definition (SD) video, and most commonly involves display resolutions of 1,280x720 pixels (720p) or 1,920x1,080 pixels (1080i / 1080p). A display may be a 3D (3-Dimensions) display, which is a display device capable of conveying a stereoscopic perception of 3-D depth to the viewer. The basic technique is to present offset images that are displayed separately to the left and right eye. Both of these 2- D offset images are then combined in the brain to give the perception of 3-D depth. The display may present the information as scrolling, static, bold, or flashing.

[0150] Sounder. The Output Component 34 may include a sounder that converts electrical energy to sound waves transmitted through the air, an elastic solid material, or a liquid, usually by means of a vibrating or moving ribbon or diaphragm. The sound may be audible or inaudible (or both) and may be omnidirectional, unidirectional, bidirectional, or provide other directionality or polar patterns. The sounder may be an electromagnetic loudspeaker, a piezoelectric speaker, an electrostatic loudspeaker (ESL), a ribbon or planar magnetic loudspeaker, or a bending wave loudspeaker. Further, the sounder may be an electromechanical type, such as an electric bell, a buzzer (or beeper), a chime, a whistle, or a ringer, and may be either an electromechanical or ceramic -based piezoelectric sounder. The sounder may emit a single or multiple tones and can be in continuous or intermittent operation.

[0151] The sounder may be operative for converting an electrical energy to omnidirectional, unidirectional, or bidirectional pattern of emitted, audible or inaudible, sound waves. Any sounder herein may be audible and may be an electromagnetic loudspeaker, a piezoelectric speaker, an electrostatic loudspeaker (ESL), a ribbon or planar magnetic loudspeaker, or a bending wave loudspeaker. Any sounder herein may be operative to emit a single or multiple tones, or may be operative to continuous or intermittent operation. Any sounder herein may be an electromechanical or a ceramic-based, and may be an electric bell, a buzzer (or beeper), a chime, a whistle, or a ringer. Any sound herein may be audible, and any sounder herein may be a loudspeaker, and any apparatus or device herein may be operative to store and play one or more digital audio content files.

[0152] Headphones. Headphones are a pair of small loudspeakers that are designed to be worn on or around the head over a user's ears. They are electroacoustic transducers, which convert an electrical signal to a corresponding sound in the user's ear, and are designed to allow a single user to listen to an audio source privately, in contrast to a loudspeaker, which emits sound into the open air, allowing anyone nearby to listen. Headphones are also known as ear-speakers or earphones. Circumaural and supra-aural headphones use a band over the top of the head to hold the speakers in place. The other type, known as earbuds or earphones, consists of individual units that plug into the user's ear canal. In the context of telecommunication, a headset is a combination of headphones and a microphone. Generally, headphone form factors can be divided into four separate categories: circumaural, supra-aural, earbud, and in-ear.

[0153] Circumaural headphones typically have large pads that surround the outer ear. Circumaural headphones (sometimes called full size headphones) have circular or ellipsoid earpads that encompass the ears. Since such headphones completely surround the ear, circumaural headphones can be designed to fully seal against the head to attenuate external noise. Supra-aural headphones typically have pads that press against the ears, rather than around them. This type of headphone generally tends to be smaller and lighter than circumaural headphones, resulting in less attenuation of outside noise. Supra-aural headphones can also lead to discomfort due to the pressure on the ear as compared to circumaural headphones that sit around the ear. Comfort may vary due to the earcup material.

[0154] Both circumaural and supra-aural headphones can be further differentiated by the type of earcups: Open-back headphones have the back of the earcups open. While the opening leaks sound out of the headphone, and also lets more ambient sounds into the headphone, it gives a more natural or speaker-like sound and a more spacious "soundstage" - the perception of distance from the source. Closed-back (or sealed) styles have the back of the earcups closed, so they usually block some of the ambient noise, but have a smaller soundstage, giving the wearer a perception that the sound is coming from within their head. Closed-back headphones tend to be able to produce stronger low frequencies than open-back headphones. Semi-open headphones have a design that can be considered as a compromise between open-back headphones and closed-back headphones. This may imply that the result combines all the positive properties of both designs. Where the open-back approach has hardly any measure to block sound at the outer side of the diaphragm and the closed-back approach really has a closed chamber at the outer side of the diaphragm, a semi-open headphone can have a chamber to partially block sound while letting some sound through via openings or vents.

[0155] Earphones (popularly called "earbuds" in recent years) are very small headphones that are fitted directly in the outer ear, facing but not inserted in the ear canal. Earphones are portable and convenient, but many people consider them uncomfortable and prone to falling out. They provide hardly any acoustic isolation and leave room for ambient noise to seep in; users may turn up the volume dangerously high to compensate, at the risk of causing hearing loss. On the other hand, they let the user be better aware of their surroundings. They are sold at times with foam pads for comfort.

[0156] In-ear headphones, also known as In-Ear Monitors (IEMs) or canalphones, are small headphones with similar portability to earbuds that are inserted in the ear canal itself, thus providing isolation from outside noise. IEMs are higher-quality in-ear headphones and are used by audio engineers and musicians as well as audiophiles. Because in-ear headphones engage the ear canal, they can be less prone to falling out, and they block out environmental noise. Lack of sound from the environment can be a problem when sound is a necessary cue for safety or other reasons, such as when walking, driving, or riding near or in vehicular traffic. Generic or customfitting ear canal plugs are made from silicone rubber, elastomer, or foam. Custom in-ear headphones use castings of the ear canal to create custom-molded plugs that provide added comfort and noise isolation.

[0157] Smartphone. A mobile phone (also known as a cellular phone, cell phone, smartphone, or hand phone) is a device which can make and receive telephone calls over a radio link whilst moving around a wide geographic area, by connecting to a cellular network provided by a mobile network operator. The calls are to and from the public telephone network, which includes other mobiles and fixed-line phones across the world. The Smartphones are typically hand-held and may combine the functions of a personal digital assistant (PDA), and may serve as portable media players and camera phones with high-resolution touch- screens, web browsers that can access, and properly display, standard web pages rather than just mobile-optimized sites, GPS navigation, Wi-Fi and mobile broadband access. In addition to telephony, smartphones may support a wide variety of other services such as text messaging, MMS, email, Internet access, short-range wireless communications (infrared, Bluetooth), business applications, gaming, and photography.

[0158] An example of a contemporary smartphone is model iPhone 6 available from Apple Inc., headquartered in Cupertino, California, U.S.A, and described in iPhone 6 technical specification (retrieved 10 / 2015 from www.apple.com / iphone-6 / specs / ), and in a User Guide dated 2015 (019-00155 / 2015-06) by Apple Inc. entitled: “iPhone User Guide For iOS 8.4 Software”, which are both incorporated in their entirety for all purposes as if fully set forth herein. Another example of a smartphone is Samsung Galaxy S6 available from Samsung Electronics headquartered in Suwon, South-Korea, described in the user manual numbered English (EU), 03 / 2015 (Rev. 1.0) entitled: “SM-G925F SM-G925FQ SM-G925I User Manual” and having features and specification described in “Galaxy S6 Edge - Technical Specification” (retrieved 10 / 2015 from www.samsung.com / us / explore / galaxy-s-6-features-and-specs), which are both incorporated in their entirety for all purposes as if fully set forth herein.

[0159] A mobile operating system (also referred to as mobile OS), is an operating system that operates a smartphone, tablet, PDA, or another mobile device. Modem mobile operating systems combine the features of a personal computer operating system with other features, including a touchscreen, cellular, Bluetooth, Wi-Fi, GPS mobile navigation, camera, video camera, speech recognition, voice recorder, music player, near field communication and infrared blaster. Popular mobile OSs are Android, Symbian, Apple iOS, BlackBerry, MeeGo, Windows Phone, and Bada. Mobile devices with mobile communications capabilities (e.g. smartphones) typically contain two mobile operating systems - a main user-facing software platform is supplemented by a second low-level proprietary real-time operating system that operates the radio and other hardware.

[0160] Android is an open-source and Linux-based mobile operating system (OS) based on the Linux kernel that is currently offered by Google. With a user interface based on direct manipulation, Android is designed primarily for touchscreen mobile devices such as smartphones and tablet computers, with specialized user interfaces for televisions (Android TV), cars (Android Auto), and wrist watches (Android Wear). The OS uses touch inputs that loosely correspond to real-world actions, such as swiping, tapping, pinching, and reverse pinching to manipulate on-screen objects and a virtual keyboard. Despite being primarily designed for touchscreen input, it also has been used in game consoles, digital cameras, and other electronics. The response to user input is designed to be immediate and provides a fluid touch interface, often using the vibration capabilities of the device to provide haptic feedback to the user. Internal hardware such as accelerometers, gyroscopes, and proximity sensors are used by some applications to respond to additional user actions, for example adjusting the screen from portrait to landscape depending on how the device is oriented, or allowing the user to steer a vehicle in a racing game by rotating the device by simulating control of a steering wheel. Android devices boot to the homescreen, the primary navigation and information point on the device, which is similar to the desktop found on PCs. Android homescreens are typically made up of app icons and widgets; app icons launch the associated app, whereas widgets display live, auto-updating content such as the weather forecast, the user's email inbox, or a news ticker directly on the homescreen. A homescreen may be made up of several pages that the user can swipe back and forth between, though Android's homescreen interface is heavily customizable, allowing the user to adjust the look and feel of the device to their tastes. Third-party apps available on Google Play and other app stores can extensively re-theme the homescreen and even mimic the look of other operating systems, such as Windows Phone. The Android OS is described in a publication entitled: “Android Tutorial”, downloaded from tutorialspoint.com on July 2014, which is incorporated in its entirety for all purposes as if fully set forth herein. iOS (previously iPhone OS) from Apple Inc. (headquartered in Cupertino, California, U.S.A.) is a mobile operating system distributed exclusively for Apple hardware. The user interface of the iOS is based on the concept of direct manipulation, using multi-touch gestures. Interface control elements consist of sliders, switches, and buttons. Interaction with the OS includes gestures such as swipe, tap, pinch, and reverse pinch, all of which have specific definitions within the context of the iOS operating system and its multi-touch interface. Internal accelerometers are used by some applications to respond to shaking the device (one common result is the undo command) or rotating it in three dimensions (one common result is switching from portrait to landscape mode). The iOS OS is described in a publication entitled: “IOS Tutorial”, downloaded from tutorialspoint.com on July 2014, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0161] Antenna. An antenna (plural antennae or antennas), or aerial, is an electrical device that converts electric power into radio waves, and vice versa, and is usually used with a radio transmitter or radio receiver. In transmission, a radio transmitter supplies an electric current oscillating at radio frequency (i.e. a high frequency Alternating Current (AC)) to the antenna's terminals, and the antenna radiates the energy from the current as electromagnetic waves (radio waves). In reception, an antenna intercepts some of the power of an electromagnetic wave in order to produce a tiny voltage at its terminals that is applied to a receiver to be amplified.

[0162] Typically an antenna consists of an arrangement of metallic conductors (elements), electrically connected (often through a transmission line) to the receiver or transmitter. An oscillating current of electrons forced through the antenna by a transmitter will create an oscillating magnetic field around the antenna elements, while the charge of the electrons also creates an oscillating electric field along the elements. These time-varying fields radiate away from the antenna into space as a moving transverse electromagnetic field wave. Conversely, during reception, the oscillating electric and magnetic fields of an incoming radio wave exert force on the electrons in the antenna elements, causing them to move back and forth, creating oscillating currents in the antenna. Antennas can be designed to transmit and receive radio waves in all horizontal directions equally (omnidirectional antennas), or preferentially in a particular direction (directional or high gain antennas). In the latter case, an antenna may also include additional elements or surfaces with no electrical connection to the transmitter or receiver, such as parasitic elements, parabolic reflectors, or horns, which serve to direct the radio waves into a beam or other desired radiation pattern.

[0163] Metadata. The term “metadata”, as used herein, refers to data that describes characteristics, attributes, or parameters of other data, in particular, files (such as program files) and objects. Such data typically includes structured information that describes, explains, locates, and otherwise makes it easier to retrieve and use an information resource. Metadata typically includes structural metadata, relating to the design and specification of data structures or "data about the containers of data"; and descriptive metadata about individual instances of application data or the data content. Metadata may include the means of creation of the data, the purpose of the data, time and date of creation, the creator or author of the data, the location on a computer network where the data were created, and the standards used.

[0164] For example, metadata associated with a computer word processing file may include the title of the document, the name of the author, the company to whom the document belongs, the dates that the document was created and last modified, keywords which describe the document, and other descriptive data. While some of this information may also be included in the document itself (e.g., title, author, and data), metadata may be a separate collection of data that may be stored separately from, but associated with, the actual document. One common format for documenting metadata is extensible Markup Language (XML). XML provides a formal syntax, which supports the creation of arbitrary descriptions, sometimes called “tags.” An example of a metadata entry might be <title>War and Peace< / title>, where the bracketed words delineate the beginning and end of the group of characters that constitute the title of the document that is described by the metadata. In the example of the word processing file, the metadata (sometimes referred to as “document properties”) is entered manually by the author, the editor, or the document manager. The metadata concept is further described in a National Information Standards Organization (NISO) Booklet entitled: “Understanding Metadata” (ISBN: 1-880124- 62-9), in the IETF RFC 5013 entitled: “The Dublin Core Metadata Element Set”, and in the IETF RFC 2731 entitled: “Encoding Dublin Core Metadata in HTME”, which are all incorporated in their entirety for all purposes as if fully set forth herein. An extraction of metadata from files or objects is described in a U.S. Patent 8,700,626 to Bedingfield, entitled: “Systems, Methods and Computer Products for Content-Derived Metadata” , and in a U.S. Patent Application Publication 2012 / 0278705 to Yang et al., entitled: “System and Method for Automatically Extracting Metadata from Unstructured Electronic Documents” , which are both incorporated in their entirety for all purposes as if fully set forth herein.

[0165] Metadata can be stored either internally in the same file, object, or structure as the data (this is also called internal or embedded metadata), or externally in a separate file or field separated from the described data. A data repository typically stores the metadata detached from the data, but can be designed to support embedded metadata approaches. Metadata can be stored in either human-readable or binary form. Storing metadata in a human-readable format such as XML can be useful because users can understand and edit it without specialized tools, however, these formats are rarely optimized for storage capacity, communication time, and processing speed. A binary metadata format enables efficiency in all these respects but requires special libraries to convert the binary information into a human-readable content.

[0166] XML. An Extensible Markup Language (XML) is a markup language and file format for storing, transmitting, and reconstructing arbitrary data. It defines a set of rules for encoding documents in a format that is both human-readable and machine-readable. It is a textual data format with strong support via Unicode for different human languages. Although the design of XML focuses on documents, the language is widely used for the representation of arbitrary data structures such as those used in web services. Several schema systems exist to aid in the definition of XML-based languages, while programmers have developed many application programming interfaces (APIs) to aid the processing of XML data.

[0167] The main purpose of XML is serialization, i.e., storing, transmitting, and reconstructing arbitrary data. For two disparate systems to exchange information, they need to agree upon a file format. XML standardizes this process. IETF RFC 7303 (which supersedes the older RFC 3023), provides rules for the construction of media types for use in XML message. It defines three media types: application / xml (text / xml is an alias), application / xml-extemal-parsed- entity (text / xml-extemal-parsed-entity is an alias) and application / xml-dtd. They are used for transmitting raw XML files without exposing their internal semantics. RFC 7303 further recommends that XML-based languages be given media types ending in +xml, for example, image / svg+xml for SVG.

[0168] The Internet Engineering Task Force (IETF) Request for Comments (RFC) 7303 dated July 2014, entitled: “XML Media Types”, which is incorporated in its entirety for all purposes as if fully set forth herein, standardizes three media types: application / xml, application / xml- extemal-parsed-entity, and application / xml-dtd for use in exchanging network entities that are related to the Extensible Markup Language (XML), while defining text / xml and text / xml- extemal-parsed-entity as aliases for the respective application types. This specification also standardizes the '+xml' suffix for naming media types outside of these five types when those media types represent XML MIME entities.

[0169] JSON. JSON (JavaScript Object Notation) is an open standard file format and data interchange format that uses human-readable text to store and transmit data objects consisting of attribute-value pairs and arrays (or other serializable values). It is a common data format with diverse uses in electronic data interchange, including that of web applications with servers. JSON is a language-independent data format. It was derived from JavaScript, but many modem programming languages include code to generate and parse JSON-format data. JSON filenames use the extension ‘.json’. The Internet Engineering Task Force (IETF) Request for Comments (RFC) 8259 dated December 2017, which is incorporated in its entirety for all purposes as if fully set forth herein and entitled: “77ic JavaScript Object Notation (JSON) Data Interchange Format”, describes a lightweight, text-based, language-independent data interchange format. It was derived from the ECMAScript Programming Language Standard. JSON defines a small set of formatting rules for the portable representation of structured data.

[0170] CSV. Comma-Separated Values (CSV) is a text file format that uses commas to separate values. A CSV file stores tabular data (numbers and text) in plain text, where each line of the file typically represents one data record. Each record consists of the same number of fields, and these are separated by commas in the CSV file. If the field delimiter itself may appear within a field, fields can be surrounded with quotation marks. The CSV file format is one type of delimiter- separated file format. Delimiters frequently used include the comma, tab, space, and semicolon. Delimiter- separated files are often given a ".csv" extension even when the field separator is not a comma. Many applications or libraries that consume or produce CSV files have options to specify an alternative delimiter.

[0171] The Internet Engineering Task Force (IETF) Request for Comments (RFC) 7111 dated January 2014 [ISSN: 2070-1721], which is incorporated in its entirety for all purposes as if fully set forth herein and entitled: “URI Fragment Identifiers for the text / csv Media Type”, defines URI fragment identifiers for text / csv MIME entities. These fragment identifiers make it possible to refer to parts of a text / csv MIME entity identified by row, column, or cell. Fragment identification can use single items or ranges. SQL. Structured Query Language (SQL) is a widely-used programming language for working with relational databases, designed for managing data held in a relational database management system (RDBMS), or for stream processing in a relational data stream management system (RDSMS).SQL consists of a data definition language and a data manipulation language. The scope of SQL includes data insert, query, update and delete, schema creation and modification, and data access control. Although SQL is often described as, and largely is, a declarative language (4GL), it also includes procedural elements. SQL is designed for querying data contained in a relational database and is a set-based, declarative query language. The SQL is standardized as ISO / IEC 9075:2011 standard: "Information technology - Database languages - SQL". The ISO / IEC 9075 standard is complemented by ISO / IEC 13249 standard: “SQL Multimedia and Application Packages”, which defines interfaces and packages based on SQL. The aim is unified access to typical database applications like text, pictures, data mining, or spatial data. SQL is described in the tutorial entitled: “Oracle / SQL Tutorial” by Michael Gertz of the University of California, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0172] SSML. Speech Synthesis Markup Language (SSML) is an XML-based markup language for speech synthesis applications. It is a recommendation of the W3C's Voice Browser Working Group and is often embedded in VoiceXML scripts to drive interactive telephony systems. However, it also may be used alone, such as for creating audio books. For desktop applications, other markup languages are popular, including Apple's embedded speech commands, and Microsoft's SAPI Text-To-Speech (TTS) markup, also an XML language. It is also used to produce sounds via Azure Cognitive Services' Text-To-Speech API or when writing third-party skills for Google Assistant or Amazon Alexa. SSML is based on the Java Speech Markup Language (JSML) developed by Sun Microsystems, although the current recommendation was developed mostly by speech synthesis vendors. It covers virtually all aspects of synthesis, although some areas have been left unspecified, so each vendor accepts a different variant of the language. Also, in the absence of markup, the synthesizer is expected to do its own interpretation of the text. SSML specifies a fair amount of markup for prosody, which is not apparent in the above example. This includes markup for pitch, contour, pitch range, rate, duration, and volume.

[0173] Speech Synthesis Markup Language (SSML) is a standard to enable access to the Web using spoken interaction, and is described in W3C Recommendation dated 7 September 2010, entitled: “Speech Synthesis Markup Language (SSML) Version 1.1”, which is incorporated in its entirety for all purposes as if fully set forth herein. The Speech Synthesis Markup Language Specification is one of these standards and is designed to provide a rich, XML-based markup language for assisting the generation of synthetic speech in Web and other applications. The essential role of the markup language is to provide authors of synthesizable content a standard way to control aspects of speech such as pronunciation, volume, pitch, rate, etc. across different synthesis-capable platforms.

[0174] JSML. Java Speech API Markup Language (JSML) is an XML-based markup language for annotating text input to speech synthesizers. JSML is used within the Java Speech API, is an XML application, and conforms to the requirements of well-formed XML documents. Java Speech API Markup Language and JSpeech Markup Language are identical apart from the change in name, which is made to protect Sun trademarks. JSML is primarily an XML text format used by Java applications to annotate text input to speech synthesizers. Elements of JSML provide speech synthesizers with detailed information on how to speak text in a naturalized fashion. JSML defines elements that define a document's structure, the pronunciation of certain words and phrases, features of speech such as emphasis and intonation, etc. JSML is designed in the Java fashion to be simple to learn and use, to be portable across different synthesizers and computing platforms, and although designed for use within is also applicable to a wide range of languages.

[0175] VXML. VoiceXML (VXML) is a digital document standard for specifying interactive media and voice dialogs between humans and computers. It is used for developing audio and voice response applications, such as banking systems and automated customer service portals. VoiceXML applications are developed and deployed in a manner analogous to how a web browser interprets and visually renders the Hypertext Markup Language (HTML) it receives from a web server. VoiceXML documents are interpreted by a voice browser and in common deployment architectures, users interact with voice browsers via the public switched telephone network (PSTN). The VoiceXML document format is based on Extensible Markup Language (XML). VoiceXML applications are commonly used in many industries and segments of commerce. These applications include order inquiry, package tracking, driving directions, emergency notification, wake-up, flight tracking, voice access to email, customer relationship management, prescription refilling, audio news magazines, voice dialling, real-estate information, and national directory assistance applications. VoiceXML has tags that instruct the voice browser to provide speech synthesis, automatic speech recognition, dialog management, and audio playback. The following is an example of a VoiceXML document:

[0176] Voice Extensible Markup Language (VoiceXML) is described in W3C Recommendation dated 19 June 2007, entitled: "Voice Extensible Markup Language (VoiceXML) 2.1”, which is incorporated in its entirety for all purposes as if fully set forth herein. VoiceXML 2.1 specifies a set of features commonly implemented by Voice Extensible Markup Language platforms, and is designed to be fully backward-compatible with VoiceXML 2.0 [VXML2].

[0177] PLS. The Pronunciation Lexicon Specification (PLS) is used to define how words are pronounced, and the generated pronunciation information is meant to be used by both speech recognizers and speech synthesizers in voice browsing applications. The Pronunciation Lexicon Specification (PLS) is designed to enable interoperable specification of pronunciation information for both speech recognition and speech synthesis engines within voice browsing applications. The language is intended to be easy to use by developers while supporting the accurate specification of pronunciation information for international use.

[0178] The language allows one or more pronunciations for a word or phrase to be specified using a standard pronunciation alphabet or if necessary, using vendor- specific alphabets. Pronunciations are grouped together into a PLS document, which may be referenced from other markup languages, such as the Speech Recognition Grammar Specification SRGS and the Speech Synthesis Markup Language SSML.

[0179] Pronunciation Lexicon Specification (PLS) is described in W3C Recommendation dated 14 October 2008, entitled: ” Pronunciation Lexicon Specification (PLS) Version 1.0”, which is incorporated in its entirety for all purposes as if fully set forth herein. This document defines the syntax for specifying pronunciation lexicons to be used by Automatic Speech Recognition and Speech Synthesis engines in voice browser applications.

[0180] ToBI. ToBI (Tones and Break Indices) is a set of conventions for transcribing and annotating the prosody of speech, and is sometimes used to refer to the conventions used for describing American English specifically. A full ToBI transcription consists of six parts: (a) an audio recording, (b) an electronic print-out or paper record of the F0 (fundamental pitch), (c) a tones tier, with an analysis of the tonal events in terms of H and L, (d) a words-tier with the words of the utterance in ordinary writing, (e) a break-index tier showing the strength of the junctures, and (f) a miscellaneous tier with comments.

[0181] Tonal events include pitch accents, phrase accents, and boundary tones. Pitch accents, written as H* or L* (high and low tones, respectively), are typically realized on words that carry the most information in a sentence. For example, in the sentence "Mary went to the store to get some milk", a natural pronunciation would include pitch accents on "Mary", "store", and "milk". Other kinds of pitch accents include L*+H (a syllable which starts with a low accent and then rises) and L+H* (again low-high on one syllable, but with the second part accented). Phrase accents, written H- or L-, are the tones between a pitch accent and a boundary tone. For example, the intonation at the end of a question might be H*L-H%, indicating that the pitch starts high, falls to a low, and rises again; or L*H-H%, indicating that the pitch starts low, then rises steadily to a high.

[0182] Boundary tones, written with H% and L%, are affiliated not to words but to phrase edges. For example, the sentence "Mary went to the store" can be pronounced as a statement or a question ("Mary went to the store." vs. "Mary went to the store?"). The contrast between the statement and the question is signaled by a boundary tone at the end of the phrase: a low boundary tone causes a falling pitch contour, signaling the statement, whereas a high boundary tone causes a rising pitch contour, signaling the question. Break indices are numbers indicating how strong the break is between words: 0 = clitic boundary, e.g., who's; 1 = normal word boundary; 2 = perceived juncture with no intonation effect, or apparent intonational boundary without a pause or any other clues; 3 = intermediate phrase, marked with H- or L-; and 4 = full intonation phrase, marked L% or H%, at the end of a phrase or sentence. The English ToBI standard distinguishes four or five levels of boundary strength, corresponding roughly to breaks between constituents at different levels of the Prosodic Hierarchy. One signal of boundary strength is lengthening of the preceding syllable: the stronger the boundary, the more lengthening of the preceding syllable. In some versions, level 2 is omitted.

[0183] The ToBI conventions are and how the ToBI transcription system works, as well as discussion of the strengths and weaknesses of the ToBI system, are described in chapter 4 published May 2021 [581-95327_ch01_lP.indd] by Sun-Ah Jun entitled: “The ToBI Transcription System: Conventions, Strengths, and Challenges” , which is attached to this document, and is incorporated in its entirety for all purposes as if fully set forth herein. The chapter also describes some of the recent developments in the prosodic transcription system that attempt to address the ToBI system’s known limitations. Before introducing how the ToBI system works, section 4.2 offers a brief description of the theoretical background and framework on which ToBI is based. Section 4.3 presents the ToBI conventions and how the ToBI system has been applied to typologically various languages, showing the workings of the phonological theory it has adopted. Section 4.4 presents the strengths of the ToBI system; section 4.5 discusses the problems and challenges ToBI users face, as well as recent developments that have been made in response to such challenges; and section 4.6 concludes the chapter.

[0184] Prosodic boundaries in speech are of great relevance to both speech synthesis and audio annotation. Examples of experiments for predicting ToBI boundaries and stress as part of the wav2vec 2.0 framework as applied to the task of detecting these boundaries in speech signal, using only acoustic information, is described in an article by Marie Kunesova and Marketa Rezackova published 29 Sep 2022 in New Technologies for the Information Society and Department of Cybernetics, Faculty of Applied Sciences, University of West Bohemia, Pilsen, Czech Republic [arXiv:2209.15032vl [eess.AS]] entitled: "Detection of Prosodic Boundaries in Speech Using Wav2Vec 2.0”, which is incorporated in its entirety for all purposes as if fully set forth herein. The article tests the approach on a set of recordings of Czech broadcast news, labeled by phonetic experts, and compare it to an existing text-based predictor, which uses the transcripts of the same data. Despite using a relatively small amount of labeled data, the wav2vec2 model achieves an accuracy of 94% and Fl measure of 83% on within- sentence prosodic boundaries (or 95% and 89% on all prosodic boundaries), outperforming the text-based approach. However, by combining the outputs of the two different models the results can be improved even further.

[0185] Prosodic event detection plays an important role in spoken language processing tasks and Computer- Assisted Pronunciation Training (CAPT) systems. Traditional methods for the detection of sentence stress and phrase boundaries rely on machine learning methods that model limited contextual information and account little for interaction between these two prosodic events. A hierarchical network for modeling the contextual factors at the granularity of phoneme, syllable and word based on bidirectional Long-Short-Term Memory (BLSTM) is proposed in a paper by Binghuai Lin, Liyuan Wang, Xiaoli Feng, and Jinsong Zhang presented in INTERSPEECH 2020 October 25-29, 2020, Shanghai, China

[0186] [http: / / dx.doi.org / 10.21437 / Interspeech.2020-1284] entitled: “ Joint detection of sentence stress and phrase boundary for prosody”, which is incorporated in its entirety for all purposes as if fully set forth herein. To account for the inherent connection between sentence stress and phrase boundaries, the paper performs a joint modeling of these two important prosodic events with a Multitask Learning Framework (MTL) which shares common prosodic features. The paper evaluates the network performance based on Aix-Machine Readable Spoken English Corpus (AixMARSEC). Experimental results show our proposed method obtains the Fl-measure of 90% for sentence stress detection and 91% for phrase boundary detection, which outperforms the baseline utilizing conditional random field (CRF) by about 4% and 9% respectively.

[0187] ToBI is a prosody labeling system that transcribes American English prosody in terms of phonological tones and break indices. Previous works on automatic ToBI transcription require additional information such as word boundaries and use modular feature extraction with separately optimized feature detectors and classifiers. The problem of pitch accent detection and prosody boundary detection using the Wav2vec 2.0 model with only acoustic information is described in an article by Wanyue Zhail and Mark Hasegawa- Johnson presented in INTERSPEECH 2023 20-24 August 2023, Dublin, Ireland [10.21437 / Interspeech.2023-477] entitled: “Wc / v27?)B / : a new approach to automatic ToBI transcription” , which is incorporated in its entirety for all purposes as if fully set forth herein. The model is trained on the Boston University Radio News Corpus and evaluated on both the Boston University Radio News Corpus and the Boston Directions Corpus. The paper shows that it achieves an Fl score of 0.82 on pitch accent detection and 0.86 on phrase boundary detection.

[0188] Expressive reading, considered the defining attribute of oral reading fluency, comprises the prosodic realization of phrasing and prominence. In the context of evaluating oral reading, it helps to establish the speaker’s comprehension of the text. A labeled dataset of children’s reading recordings is considered for the speaker-independent detection of prominent words using acoustic-prosodic and lexico- syntactic features, as described in a paper by Kamini Sabu, Mithilesh Vaidya, and Preeti Rao, published 28 Jan 2022 [arXiv:2104.05488v3], entitled: “CNN Encoding of Acoustic Parameters for Prominence Detection” , which is incorporated in its entirety for all purposes as if fully set forth herein. A previous well-tuned random forest ensemble predictor is replaced by an RNN sequence classifier to exploit potential context dependency across the longer utterance. Further, deep learning is applied to obtain word-level features from low-level acoustic contours of fundamental frequency, intensity and spectral shape in an end-to-end fashion. Performance comparisons are presented across the different feature types and across different feature learning architectures for prominent word prediction to draw insights wherever possible.

[0189] EmotionML. An Emotion Markup Language (EML or EmotionML) has first been defined by the W3C Emotion Incubator Group (EmoXG) as a general- purpose emotion annotation and representation language, which should be usable in a large variety of technological contexts where emotions need to be represented. Emotion-oriented computing (or "affective computing") is gaining importance as interactive technological systems become more sophisticated. Representing the emotional states of a user or the emotional states to be simulated by a user interface requires a suitable representation format; in this case a markup language is used. Emotion Markup Language (EmotionML) version 1.0 was published 22 May 2014 by W3C Recommendation (available at http: / / www.w3.org / TR / 2014 / REC-emotionml-20140522 / ) as "Emotion Markup Language L0", and is incorporated in its entirety for all purposes as if fully set forth herein. The specification of Emotion Markup Language 1.0 aims to strike a balance between practical applicability and scientific well-founded. The language is conceived as a "plug-in" language suitable for use in three different areas: (1) manual annotation of data; (2) automatic recognition of emotion-related states from user behavior; and (3) generation of emotion- related system behavior.

[0190] Concrete implementations that utilize EmotionML are described in a paper by Steiner, and Tim Llewellyn, published in “In Proceedings of the 5th International Workshop on Emotion, Sentiment, Social Signals and Linked Open Data (ES3LOD)” [vol. 80. 2014], entitled: “ Application of emotionml", which is incorporated in its entirety for all purposes as if fully set forth herein.

[0191] Cloud. The term “Cloud” or "Cloud computing" as used herein is defined as a technology infrastructure facilitating supplement, consumption, and delivery of IT services, and generally refers to any group of networked computers capable of delivering computing services (such as computations, applications, data access, and data management and storage resources) to end users. This disclosure does not limit the type (such as public or private) of the cloud as well as the underlying system architecture used by the cloud. The IT services are Internet-based and may involve elastic provisioning of dynamically scalable and time-virtualized resources. Although such virtualization environments can be privately deployed and used within local area or wide area networks owned by an enterprise, a number of “cloud service providers” host virtualization environments accessible through the public internet (the “public cloud”) that is generally open to anyone, or through private IP or another type of network accessible only by entities given access to it (a “private cloud.”). Using a cloud-based control server or using the system above may allow for reduced capital or operational expenditures. The users may further access the system using a web browser regardless of their location or what device they are using, and the virtualization technology allows servers and storage devices to be shared and utilization be increased. Examples of public cloud providers include Amazon AWS, Microsoft Azure, and Google GCP. Comparison of service features such as computation, storage, and infrastructure of the three cloud service providers (AWS, Microsoft Azure, and GoogleGCP) is disclosed in an article entitled: “Highlight the Features of AWS, GCP and Microsoft Azure that Have an Impact when Choosing a Cloud Service Provider'’ by Muhammad Ayoub Kamal, Hafiz Wahab Raza, Muhammad Mansoor Alam, and Mazliham Mohd Su’ud, published January 2020 in ‘International Journal of Recent Technology and Engineering (URTE)’ ISSN: 2277-3878, Volume-8by Blue Eyes Intelligence Engineering & Sciences Publication [DGI:10.35940 / jute.D8573.018520], which is incorporated in its entirety for all purposes as if fully set forth herein.

[0192] The term "Software as a Service (SaaS)" as used herein in this application, is defined as a model of software deployment whereby a provider licenses a Software Application (SA) to customers for use as a service on demand. Similarly, an “Infrastructure as a Service” (laaS) allows enterprises to access virtualized computing systems through the public Internet. The term "customer" as used herein in this application, is defined as a business entity that is served by an SA, provided on the SaaS platform. A customer may be a person or an organization and may be represented by a user that is responsible for the administration of the application in aspects of permissions configuration, user-related configuration, and data security policy. The service is supplied and consumed over the Internet, thus eliminating requirements to install and run applications locally on a site of a customer as well as simplifying maintenance and support. Particularly it is advantageous in massive business applications. Licensing is a common form of billing for the service and it is paid periodically. SaaS is becoming ever more common as a form of SA delivery over the Internet and is being facilitated in a technology infrastructure called "Cloud Computing". In this form of SA delivery, where the SA is controlled by a service provider, a customer may experience stability and data security issues. In many cases, the customer is a business organization that is using the SaaS for business purposes such as business software; hence, stability and data security are primary requirements. As part of a cloud service arrangement, any computer system may also be emulated using software running on a hardware computer system. This virtualization allows for multiple instances of a computer system, each referred to as virtual machine, to run on a single machine. Each virtual machine behaves like a computer system running directly on hardware. It is isolated from the other virtual machines, as would two hardware computers. Each virtual machine comprises an instance of an operating system (the “guest operating system”). There is a host operating system running directly on the hardware that supports the software that emulates the hardware, and the emulation software is referred to as a hypervisor.

[0193] The term “cloud-based” generally refers to a hosted service that is remotely located from a data source and configured to receive, store, and process data delivered by the data source over a network. Cloud-based systems may be configured to operate as a public cloud-based service, a private cloud-based service, or a hybrid cloud-based service. A “public cloud-based service” may include a third-party provider that supplies one or more servers to host multi-tenant services. Examples of a public cloud-based service include Amazon Web Services® (AWS®), Microsoft® Azure™, and Google® Compute Engine™ (GCP) as examples. In contrast, a “private” cloud-based service may include one or more servers that host services provided to a single subscriber (enterprise) and a hybrid cloud-based service may be a combination of certain functionality from a public cloud-based service and a private cloud-based service. Cloud computing and virtualization are described in a book entitled “Cloud Computing and Virtualization” authored by Dac-Nhuong Le (Faculty of Information Technology, Haiphong University, Haiphong, Vietnam), Raghvendra Kumar (Department of Computer Science and Engineering, LNCT, Jabalpur, India), Gia Nhu Nguyen (Graduate School, Duy Tan University, Da Nang, Vietnam), and Jyotir Moy Chatterjee (Department of Computer Science and Engineering at GD-RCET, Bhilai, India), and published 2018 by John Wiley & Sons, Inc. [ISBN 978-1-119-48790-6], which is incorporated in its entirety for all purposes as if fully set forth herein. The book describes the adoption of virtualization in data centers creates the need for a new class of networks designed to support the elasticity of resource allocation, increasing mobile workloads, and the shift to the production of virtual workloads, requiring maximum availability. Building a network that spans both physical servers and virtual machines with consistent capabilities demands a new architectural approach to designing and building the IT infrastructure. Performance, elasticity, and logical addressing structures must be considered as well as the management of the physical and virtual networking infrastructure. Once deployed, a network that is virtualization-ready can offer many revolutionary services over a common shared infrastructure. Virtualization technologies from VMware, Citrix, and Microsoft encapsulate existing applications and extract them from the physical hardware. Unlike physical machines, virtual machines are represented by a portable software image, which can be instantiated on physical hardware at a moment’s notice. With virtualization, comes elasticity where computer capacity can be scaled up or down on demand by adjusting the number of virtual machines actively executing on a given physical server. Additionally, virtual machines can be migrated while in service from one physical server to another.

[0194] Extending this further, virtualization creates “location freedom” enabling virtual machines to become portable across an ever-increasing geographical distance. As cloud architectures and multi-tenancy capabilities continue to develop and mature, there is an economy of scale that can be realized by aggregating resources across applications, business units, and separate corporations to a common shared, yet segmented, infrastructure. Elasticity, mobility, automation, and density of virtual machines demand new network architectures focusing on high performance, addressing portability, and the innate understanding of the virtual machine as the new building block of the data center. Consistent network- supported and virtualization-driven policies and controls are necessary for the visibility to virtual machines’ state and location as they are created and moved across a virtualized infrastructure.

[0195] Virtualization technologies in data center environments are described in an eBook authored by Gustavo Alessandro Andrade Santana and published 2014 by Cisco Systems, Inc. (Cisco Press) [ISBN-13: 978-1-58714-324-3] entitled: “Data Center Virtualization Fundamentals” , which is incorporated in its entirety for all purposes as if fully set forth herein. PowerVM technology for virtualization is described in IBM RedBook entitled: “IBM PowerVM Virtualization - Introduction and Configuration” published by IBM Corporation June 2013, and virtualization basics is described in a paper by IBM Corporation published 2009 entitled: “Power Systems - Introduction to virtualization” , which are both incorporated in their entirety for all purposes as if fully set forth herein.

[0196] Server. The Internet architecture employs a client-server model, among other arrangements. The terms 'server' or 'server computer' herein relate to a device or computer (or a plurality of computers) connected to the Internet and is used for providing facilities or services to other computers or other devices (referred to in this context as 'clients') connected to the Internet. A server is commonly a host that has an IP address and executes a 'server program', and typically operates as a socket listener. Many servers have dedicated functionality such as web server, Domain Name System (DNS) server (described in RFC 1034 and RFC 1035), Dynamic Host Configuration Protocol (DHCP) server (described in RFC 2131 and RFC 3315), mail server, File Transfer Protocol (FTP) server and database server. Similarly, the term 'client' is used herein to include, but not limited to, a program or a device or a computer (or a series of computers) executing this program, which accesses a server over the Internet for a service or a resource. Clients commonly initiate connections that a server may accept. For non-limiting examples, web browsers are clients that connect to web servers for retrieving web pages, and email clients connect to mail storage servers for retrieving mails.

[0197] A server device (in server / client architecture) typically offers information resources, services, and applications to clients, using a server dedicated or oriented operating system. A server device may consist of, be based on, include, or be included in a work-station. Current popular server operating systems are based on Microsoft Windows (by Microsoft Corporation, headquartered in Redmond, Washington, U.S.A.), Unix, and Linux-based solutions, such as the ‘Windows Server 2012’ server operating system, which is a part of the Microsoft ‘Windows Server’ OS family, that was released by Microsoft in 2012. ‘Windows Server 2012’ provides enterprise-class datacenter and hybrid cloud solutions that are simple to deploy, cost-effective, application-specific, and user-centric, and is described in Microsoft publication entitled: “Inside- Out Windows Server 2012”, by William R. Stanek, published 2013 by Microsoft Press, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0198] Unix operating system is widely used in servers. It is a multitasking, multiuser computer operating system that exists in many variants and is characterized by a modular design that is sometimes called the "Unix philosophy", meaning the OS provides a set of simple tools, which each performs a limited, well-defined function, with a unified filesystem as the primary means of communication, and a shell scripting and command language to combine the tools to perform complex workflows. Unix was designed to be portable, multi-tasking, and multi-user in a timesharing configuration, and Unix systems are characterized by various concepts: the use of plain text for storing data, a hierarchical file system, treating devices and certain types of Inter-Process Communication (IPC) as files, the use of a large number of software tools, and small programs that can be strung together through a command line interpreter using pipes, as opposed to using a single monolithic program that includes all of the same functionality. Unix operating system consists of many utilities along with the master control program, the kernel. The kernel provides services to start and stop programs, handles the file system and other common "low level" tasks that most programs share, and schedules access to avoid conflicts when programs try to access the same resource, or device simultaneously. To mediate such access, the kernel has special rights, reflected in the division between user-space and kernel-space. Unix is described in a publication entitled: “UNIX Tutorial” by tutorialspoint.com, downloaded on July 2014, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0199] Client. The term ‘client’ typically refers to an application (or a device executing the application) used for retrieving or rendering resources, or resource manifestations, such as a web browser, an e-mail reader, or a Usenet reader, while the term ‘server’ typically refers to an application (or a device executing the application) used for supplying resources or resource manifestations, and typically offers (or hosts) various services to other network computers and users. These services are usually provided through ports or numbered access points beyond the server's network address. Each port number is usually associated with a maximum of one running program, which is responsible for handling requests to that port. A daemon, being a user program, can in turn access the local hardware resources of that computer by passing requests to the operating system kernel.

[0200] A client device (in server / client architecture) typically receives information resources, services, and applications from servers, and is using a client dedicated or oriented operating system. The client device may consist of, may be based on, may include, or may be included in, any workstation or a computer system. Current popular client operating systems are based on Microsoft Windows (by Microsoft Corporation, headquartered in Redmond, Washington, U.S.A.), which is a series of graphical interface operating systems developed, marketed, and sold by Microsoft. Microsoft Windows is described in Microsoft publications entitled: “Windows Internals - Part 1” and “Windows Internals - Part 2”, by Mark Russinovich, David A. Solomon, and Alex loescu, published by Microsoft Press in 2012, which are both incorporated in their entirety for all purposes as if fully set forth herein. Windows 8 is a personal computer operating system developed by Microsoft as part of Windows NT family of operating systems, that was released for general availability on October 2012, and is described in Microsoft Press 2012 publication entitled: “Introducing Windows 8 - An Overview for IT Professionals” by Jerry Honeycutt, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0201] Chrome OS is a Linux kernel-based operating system designed by Google Inc. out of Mountain View, California, U.S.A., to work primarily with web applications. The user interface takes a minimalist approach and consists almost entirely of just the Google Chrome web browser; since the operating system is aimed at users who spend most of their computer time on the Web, the only "native" applications on Chrome OS are a browser, media player and file manager, and hence the Chrome OS is almost a pure web thin client OS.

[0202] The Chrome OS is described as including a three-tier architecture: firmware, browser and window manager, and system-level software and userland services. The firmware contributes to fast boot time by not probing for hardware, such as floppy disk drives, that are no longer common on computers, especially netbooks. The firmware also contributes to security by verifying each step in the boot process and incorporating system recovery. The system-level software includes the Linux kernel that has been patched to improve boot performance. The userland software has been trimmed to essentials, with management by Upstart, which can launch services in parallel, re-spawn crashed jobs, and defer services in the interest of faster booting. The Chrome OS user guide is described in the Samsung Electronics Co., Ltd. presentation entitled: “Google™ Chrome OS USER GUIDE” published 2011, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0203] RTOS. A Real-Time Operating System (RTOS) is an Operating System (OS) intended to serve real-time applications that process data as it comes in, typically without buffer delays. Processing time requirements (including any OS delay) are typically measured in tenths of seconds or shorter increments of time, and is a time bound system which has well defined fixed time constraints. Processing is commonly to be done within the defined constraints, or the system will fail. They either are event driven or time sharing, where event driven systems switch between tasks based on their priorities while time sharing systems switch the task based on clock interrupts. A key characteristic of an RTOS is the level of its consistency concerning the amount of time it takes to accept and complete an application's task; the variability is jitter. A hard realtime operating system has less jitter than a soft real-time operating system. The chief design goal is not high throughput, but rather a guarantee of a soft or hard performance category. An RTOS that can usually or generally meet a deadline is a soft real-time OS, but if it can meet a deadline deterministically it is a hard real-time OS. An RTOS has an advanced algorithm for scheduling, and includes a scheduler flexibility that enables a wider, computer- system orchestration of process priorities. Key factors in a real-time OS are minimal interrupt latency and minimal thread switching latency; a real-time OS is valued more for how quickly or how predictably it can respond than for the amount of work it can perform in a given period of time.

[0204] Common designs of RTOS include event-driven, where tasks are switched only when an event of higher priority needs servicing; called preemptive priority, or priority scheduling, and time-sharing, where task are switched on a regular clocked interrupt, and on events; called round robin. Time sharing designs switch tasks more often than strictly needed, but give smoother multitasking, giving the illusion that a process or user has sole use of a machine. In typical designs, a task has three states: Running (executing on the CPU); Ready (ready to be executed); and Blocked (waiting for an event, I / O for example). Most tasks are blocked or ready most of the time because generally only one task can run at a time per CPU. The number of items in the ready queue can vary greatly, depending on the number of tasks the system needs to perform and the type of scheduler that the system uses. On simpler non-preemptive but still multitasking systems, a task has to give up its time on the CPU to other tasks, which can cause the ready queue to have a greater number of overall tasks in the ready to be executed state (resource starvation).

[0205] RTOS concepts and implementations are described in an Application Note No. RES05B00008-0100 / Rec. 1.00 published January 2010 by Renesas Technology Corp, entitled: ”R8C Family - General RTOS Concepts”, in JAJA Technology Review article published February 2007 [1535-5535 / S32.00] by The Association for Laboratory Automation [doi: 10.1016 / j.jala.2006.10.016] entitled: “An Overview of Real-Time Operating Systems”, and in Chapter 2 entitled: “ Basic Concepts of Real Time Operating Systems” of a book published 2009 [ISBN - 978-1-4020-9435-4] by Springer Science + Business Media B.V. entitled: “Hardware-Dependent Software - Principles and Practice”, which are all incorporated in their entirety for all purposes as if fully set forth herein.

[0206] QNX. One example of RTOS is QNX, which is a commercial Unix-like real-time operating system, aimed primarily at the embedded systems market. QNX was one of the first commercially successful microkernel operating systems and is used in a variety of devices including cars and mobile phones. As a microkernel-based OS, QNX is based on the idea of running most of the operating system kernel in the form of a number of small tasks, known as Resource Managers. In the case of QNX, the use of a microkernel allows users (developers) to turn off any functionality they do not require without having to change the OS itself; instead, those services will simply not run.

[0207] FreeRTOS. FreeRTOS™ is a free and open-source Real-Time Operating system developed by Real Time Engineers Ltd., designed to fit on small embedded systems and implements only a very minimalist set of functions: very basic handle of tasks and memory management, and just sufficient API concerning synchronization. Its features include characteristics such as preemptive tasks, support for multiple microcontroller architectures, a small footprint (4.3Kbytes on an ARM7 after compilation), written in C, and compiled with various C compilers. It also allows an unlimited number of tasks to run at the same time, and no limitation about their priorities as long as used hardware can afford it.

[0208] FreeRTOS™ provides methods for multiple threads or tasks, mutexes, semaphores and software timers. A tick-less mode is provided for low power applications, and thread priorities are supported. Four schemes of memory allocation are provided: allocate only; allocate and free with a very simple, fast, algorithm; a more complex but fast allocate and free algorithm with memory coalescence; and C library allocate and free with some mutual exclusion protection. While the emphasis is on compactness and speed of execution, a command line interface and POSIX-like IO abstraction add-ons are supported. FreeRTOS™ implements multiple threads by having the host program call a thread tick method at regular short intervals.

[0209] The thread tick method switches tasks depending on priority and a round-robin scheduling scheme. The usual interval is 1 / 1000 of a second to 1 / 100 of a second, via an interrupt from a hardware timer, but this interval is often changed to suit a particular application. FreeRTOS™ is described in a paper by Nicolas Melot (downloaded 7 / 2015) entitled: “Study of an operating system: FreeRTOS - Operating systems for embedded devices”, in a paper (dated September 23, 2013) by Dr. Richard Wall entitled: “Carebot PIC32 MX7ck implementation of Free RTOS”, FreeRTOS™ modules are described in web pages entitled: “FreeRTOS™ Modules” published in the www, freertos.org web-site dated 26.11.2006, and FreeRTOS kernel is described in a paper published 1 April 07 by Rich Goyette of Carleton University as part of ‘SYSC5701: Operating System Methods for Real-Time Applications’, entitled: “An Analysis and Description of the Inner Workings of the FreeRTOS Kernel”, which are all incorporated in their entirety for all purposes as if fully set forth herein.

[0210] SafeRTOS. SafeRTOS was constructed as a complementary offering to FreeRTOS, with common functionality but with a uniquely designed safety-critical implementation. When the FreeRTOS functional model was subjected to a full HAZOP, weakness with respect to user misuse and hardware failure within the functional model and API were identified and resolved. Both SafeRTOS and FreeRTOS share the same scheduling algorithm, have similar APIs, and are otherwise very similar, but they were developed with differing objectives. SafeRTOS was developed solely in the C language to meet requirements for certification to IEC61508. SafeRTOS is known for its ability to reside solely in the on-chip read only memory of a microcontroller for standards compliance. When implemented in hardware memory, SafeRTOS code can only be utilized in its original configuration, so certification testing of systems using this OS need not re-test this portion of their designs during the functional safety certification process.

[0211] VxWorks. VxWorks is an RTOS developed as proprietary software and designed for use in embedded systems requiring real-time, deterministic performance and, in many cases, safety and security certification, for industries, such as aerospace and defense, medical devices, industrial equipment, robotics, energy, transportation, network infrastructure, automotive, and consumer electronics. VxWorks supports Intel architecture, POWER architecture, and ARM architectures. The VxWorks may be used in multicore asymmetric multiprocessing (AMP), symmetric multiprocessing (SMP), and mixed modes and multi-OS (via Type 1 hypervisor) designs on 32- and 64-bit processors. VxWorks comes with the kernel, middleware, board support packages, Wind River Workbench development suite and complementary third-party software and hardware technologies. In its latest release, VxWorks 7, the RTOS has been reengineered for modularity and upgradeability so the OS kernel is separate from middleware, applications and other packages. Scalability, security, safety, connectivity, and graphics have been improved to address Internet of Things (loT) needs. pC / OS. Micro-Controller Operating Systems (MicroC / OS, stylized as pC / OS) is a realtime operating system (RTOS) that is a priority-based preemptive real-time kernel for microprocessors, written mostly in the programming language C, and is intended for use in embedded systems. MicroC / OS allows defining several functions in C, each of which can execute as an independent thread or task. Each task runs at a different priority, and runs as if it owns the central processing unit (CPU). Lower priority tasks can be preempted by higher priority tasks at any time. Higher priority tasks use operating system (OS) services (such as a delay or event) to allow lower priority tasks to execute. OS services are provided for managing tasks and memory, communicating between tasks, and timing.

[0212] Database. A database is generally a grouping of data values typically stored in a computer memory and organized for convenient data value access by a database management system. More specific to the present patent application, a database is a defined data structure, generally stored in a computer memory, comprised of database tables, common database tables, database columns, database indexes, common unique indexes, foreign key constraints, common index master constraints and other database objects defined using a computer-based database management system. A database may be an organized collection of data, typically managed by a DataBase Management System (DBMS) that organizes the storage of data and performs other functions such as the creation, maintenance, and usage of the database storage structures. The data is typically organized to model aspects of reality in a way that supports processes requiring information. Databases commonly also provide users with a user interface and front-end that enables the users to query the database, often in complex manners that require processing and organization of the data. The term "database" is used herein to refer to a database, or to both a database and the DBMS used to manipulate it. Database Management Systems (DBMS) are typically computer software applications that interact with the user, other applications, and the database itself to capture and analyze data, typically providing various functions that allow entry, storage and retrieval of large quantities of information, as well as providing ways to manage how that information is organized. A general-purpose DBMS is designed to allow the definition, creation, querying, update, and administration of databases. Examples of DBMSs include MySQL, PostgreSQL, Microsoft SQL Server, Oracle, Sybase and IBM DB2. Database technology and application is described in a document published by Telemark University College entitled: “Introduction to Database Systems”, authored by Hans-Petter Halvorsen (dated 2014.03.03), which is incorporated in its entirety for all purposes as if fully set forth herein.

[0213] Wireless. Any embodiment herein may be used in conjunction with one or more types of wireless communication signals and / or systems, for example, Radio Frequency (RF), Infra-Red (IR), Frequency-Division Multiplexing (FDM), Orthogonal FDM (OFDM), Time-Division Multiplexing (TDM), Time-Division Multiple Access (TDMA), Extended TDMA (E-TDMA), General Packet Radio Service (GPRS), extended GPRS, Code-Division Multiple Access (CDMA), Wideband CDMA (WCDMA), CDMA 2000, single-carrier CDMA, multi-carrier CDMA, Multi-Carrier Modulation (MDM), Discrete Multi-Tone (DMT), Bluetooth (RTM), Global Positioning System (GPS), Wi-Fi, Wi-Max, ZigBee (TM), Ultra-Wideband (UWB), Global System for Mobile communication (GSM), 2G, 2.5G, 3G, 3.5G, Enhanced Data rates for GSM Evolution (EDGE), or the like. Any wireless network or wireless connection herein may be operating substantially in accordance with existing IEEE 802.11, 802.11a, 802.11b, 802.11g, 802.11k, 802.1 In, 802. Hr, 802.16, 802.16d, 802.16e, 802.20, 802.21 standards and / or future versions and / or derivatives of the above standards. Further, a network element (or a device) herein may consist of, be part of, or include, a cellular radio-telephone communication system, a cellular telephone, a wireless telephone, a Personal Communication Systems (PCS) device, a PDA device that incorporates a wireless communication device, or a mobile / portable Global Positioning System (GPS) device. Further, wireless communication may be based on wireless technologies that are described in Chapter 20: "Wireless Technologies" of the publication number 1-587005-001-3 by Cisco Systems, Inc. (7 / 99) entitled: "Internetworking Technologies Handbook" , which is incorporated in its entirety for all purposes as if fully set forth herein. Wireless technologies and networks are further described in a book published 2005 by Pearson Education, Inc. William Stallings [ISBN: 0-13-191835-4] entitled: “Wireless Communications and Networks - second Edition”, which is incorporated in its entirety for all purposes as if fully set forth herein.

[0214] Wireless networking typically employs an antenna (a.k.a. aerial), which is an electrical device that converts electric power into radio waves, and vice versa, connected to a wireless radio transceiver. In transmission, a radio transmitter supplies an electric current oscillating at radio frequency to the antenna terminals, and the antenna radiates the energy from the current as electromagnetic waves (radio waves). In reception, an antenna intercepts some of the power of an electromagnetic wave in order to produce a low-voltage at its terminals that is applied to a receiver to be amplified. Typically an antenna consists of an arrangement of metallic conductors (elements), electrically connected (often through a transmission line) to the receiver or transmitter. An oscillating current of electrons forced through the antenna by a transmitter will create an oscillating magnetic field around the antenna elements, while the charge of the electrons also creates an oscillating electric field along the elements. These time-varying fields radiate away from the antenna into space as a moving transverse electromagnetic field wave. Conversely, during reception, the oscillating electric and magnetic fields of an incoming radio wave exert force on the electrons in the antenna elements, causing them to move back and forth, creating oscillating currents in the antenna. Antennas can be designed to transmit and receive radio waves in all horizontal directions equally (omnidirectional antennas), or preferentially in a particular direction (directional or high gain antennas). In the latter case, an antenna may also include additional elements or surfaces with no electrical connection to the transmitter or receiver, such as parasitic elements, parabolic reflectors, or horns, which serve to direct the radio waves into a beam or other desired radiation pattern.

[0215] ISM. The Industrial, Scientific and Medical (ISM) radio bands are radio bands (portions of the radio spectrum) reserved internationally for the use of radio frequency (RF) energy for industrial, scientific and medical purposes other than telecommunications. In general, communications equipment operating in these bands must tolerate any interference generated by ISM equipment, and users have no regulatory protection from ISM device operation. The ISM bands are defined by the fTU-R in 5.138, 5.150, and 5.280 of the Radio Regulations. Individual countries use of the bands designated in these sections may differ due to variations in national radio regulations. Because communication devices using the ISM bands must tolerate any interference from ISM equipment, unlicensed operations are typically permitted to use these bands, since unlicensed operation typically needs to be tolerant of interference from other devices anyway. The ISM bands share allocations with unlicensed and licensed operations; however, due to the high likelihood of harmful interference, licensed use of the bands is typically low. In the United States, uses of the ISM bands are governed by Part 18 of the Federal Communications Commission (FCC) rules, while Part 15 contains the rules for unlicensed communication devices, even those that share ISM frequencies. In Europe, the ETSI is responsible for governing ISM bands.

[0216] Commonly used ISM bands include a 2.45 GHz band (also known as 2.4 GHz band) that includes the frequency band between 2.400 GHz and 2.500 GHz, a 5.8 GHz band that includes the frequency band 5.725 - 5.875 GHz, a 24GHz band that includes the frequency band 24.000 - 24.250 GHz, a 61 GHz band that includes the frequency band 61.000 - 61.500 GHz, a 122 GHz band that includes the frequency band 122.000 - 123.000 GHz, and a 244 GHz band that includes the frequency band 244.000 - 246.000 GHz.

[0217] ZigBee. ZigBee is a standard for a suite of high-level communication protocols using small, low-power digital radios based on an IEEE 802 standard for Personal Area Network (PAN). Applications include wireless light switches, electrical meters with in-home displays, and other consumer and industrial equipment that require a short-range wireless transfer of data at relatively low rates. The technology defined by the ZigBee specification is intended to be simpler and less expensive than other WPANs, such as Bluetooth. ZigBee is targeted at Radio- Frequency (RF) applications that require a low data rate, long battery life, and secure networking. ZigBee has a defined rate of 250 kbps suited for periodic or intermittent data or a single signal transmission from a sensor or input device.

[0218] ZigBee builds upon the physical layer and medium access control defined in IEEE standard 802.15.4 (2003 version) for low-rate WPANs. The specification further discloses four main components: network layer, application layer, ZigBee Device Objects (ZDOs), and manufacturer-defined application objects, which allow for customization and favor total integration. The ZDOs are responsible for several tasks, which include the keeping of device roles, management of requests to join a network, device discovery, and security. Because ZigBee nodes can go from sleep to active mode in 30 ms or less, the latency can be low and devices can be responsive, particularly compared to Bluetooth wake-up delays, which are typically around three seconds. ZigBee nodes can sleep most of the time, thus the average power consumption can be lower, resulting in longer battery life.

[0219] There are three defined types of ZigBee devices: ZigBee Coordinator (ZC), ZigBee Router (ZR), and ZigBee End Device (ZED). ZigBee Coordinator (ZC) is the most capable device and forms the root of the network tree and might bridge to other networks. There is exactly one defined ZigBee coordinator in each network since it is the device that started the network originally. It can store information about the network, including acting as the Trust Center & repository for security keys. ZigBee Router (ZR) may be running an application function as well as may be acting as an intermediate router, passing on data from other devices. ZigBee End Device (ZED) contains functionality to talk to a parent node (either the coordinator or a router). This relationship allows the node to be asleep a significant amount of time, thereby giving long battery life. A ZED requires the least amount of memory and therefore can be less expensive to manufacture than a ZR or ZC.

[0220] The protocols build on recent algorithmic research (Ad-hoc On-demand Distance Vector, neuRFon) to automatically construct a low-speed ad-hoc network of nodes. In most large network instances, the network will be a cluster of clusters. It can also form a mesh or a single cluster. The current ZigBee protocols support beacon and non-beacon enabled networks. In non-beacon-enabled networks, an unslotted CSMA / CA channel access mechanism is used. In this type of network, ZigBee Routers typically have their receivers continuously active, requiring a more robust power supply. However, this allows for heterogeneous networks in which some devices receive continuously, while others only transmit when an external stimulus is detected.

[0221] In beacon-enabled networks, the special network nodes called ZigBee Routers transmit periodic beacons to confirm their presence to other network nodes. Nodes may sleep between the beacons, thus lowering their duty cycle and extending their battery life. Beacon intervals depend on the data rate; they may range from 15.36 milliseconds to 251.65824 seconds at 250 Kbit / s, from 24 milliseconds to 393.216 seconds at 40 Kbit / s, and from 48 milliseconds to 786.432 seconds at 20 Kbit / s. In general, the ZigBee protocols minimize the time the radio is on to reduce power consumption. In beaconing networks, nodes only need to be active while a beacon is being transmitted. In non-beacon-enabled networks, power consumption is decidedly asymmetrical: some devices are always active while others spend most of their time sleeping.

[0222] Except for the Smart Energy Profile 2.0, current ZigBee devices conform to the IEEE 802.15.4-2003 Low-Rate Wireless Personal Area Network (LR-WPAN) standard. The standard specifies the lower protocol layers — the PHYsical layer (PHY), and the Media Access Control (MAC) portion of the Data Link Layer (DLL). The basic channel access mode is "Carrier Sense, Multiple Access / Collision Avoidance" (CSMA / CA), that is, the nodes talk in the same way that people converse; they briefly check to see that no one is talking before they start. There are three notable exceptions to the use of CSMA. Beacons are sent on a fixed time schedule, and do not use CSMA. Message acknowledgments also do not use CSMA. Finally, devices in Beacon Oriented networks that have low latency real-time requirements may also use Guaranteed Time Slots (GTS), which by definition do not use CSMA.

[0223] Z-Wave. Z-Wave is a wireless communications protocol by the Z-Wave Alliance (http: / / www.z-wave.com) designed for home automation, specifically for remote control applications in residential and light commercial environments. The technology uses a low -power RF radio embedded or retrofitted into home electronics devices and systems, such as lighting, home access control, entertainment systems, and household appliances. Z-Wave communicates using a low-power wireless technology designed specifically for remote control applications. Z- Wave operates in the sub-gigahertz frequency range, around 900 MHz. This band competes with some cordless telephones and other consumer electronics devices but avoids interference with WiFi and other systems that operate on the crowded 2.4 GHz band. Z-Wave is designed to be easily embedded in consumer electronics products, including battery-operated devices such as remote controls, smoke alarms, and security sensors.

[0224] Z-Wave is a mesh networking technology where each node or device on the network is capable of sending and receiving control commands through walls or floors, and uses intermediate nodes to route around household obstacles or radio dead spots that might occur in the home. Z-Wave devices can work individually or in groups, and can be programmed into scenes or events that trigger multiple devices, either automatically or via remote control. The Z- wave radio specifications include bandwidth of 9,600 bit / s or 40 Kbit / s, fully interoperable, GFSK modulation, and a range of approximately 100 feet (or 30 meters) assuming "open air" conditions, with reduced range indoors depending on building materials, etc. The Z-Wave radio uses the 900 MHz ISM band: 908.42 MHz (United States); 868.42 MHz (Europe); 919.82 MHz (Hong Kong); and 921.42 MHz (Australia / New Zealand).

[0225] Z-Wave uses a source-routed mesh network topology and has one or more master controllers that control routing and security. The devices can communicate to another by using intermediate nodes to actively route around, and circumvent household obstacles or radio dead spots that might occur. A message from node A to node C can be successfully delivered even if the two nodes are not within range, providing that a third node B can communicate with nodes A and C. If the preferred route is unavailable, the message originator will attempt other routes until a path is found to the "C" node. Therefore, a Z-Wave network can span much farther than the radio range of a single unit; however, with several of these hops, a delay may be introduced between the control command and the desired result. In order for Z-Wave units to be able to route unsolicited messages, they cannot be in sleep mode. Therefore, most battery-operated devices are not designed as repeater units. A Z-Wave network can consist of up to 232 devices with the option of bridging networks if more devices are required.

[0226] WWAN. Any wireless network herein may be a Wireless Wide Area Network (WWAN) such as a wireless broadband network, and the WWAN port may be an antenna and the WWAN transceiver may be a wireless modem. The wireless network may be a satellite network, the antenna may be a satellite antenna, and the wireless modem may be a satellite modem. The wireless network may be a WiMAX network such as according to, compatible with, or based on, IEEE 802.16-2009, the antenna may be a WiMAX antenna, and the wireless modem may be a WiMAX modem. The wireless network may be a cellular telephone network, the antenna may be a cellular antenna, and the wireless modem may be a cellular modem. The cellular telephone network may be a Third Generation (3G) network, and may use UMTS W- CDMA, UMTS HSPA, UMTS TDD, CDMA2000 IxRTT, CDMA2000 EV-DO, or GSM EDGE-Evolution. The cellular telephone network may be a Fourth Generation (4G) network and may use or be compatible with HSPA+, Mobile WiMAX, LTE, LTE- Advanced, MBWA, or may be compatible with, or based on, IEEE 802.20-2008.

[0227] WLAN. Wireless Local Area Network (WLAN), is a popular wireless technology that makes use of the Industrial, Scientific and Medical (ISM) frequency spectrum. In the US, three of the bands within the ISM spectrum are the A band, 902-928 MHz; the B band, 2.4-2.484 GHz (a.k.a. 2.4 GHz); and the C band, 5.725-5.875 GHz (a.k.a. 5 GHz). Overlapping and / or similar bands are used in different regions such as Europe and Japan. In order to allow interoperability between equipment manufactured by different vendors, few WLAN standards have evolved, as part of the IEEE 802.11 standard group, branded as WiFi (www.wi-fi.org). IEEE 802.11b describes a communication using the 2.4GHz frequency band and supporting communication rate of UMb / s, IEEE 802.11a uses the 5GHz frequency band to carry 54MB / s and IEEE 802.11g uses the 2.4 GHz band to support 54Mb / s. The WiFi technology is further described in a publication entitled: “WzFz Technology” by Telecom Regulatory Authority, published on July 2003, which is incorporated in its entirety for all purposes as if fully set forth herein. The IEEE 802 defines an ad-hoc connection between two or more devices without using a wireless access point: the devices communicate directly when in range. An ad hoc network offers peer-to-peer layout and is commonly used in situations such as a quick data exchange or a multiplayer LAN game because the setup is easy and an access point is not required.

[0228] A node / client with a WLAN interface is commonly referred to as STA (Wireless Station / Wireless client). The STA functionality may be embedded as part of the data unit, or may be a dedicated unit, referred to as a bridge, coupled to the data unit. While STAs may communicate without any additional hardware (ad-hoc mode), such a network usually involves Wireless Access Point (a.k.a. WAP or AP) as a mediation device. The WAP implements the Basic Stations Set (BSS) and / or ad-hoc mode based on Independent BSS (IBSS). STA, client, bridge, and WAP will be collectively referred to hereon as WLAN unit. Bandwidth allocation for IEEE 802.11g wireless in the U.S. allows multiple communication sessions to take place simultaneously, where eleven overlapping channels are defined spaced 5MHz apart, spanning from 2412 MHz as the center frequency for channel number 1, via channel 2 centered at 2417 MHz and 2457 MHz as the center frequency for channel number 10, up to channel 11 centered at 2462 MHz. Each channel bandwidth is 22MHz, symmetrically (+ / -11 MHz) located around the center frequency. In the transmission path, first, the baseband signal (IF) is generated based on the data to be transmitted, using 256 QAM (Quadrature Amplitude Modulation) based OFDM (Orthogonal Frequency Division Multiplexing) modulation technique, resulting in a 22 MHz (single channel wide) frequency band signal. The signal is then up-converted to the 2.4 GHz (RF) and placed in the center frequency of the required channel, and transmitted to the air via the antenna. Similarly, the receiving path comprises a received channel in the RF spectrum, down-converted to the baseband (IF) wherein the data is then extracted.

[0229] In order to support multiple devices and use a permanent solution, a Wireless Access Point (WAP) is typically used. A Wireless Access Point (WAP, or Access Point - AP) is a device that allows wireless devices to connect to a wired network using Wi-Fi, or related standards. The WAP usually connects to a router (via a wired network) as a standalone device, but can also be an integral component of the router itself. Using Wireless Access Point (AP) allows users to add devices that access the network with little or no cables. A WAP normally connects directly to a wired Ethernet connection, and the AP then provides wireless connections using radio frequency links for other devices to utilize that wired connection. Most APs support the connection of multiple wireless devices to one wired connection. Wireless access typically involves special security considerations, since any device within a range of the WAP can attach to the network. The most common solution is wireless traffic encryption. Modem access points come with built-in encryption such as Wired Equivalent Privacy (WEP) and Wi-Fi Protected Access (WPA), typically used with a password or a passphrase. Authentication in general, and a WAP authentication in particular, is used as the basis for authorization, which determines whether a privilege may be granted to a particular user or process, privacy, which keeps information from becoming known to non-participants, and non-repudiation, which is the inability to deny having done something that was authorized to be done based on the authentication. An authentication in general, and a WAP authentication in particular, may use an authentication server that provides a network service that applications may use to authenticate the credentials, usually account names and passwords of their users. When a client submits a valid set of credentials, it receives a cryptographic ticket that it can subsequently be used to access various services. Authentication algorithms include passwords, Kerberos, and public key encryption.

[0230] Prior art technologies for data networking may be based on single carrier modulation techniques, such as AM (Amplitude Modulation), FM (Frequency Modulation), and PM (Phase Modulation), as well as bit encoding techniques such as QAM (Quadrature Amplitude Modulation) and QPSK (Quadrature Phase Shift Keying). Spread spectrum technologies, to include both DSSS (Direct Sequence Spread Spectrum) and FHSS (Frequency Hopping Spread Spectrum) are known in the art. Spread spectrum commonly employs Multi-Carrier Modulation (MCM) such as OFDM (Orthogonal Frequency Division Multiplexing). OFDM and other spread spectrum are commonly used in wireless communication systems, particularly in WLAN networks.

[0231] Bluetooth. Bluetooth is a wireless technology standard for exchanging data over short distances (using short-wavelength UHF radio waves in the ISM band from 2.4 to 2.485 GHz) from fixed and mobile devices, and building personal area networks (PANs). It can connect several devices, overcoming problems of synchronization. A Personal Area Network (PAN) may be according to, compatible with, or based on, Bluetooth™ or IEEE 802.15.1-2005 standard. A Bluetooth controlled electrical appliance is described in U.S. Patent Application No. 2014 / 0159877 to Huang entitled: “Bluetooth Controllable Electrical Appliance” , and an electric power supply is described in U.S. Patent Application No. 2014 / 0070613 to Garb et al. entitled: “Electric Power Supply and Related Methods”, which are both incorporated in their entirety for all purposes as if fully set forth herein. Any Personal Area Network (PAN) may be according to, compatible with, or based on, Bluetooth™ or IEEE 802.15.1-2005 standard. A Bluetooth controlled electrical appliance is described in U.S. Patent Application No. 2014 / 0159877 to Huang entitled: “Bluetooth Controllable Electrical Appliance” , and an electric power supply is described in U.S. Patent Application No. 2014 / 0070613 to Garb et al. entitled: “Electric Power Supply and Related Methods”, which are both incorporated in their entirety for all purposes as if fully set forth herein.

[0232] Bluetooth operates at frequencies between 2,402 and 2,480 MHz, or 2,400 and 2,483.5 MHz including guard bands 2 MHz wide at the bottom end and 3.5 MHz wide at the top. This is in the globally unlicensed (but not unregulated) Industrial, Scientific and Medical (ISM) 2.4 GHz short-range radio frequency band. Bluetooth uses a radio technology called frequency-hopping spread spectrum. Bluetooth divides transmitted data into packets, and transmits each packet on one of 79 designated Bluetooth channels. Each channel has a bandwidth of 1 MHz. It usually performs 800 hops per second, with Adaptive Frequency- Hopping (AFH) enabled. Bluetooth low energy uses 2 MHz spacing, which accommodates 40 channels. Bluetooth is a packet-based protocol with a master-slave structure. One master may communicate with up to seven slaves in a piconet. All devices share the master's clock. Packet exchange is based on the basic clock, defined by the master, which ticks at 312.5 ps intervals. Two clock ticks make up a slot of 625 ps, and two slots make up a slot pair of 1250 ps. In the simple case of single-slot packets the master transmits in even slots and receives in odd slots. The slave, conversely, receives in even slots and transmits in odd slots. Packets may be 1, 3, or 5 slots long, but in all cases the master's transmission begins in even slots and the slave’s in odd slots.

[0233] Bluetooth Eow Energy. Bluetooth low energy (Bluetooth EE, BLE, marketed as Bluetooth Smart) is a wireless personal area network technology designed and marketed by the Bluetooth Special Interest Group (SIG) aimed at applications in the healthcare, fitness, beacons, security, and home entertainment industries. Compared to Classic Bluetooth, Bluetooth Smart is intended to provide considerably reduced power consumption and cost while maintaining a similar communication range. Bluetooth low energy is described in a Bluetooth SIG published Dec. 2, 2014 standard Covered Core Package version: 4.2, entitled: “Master Table of Contents & Compliance Requirements - Specification Volume 0”, and in an article published 2012 in Sensors [ISSN 1424-8220] by Carles Gomez et al. [Sensors 2012, 12, 11734-11753; doi:10.3390 / sl20211734] entitled: “Overview and Evaluation of Bluetooth Low Energy: An Emerging Low-Power Wireless Technology” , which are both incorporated in their entirety for all purposes as if fully set forth herein.

[0234] Bluetooth Smart technology operates in the same spectrum range (the 2.400 GHz- 2.4835 GHz ISM band) as Classic Bluetooth technology, but uses a different set of channels. Instead of the Classic Bluetooth 79 1-MHz channels, Bluetooth Smart has 40 2-MHz channels. Within a channel, data is transmitted using Gaussian frequency shift modulation, similar to Classic Bluetooth's Basic Rate scheme. The bit rate is IMbit / s, and the maximum transmit power is 10 mW. Bluetooth Smart uses frequency hopping to counteract narrowband interference problems. Classic Bluetooth also uses frequency hopping but the details are different; as a result, while both FCC and ETSI classify Bluetooth technology as an FHSS scheme, Bluetooth Smart is classified as a system using digital modulation techniques or a direct- sequence spread spectrum. All Bluetooth Smart devices use the Generic Attribute Profile (GATT). The application programming interface offered by a Bluetooth Smart aware operating system will typically be based around GATT concepts.

[0235] Cellular. Cellular telephone network may be according to, compatible with, or may be based on, a Third Generation (3G) network that uses UMTS W-CDMA, UMTS HSPA, UMTS TDD, CDMA2000 IxRTT, CDMA2000 EV-DO, or GSM EDGE-Evolution. The cellular telephone network may be a Fourth Generation (4G) network that uses HSPA+, Mobile WiMAX, LTE, LTE- Advanced, MBWA, or may be based on or compatible with IEEE 802.20- 2008.

[0236] Mobile systems and methods that overcomes the deficiencies of prior art speech-based interfaces for telematics applications through the use of a complete speech-based information query, retrieval, presentation and local or remote command environment, are disclosed in U.S. Patent No. 7,693,720 to Kennewick et al., entitled: “Mobile systems and methods for responding to natural language speech utterance”, which is incorporated in its entirety for all purposes as if fully set forth herein. This environment makes significant use of context, prior information, domain knowledge, and user specific profile data to achieve a natural environment for one or more users making queries or commands in multiple domains. Through this integrated approach, a complete speech-based natural language query and response environment can be created. The invention creates, stores and uses extensive personal profile information for each user, thereby improving the reliability of determining the context and presenting the expected results for a particular question or command. The invention may organize domain specific behavior and information into agents, that are distributable or updateable over a wide area network. The invention can be used in dynamic environments such as those of mobile vehicles to control and communicate with both vehicle systems and remote systems and devices.

[0237] Systems and methods for receiving natural language queries and / or commands and execute the queries and / or commands are disclosed in U.S. Patent No. 8,015,006 to Kennewick et al., entitled: “Systems and methods for processing natural language speech utterances with context-specific domain agents”, which is incorporated in its entirety for all purposes as if fully set forth herein. The systems and methods overcome the deficiencies of prior art speech query and response systems through the application of a complete speech-based information query, retrieval, presentation and command environment. This environment makes significant use of context, prior information, domain knowledge, and user specific profile data to achieve a natural environment for one or more users making queries or commands in multiple domains. Through this integrated approach, a complete speech-based natural language query and response environment can be created. The systems and methods create, store and use extensive personal profile information for each user, thereby improving the reliability of determining the context and presenting the expected results for a particular question or command.

[0238] A mobile system that includes speech-based and non- speech-based interfaces for telematics applications is disclosed in U.S. Patent No. 8,195,468 to Weider et al., entitled: “Mobile systems and methods of supporting natural language human-machine interactions”, which is incorporated in its entirety for all purposes as if fully set forth herein. The mobile system identifies and uses context, prior information, domain knowledge, and user specific profile data to achieve a natural environment for users that submit requests and / or commands in multiple domains. The invention creates, stores and uses extensive personal profile information for each user, thereby improving the reliability of determining the context and presenting the expected results for a particular question or command. The invention may organize domain specific behavior and information into agents, that are distributable or updateable over a wide area network.

[0239] A speakerphone system integrated in a mobile device that is automatically controlled based on the current state of the mobile device is disclosed in U.S. Patent No. 8,676,224 to Louch entitled: “Speakerphone control for mobile device”, which is incorporated in its entirety for all purposes as if fully set forth herein. In one implementation, the mobile device is controlled based on an orientation or position of the mobile device. In another implementation, the control of the speakerphone includes automatically controlling one or more graphical user interfaces associated with the speakerphone system.

[0240] Systems, methods, and computer-readable media for voice-based determination of physical and emotional characteristics of users are disclosed in U.S. Patent No. 10,096,319 to Jin et al. entitled: “Voice-based determination of physical and emotional characteristics of users”, which is incorporated in its entirety for all purposes as if fully set forth herein. Example methods may include determining first voice data, wherein the first voice data is generated by a user, determining a first real-time user status of the user using the first voice data, generating a first data tag indicative of the first real-time user status, determining first audio content for presentation at a speaker device using the first data tag and the first voice data, and causing presentation of the first audio content via a speaker of the speaker device.

[0241] A cooperative conversational voice user interface is provided in U.S. Patent No. 8.073,681 to Baldwin et al. entitled: “System and method for a cooperative conversational voice user interface''', which is incorporated in its entirety for all purposes as if fully set forth herein. The cooperative conversational voice user interface may build upon short-term and long-term shared knowledge to generate one or more explicit and / or implicit hypotheses about an intent of a user utterance. The hypotheses may be ranked based on varying degrees of certainty, and an adaptive response may be generated for the user. Responses may be worded based on the degrees of certainty and to frame an appropriate domain for a subsequent utterance. In one implementation, misrecognitions may be tolerated, and conversational course may be corrected based on subsequent utterances and / or responses.

[0242] A system and method for selecting and presenting advertisements based on natural language processing of voice-based inputs is provided in U.S. Patent No. 7,818,176 to Freeman et al. entitled: “System and method for selecting and presenting advertisements based on natural language processing of voice-based input”, which is incorporated in its entirety for all purposes as if fully set forth herein. A user utterance may be received at an input device, and a conversational, natural language processor may identify a request from the utterance. At least one advertisement may be selected and presented to the user based on the identified request. The advertisement may be presented as a natural language response, thereby creating a conversational feel to the presentation of advertisements. The request and the user's subsequent interaction with the advertisement may be tracked to build user statistical profiles, thus enhancing subsequent selection and presentation of advertisements.

[0243] Any step or method herein may be based on, or may use, any of the documents mentioned in, or incorporated herein, the background part of this document. Each of the methods or steps herein, may consist of, include, be part of, be integrated with, or be based on, a part of, or the whole of, the steps, functionalities, or structure (such as software) described in the publications that are incorporated in their entirety herein. Further, each of the components, devices, or elements herein may consist of, integrated with, include, be part of, or be based on, a part of, or the whole of, any components, systems, devices or elements described in the publications that are incorporated in their entirety herein.

[0244] In one example, it would be an advancement in the art to provide methods and systems for detection of non-verbal messages in speech that may be used for disentangling and flagging non-verbal spoken messages in speech that occur simultaneously, that are simple, intuitive, require low communication bandwidth, small, secure, cost-effective, reliable, provide lower power consumption, provide lower CPU and / or memory usage, easy to use, reduce latency, faster, has a minimum part count, minimum hardware, and / or uses existing and available components, protocols, programs, and applications for providing better quality of service, better or optimal resources allocation, or provides a better user experience.

[0245] SUMMARY

[0246] A method herein may be used for identifying non-verbal information by prosodic multilayered analysis of intonation units. The method may comprise capturing, by a first microphone, a speech or a chunk thereof (in the English language or in a non-English language) that vocalizes text data which comprises multiple words; providing or transmitting the captured speech to a trained weakly- supervised deep learning acoustic model; and generating, by the trained model, an output data. The output data may comprise the multiple words, identification of multiple Intonation Units (IUS), and one or more labels. Each of the identified multiple IUS may comprise one or more words of the multiple words, and at least one of the IUs may be associated with a label.

[0247] Alternatively or in addition, a method herein may be used for training a model to identify non-verbal information by prosodic multilayered analysis of multiple intonation units, and the method may comprise capturing, by any microphone, such as a first microphone, a speech that may vocalize text data that may comprise multiple words; sounding, by any sounder to a person, the speech captured by any microphone, such as by the first microphone; obtaining, automatically or manually, the text data or the multiple words; obtaining, by any input component, such as a first input component, from the person, in response to the sounding, an identification of multiple Intonation Units (IUs) and multiple labels associated with the identified IUs. Any IU herein, such as each of the lus, may comprise one or more words from the text data or multiple words; and any method may comprise providing the speech captured by the first microphone, the obtained identified IUs, the obtained associated multiple labels, and the obtained text data or multiple words, to a trained weakly-supervised deep learning acoustic model for training the model. Any training herein may be repeated. Any providing herein may comprise providing of the obtained multiple labels and the obtained text data as a single structured data stream that uses a structure.

[0248] Any duration herein of any capturing of a speech may be at least 0.1, 0.2, 0.3, 0.5, 0.7, 1.0, 1.1, 1.2, 1.5, 1.7, 2, 3, 5, 10, 12, 15, 20, 30, 50, or 100 second, or at least 1, 2, 3, 5, 10, 12, 15, 20, 30, 50, 100, 120, 150, 200, 300, 500, or 1,000 minutes. Alternatively or in addition, any duration herein of any capturing of a speech may be no more than 0.1, 0.2, 0.3, 0.5, 0.7, 1.0, 1.1, 1.2, 1.5, 1.7, 2, 3, 5, 10, 12, 15, 20, 30, 50, or 100 second, or at least 1, 2, 3, 5, 10, 12, 15, 20, 30, 50, 100, 120, 150, 200, 300, 500, or 1,000 minutes. At least one of any IUS herein may comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words, or each one of any IUs herein may comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words. Any label herein may be associated with a single word of any IU. Alternatively or in addition, each one of, or at least one of, any of the IUs herein may comprise less than 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words. A duration of at least one of, or of each one of, any of the IUs herein may be at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long, or may be less than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long.

[0249] Alternatively or in addition, any IU herein may comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words, and any label herein may be further associated with a single word of any IU. Alternatively or in addition, any IU herein may be associated with multiple distinct labels, such as at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 multiple distinct labels, that may be in the same category, or alternatively in different categories, or any combination thereof.

[0250] Any method herein may further comprise extracting, using any Speech-To-Text (STT) scheme, any text data from any speech captured herein, such as by the first microphone. Alternatively or in addition, any method herein may further comprise aligning any captured speech and any extracted text data. Alternatively or in addition, any method herein may further comprise providing any extracted text data to any trained weakly-supervised deep learning acoustic model, and any output data herein may be generated in response to any provided extracted text data. Any STT herein may comprise, may use, or may be based on, a Hidden Markov Models (HMMs), a Dynamic Time Warping (DTW), a Neural Network, a Deep feedforward Neural Network (DNN), Denoising Autoencoders, or any combination thereof.

[0251] Any text data herein may comprise one or more phrases, one or more sentences, one or more clauses, or any portion thereof. All the steps of any flow chart or a method herein may be performed in a single enclosure. Each of any identified IU herein may comprise a segment of the captured speech bounded by two boundaries that are characterized by a threshold speech-rate deviation. Any one or more of the steps herein, or any model herein, may comprise, may be provided as, may use, or may be interfaced by, a Software Development Kit (SDK). Alternatively or in addition, any one or more of the steps herein, or any model herein, may comprise, may be provided as, may use, or may be interfaced by, an Application Programming Interface (API). A non-transitory computer readable medium may include computer executable instructions stored thereon, and the instructions may include any of the steps, methods, or processes herein, or any model herein (or any part thereof). At least one of any method, flow charts, or steps herein, may comprise executing, by a processor, a software or firmware that may be stored in a computer readable medium.

[0252] Any microphone herein, such as the first microphone, may comprise an omnidirectional, unidirectional, or bidirectional microphone, that may be based on the sensing an incident soundbased motion of a diaphragm or a ribbon. Alternatively or in addition, any microphone herein, such as the first microphone, may comprise a condenser, an electret, a dynamic, a ribbon, a carbon, or a piezoelectric microphone. Any speech capturing herein may use multiple microphones that may include the first microphone. Any multiple microphones herein may be arranged as a directional microphones array that may be operative to estimate a number, magnitude, frequency, Direction-Of-Arrival (DOA), distance, or speed of a phenomenon impinging the microphones array. Any method or flow chart herein may comprise converting, such as by an Analog-to-Digital (A / D) converter coupled to a microphone, such as the first microphone, any captured speech to a digital data stream. Any digital data stream herein may comprise a compressed or uncompressed digital data stream, that may be in a format of Waveform Audio File Format (WAV), Windows Media Audio (WMA), MP3, Pulse-Code Modulation (PCM) format stream, or a Linear Pulse-Code Modulation (LPCM) format stream.

[0253] Any method herein may further comprise manipulating any captured speech, and any providing herein of any captured speech may comprise providing any manipulated captured speech. Any manipulating herein may comprise filtering of a noise or a background sound, and enhancing or reducing speech-related features. Alternatively or in addition, any manipulating herein may comprise amplifying, filtering, converting, range matching, resolution increasing, integrating, deviating, equalizing, compressing, de-compressing, coding, decoding, modulating, demodulating, pattern recognizing, smoothing, or any combination thereof. Alternatively or in addition, any manipulating herein may comprise performing a function that may be based on, may use, or may comprise, a discrete, continuous, monotonic, non-mono tonic, elementary, algebraic, linear, polynomial, quadratic, Cubic, Nth-root based, exponential, transcendental, quintic, quartic, logarithmic, hyperbolic, or trigonometric function.

[0254] Any manipulating herein, such as any manipulation of any captured speech as part of any pre-processing scheme, may comprise, may use, or may be based on, applying a feature engineering technique, that may comprise, may use, or may be based on, creation, transformation, extraction, selection, or any combination thereof, of one or more features or variables. Alternatively or in addition, any feature engineering technique herein may comprise, may use, or may be based on, Imputation; Categorical Imputation; Numerical Imputation; Discretization; Categorical encoding; Splitting; Outliers removal, Outliers values replacing; capping the maximum and minimum Outliers values; Variable transformation; Scaling; Min- Max Scaling; Standardization / Variance Scaling; Feature Creation; creating interaction features; dimensionality reduction; logarithmic or power variables transformations; feature selection; or any combination thereof.

[0255] Any method, scheme, or a flow chart herein, may further comprise storing of any captured speech in a memory, and may further comprise retrieving the captured speech from the memory. Any providing of any captured speech herein may comprise providing the retrieved captured speech. Any method, scheme, or a flow chart herein, may further comprise storing the output data in a memory, and any method, scheme, or a flow chart herein may further comprise storing the captured speech in the memory, associated with the output data.

[0256] Any method, scheme, or a flow chart herein may be preceded by training any model for a duration of at least 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100 or 200 minutes, or at least 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100 or 200 hours, or by training any model for a duration of less than 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100, or 200 minutes, or at least 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100 or 200 hours. Alternatively or in addition, any method, scheme, or a flow chart herein may be preceded by training any model by at least 100, 200, 300, 500, 700, 1,000, 1,200, 1,500, 2,000, 2,500, 3,000, 5,000, 7,000, 10,000, 20,000, 30,000, 40,000, 50,000, 70,000, or 100,000 words or IUS, or by less than 100, 200, 300, 500, 700, 1,000, 1,200, 1,500, 2,000, 2,500, 3,000, 5,000, 7,000, 10,000, 20,000, 30,000, 40,000, 50,000, 70,000, or 100,000 words or IUs.

[0257] Any model herein may comprise, may use, or may be based on, a machine learning model that may be primarily configured for, or trained for, speech recognition, transcription, translation, or any combination thereof, and that may consist of, may use, may comprise, or may be based on, an encoder-decoder Transformer architecture.

[0258] Alternatively or in addition, any model herein may comprise, may use, or may be based on, an open-source software that may comprise, may use, or may be based on, Mel-frequency Cepstrum or a log-Mel spectrogram. Alternatively or in addition, any model herein may comprise, may use, or may be based on, sinusoidal positional encoding or intermixing with learned positional encoding, such as Whisper model or architecture by OpenAI.

[0259] Any data herein, such as any output data or any part thereof herein, may comprise or may be based on, an annotated text, that may be in a plain text, annotated text, JavaScript Object Notation (JSON) format, Voice Extensible Markup Language (VoiceXML), an Extensible Markup Language (XML), Pronunciation Lexicon Specification (PLS), a Comma-Separated Values (CSV), a Structured Query Language (SQL), or Java Speech API Markup Language (JSML).

[0260] Alternatively or in addition, any method herein may further comprise converting any data herein, such as any output data or any part thereof, to a format that may use, or may be based on, a plain text, an annotated text, a Voice Extensible Markup Language (VoiceXML), a VoiceXML (VXML), an Extensible Markup Language (XML), JavaScript Object Notation (JSON) format, a Pronunciation Lexicon Specification (PLS), a Comma-Separated Values (CSV), a Structured Query Language (SQL), or a Java Speech API Markup Language (JSML). Any data herein, such as any output data or any part thereof herein, may comprise or may be based on, an annotated text that may use, or may be based on, to an annotated text in a format that may use, or may be based on, Tones and Break Indices (ToBI) format, INCEpTION, or Emotion Markup Language (EML or EmotionML). Alternatively or in addition, any method herein may further comprise converting any data herein, such as any output data or any part thereof, to an annotated text in a format that may use, or may be based on, to Tones and Break Indices (ToBI) format, INCEpTION, or Emotion Markup Language (EML or EmotionML).

[0261] Alternatively or in addition, any data herein, such as any output data herein, may comprise multiple labels, and each of the multiple labels may be part of a ‘Genre’ category, a ‘Prototype’ category, a ‘Discourse Function’ category, an ‘Emotion’ category, an ‘Emphasis’ category, or an ‘Attitude’ category.

[0262] Any data herein, such as any output data herein, may comprise a first label that may be in a genre category, may be associated with part of, or whole of, any captured speech, and may be associated with multiple words, a sentence, multiple sentences, a clause, a paragraph in the text data, or any combination thereof. Any label herein, such as any first label, may be associated with a sentence in the text data, and may comprise exclamation, request, command, or suggestion. Alternatively or in addition, any label herein, such as any first label, may be associated with a rhetorical mode, or may comprise narration, description, exposition, and argumentation. Alternatively or in addition, any label herein, such as any first label, may comprise an informal addressing, formal addressing, a narrative, a quotation, a casual conversation, a professional exchange, a debate, a testimonial, an instructional (such as a tutorial), an interview, an announcement, a reportage, or any combination thereof. Alternatively or in addition, any label herein may comprise a gender that may be ‘male’ or ‘female’, an age of the speaker, or both. Alternatively or in addition, any label herein may comprise a naturality indicator that may be ‘human’ or ‘machine’, respectively indicating human or machine generated speech.

[0263] Any data herein, such as any output data herein, may comprise a label, such as a first label, that may be in a prototype category that may identify a boundary in the captured speech or in the text data, and may be associated with a single IU. Alternatively or in addition, any data herein, such as any output data herein, may comprise at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the prototype category. Any label herein, such as the first label, may be associated with a tone change, or wherein the first label is associated with a punctuation, may be associated with a flattish tone, a falling tone, a rising tone, or a disfluency, in the captured speech, may be associated with a comma, a period, a question mark, or a truncated text, or may comprise a ‘continuation’, a ‘conclusion’, a ‘question’, or a ‘truncated’.

[0264] Any data herein, such as any output data herein, may comprise a first label that may be in an emphasis category that may identify, may emphasize, may intensify, or may otherwise point out a single IU. Any data herein, such as any output data herein, may comprise at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the emphasis category. Any label herein, such as the first label, may comprise a ‘contrastive’, a ‘strong’, a ‘weak’, or a ‘de-emphasis’.

[0265] Any data herein, such as any output data herein, may comprise a label, such as a first label, that may be in an emotion category that may involve intentional and non-intentional emotion or feeling exhibited in a single IU. Any data herein, such as any output data herein, may comprise at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the emotion category. Any label herein, such as the first label, may comprise an emotion that may be part of “Negative and forceful” class, such as Anger; Annoyance; Contempt; Disgust; or Irritation. Alternatively or in addition, any label herein, such as the first label, may comprise an emotion that is part of “Negative and not in control” class, such as Anxiety; Embarrassment; Fear; Helplessness; Powerlessness; or Worry. Alternatively or in addition, any label herein, such as the first label, may comprise an emotion that may be part of “Negative thoughts” class, such as Doubt; Envy; Frustration; Guilt; or Shame. Alternatively or in addition, any label herein, such as the first label, may comprise an emotion that is part of “Negative and passive” class, such as Boredom; Despair; Disappointment; Hurt; or Sadness. Alternatively or in addition, any label herein, such as the first label, may comprise an emotion that is part of an “Agitation” class, such as Stress; Shock; or Tension. Alternatively or in addition, any label herein, such as the first label, may comprise an emotion that may be part of a “Positive and lively” class, such as Amusement; Delight; Elation; Excitement; Happiness; Joy; or Pleasure. Alternatively or in addition, any label herein, such as the first label, may comprise an emotion that may be part of a “Caring” class, such as Affection; Empathy; Friendliness; or Love.

[0266] Alternatively or in addition, any label herein, such as the first label, may comprise an emotion that may be part of a “Positive thoughts” class, such as Pride; Courage; Hope; Humility; Satisfaction; or Trust. Alternatively or in addition, any label herein, such as the first label, may comprise an emotion that may be part of a “Quiet positive” class, such as Calmness; Contentment; Relaxation; Relief; or Serenity. Alternatively or in addition, any label herein, such as the first label, may comprise an emotion that may be part of a “Reactive” class, such as Politeness; or Surprise. Alternatively or in addition, any label herein, such as the first label, may comprise Disgusted; Apprehended; Hesitant; Angry; Delighted; Happy; Content; Upset; Nervous; Insecure; Confused; Enthusiastic Frustrated; Relieved; Sad; Hopeful; Jealous; Content; Anxious; Curious; Desperate; Optimistic; Pessimistic; or any combination thereof.

[0267] Any data herein, such as any output data herein, may comprise a label, such as a first label, that may be in an attitude category that may involve an attitude or sentiment about someone or something exhibited in a single IU. Any data herein, such as any output data herein, may comprise at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the attitude category. Alternatively or in addition, any label herein, such as the first label, may comprise Neutral; Negative; Positive; Hesitation (intentional or non-intentional); Inclusion; Exclusion; Solidarity or alignment; Distance; Irony; Sarcasm; Respect; Disrespect; Position of power or authority; Lack of power or authority; Modal distance; Reservation; Reserve; Reluctance; Self- deprecation; Self-humor; Shame; Indifference; Positive surprise; Negative surprise; Indignation; Protest; Admiration; Criticism; Apologetic; Agreement; Disagreement; Encouragement; Discouragement; Reassurance; Uncertainty; Certainty; Sympathy; Empathy; Anticipation; or any combination thereof. Alternatively or in addition, any label herein, such as the first label, may comprise a cognitive attitude, an affective attitude, or a conative attitude. Alternatively or in addition, any label herein, such as the first label, may comprise a negative or positive attitude.

[0268] Any data herein, such as any output data herein, may comprise a label, such as a first label, that may be in a discourse function category that may involve a relation between two or more elements in the captured speech, such as a relation between two or more IUS, two or more words, or two or more sentences. Any data herein, such as any output data herein, may comprise at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the discourse function category. Alternatively or in addition, any label herein, such as the first label, may comprise Open or closed list; Apposition; Rhetorical question; Request; Suggestion; Confirmation; Transition to another subject; Bi-partites (such as if / then; when / then; cause&effect; or contradiction (X but ¥)); Proclamation; Narration; Title; Main subject; Background; Elaboration; Conclusion; Recap; About to make a point; Making a point; Point made; Repetition; Apposition; Parentheticals; Digression; or any combination thereof.

[0269] Any method herein may be used with prosodic multilayered analysis of intonation units that may be designed for detecting non-verbal cues in speech. Once detected, these messages can obtain formal representations such as text, symbols, transformed audio, etc. Any method herein may further comprise adding multiple punctuation marks or symbols to any text data or any multiple words herein, such as in any output data. The adding may be based on, or may use, boundaries or locations of the identified multiple IUS, the associated labels, or any combination thereof. Any multiple punctuation marks or symbols herein may comprise a comma, a period (dot), a quotation mark, parenthesis marks, a dash mark, an exclamation mark, a colon mark, a semi-colon mark, or any combination thereof. Any multiple words or any the text data herein may be in a first language, and any method herein may further comprise translating any text data or any multiple words into a second language, such as by using or based on, boundaries or locations of any identified multiple IUs, any associated labels, or any combination thereof. Any method herein may further comprise embedding in any text or words herein, such as extracted text or words that may be part of any output data herein, any marks, symbols, pictograms, logograms, ideograms, or emojis, that may be based on, or may be responsive to, any labels herein, such as labels that are part of any output data herein. Any method herein may further comprise producing, using a speech synthesis or Text-To-Speech (TTS) scheme, a speech that may be annotated, combined, responsive to, or modified, with the associated labels. Any speech synthesis or Text-To-Speech (TTS) herein may use, or may be based on, a concatenative synthesis or formant synthesis,

[0270] Any microphone herein, such as any first microphone herein, may be housed in, may be attached to, or may be integrated with, a first device that may have a first enclosure, and any generating or inferring of the output data herein may be performed in a second device that may have a second enclosure.

[0271] Any device herein, such as any first device herein, may comprise a client device, and may comprise storing, operating, or using, by the client device, a client operating system. Any client operating system herein may consist of, may comprise, or may be based on, one out of Microsoft Windows 7, Microsoft Windows XP, Microsoft Windows 8, Microsoft Windows 8.1, Linux, and Google Chrome OS. Alternatively or in addition, any client operating system herein may be a mobile operating system that may comprise Android version 2.2 (Froyo), Android version 2.3 (Gingerbread), Android version 4.0 (Ice Cream Sandwich), Android Version 4.2 (Jelly Bean), Android version 4.4 (KitKat), Apple iOS version 3, Apple iOS version 4, Apple iOS version 5, Apple iOS version 6, Apple iOS version 7, Microsoft Windows® Phone version 7, Microsoft Windows® Phone version 8, Microsoft Windows® Phone version 9, or Blackberry® operating system. Alternatively or in addition, any client operating system herein may consist of, may comprise, or may be based on, is a Real-Time Operating System (RTOS), that may comprise FreeRTOS, SafeRTOS, QNX, VxWorks, or Micro-Controller Operating Systems (p,C / OS).

[0272] Any enclosure herein, such as any first enclosure herein, may comprise a hand-held enclosure or a portable enclosure. Any client device herein, may consist of, may comprise, may be part of, or may be integrated with, a notebook computer, a laptop computer, a media player, a Digital Still Camera (DSC), a Digital video Camera (DVC or digital camcorder), a Personal Digital Assistant (PDA), a cellular telephone, a digital camera, a video recorder, or a smartphone, that may comprise, or may be based on, an Apple iPhone 6 or a Samsung Galaxy S6.

[0273] Any device herein, such as any second device herein, may comprise, may consist of, or may be integrated with, a server device, that may be a dedicated device that manages network resoucres; may not be a client device and may not be not a consumer device; may be continuously online with greater availability and maximum up time to receive requests almost all of the time efficiently processes multiple requests from multiple client devices at the same time; may generate various logs associated with the client devices and traffic from / to the client devices; primarily interfaces and responds to requests from client devices; may have greater fault tolerance and higher reliability with lower failure rates; may provide scalability for increasing resources to serve increasing client demands; or any combination thereof.

[0274] Any device herein, such as any first or second device herein, may comprise, may consist of, or may be integrated with, a server device that may be storing, operating, or using, a server operating system. Any server operating system herein may consist or, may comprise, or may be based on, one out of Microsoft Windows Server®, Linux, or UNIX. Alternatively or in addition, any server operating system herein may consist or, may comprise, or may be based on, one out of Microsoft Windows Server® 2003 R2, 2008, 2008 R2, 2012, or 2012 R2 variant, Linux™ or GNU / Linux based Debian GNU / Linux, Debian GNU / kFreeBSD, Debian GNU / Hurd, Fedora™, Gentoo™, Linspire™, Mandriva, Red Hat® Linux, SuSE, and Ubuntu®, UNIX® variant Solaris™, AIX®, Mac™ OS X, FreeBSD®, OpenBSD, and NetBSD®.

[0275] Any server device herein may be cloud-based implemented as part of a public cloudbased service, such as provided by Amazon Web Services® (AWS®), Microsoft® Azure™, or Google® Compute Engine™ (GCP). Any obtaining any data herein such as any generating or inferring of any output data herein may be provided as an Infrastructure as a Service (laaS) or as a Software as a Service (SaaS).

[0276] Any method, step, process, or a flow chart herein may be used with any device, such as any first device herein, may be in a first enclosure, and any second device herein may be in a second enclosure that may communicate over a network. Any microphone herein may be housed in, may be attached to, or may be integrated with, any first device, any output data herein may be generated in any device herein, such as in any second device. Any two deices herein, such as the first and second devices herein, may be configured to communicate over the Internet. Any second device herein may comprise a server device, and any first device herein may comprise a client device, and any two devices, such as the first and second devices, may communicate using a client-server architecture.

[0277] Any two devices herein, such as the first and second devices, may communicate over a wireless network. Any method, scheme, step, or flow chart herein may comprise coupling, by a first antenna in the first device, to the wireless network; coupling, by a second antenna in the second device, to the wireless network; sending, by the first device using a first wireless transceiver that may be coupled to the first antenna, to the wireless network via the first antenna, a first data; and receiving, by the second device using a second wireless transceiver that may be coupled to the second antenna, from the wireless network via the second antenna, the first data. Any data herein, such as any first data herein, may comprise part of, or whole of, any captured speech, any representation thereof, or any function thereof.

[0278] Alternatively or in addition, any method, scheme, step, or flow chart herein may comprise sending, by the second device using the second wireless transceiver that may be coupled to the second antenna, a second data; and receiving, by the first device using the first wireless transceiver that may be coupled to the first antenna, the second data. Alternatively or in addition, any method, scheme, step, or flow chart herein may comprise sending, by the second device using the second wireless transceiver that may be coupled to the second antenna, a second data; and receiving, by a third device using a third wireless transceiver that may be coupled to a third antenna, the second data. Any data herein, such as any second data herein, may comprise part of, or whole of, any output data, any representation thereof, or any function thereof.

[0279] Any two devices herein, such as the first device and the second device, may communicates over a wired network. Alternatively or in addition, any method, scheme, step, or flow chart herein may comprise coupling, by a connector in the second device, to the wired network; transmitting, by a wired transceiver in the second device that is coupled to the connector, a first data to the wired network via the connector; and receiving, by the wired transceiver in the second device that is coupled to the connector, the first data from the wired network via the connector. Any data herein, such as any first data herein, may comprise part of, or whole of, the captured speech, any representation thereof, or any function thereof.

[0280] Any method, scheme, step, or flow chart herein may comprise storing a part of, or whole of, any output data herein, a representation thereof, or a function thereof, in a database. Alternatively or in addition, any method, scheme, step, or flow chart herein may comprise storing a part of, or whole of, any captured speech herein, a representation thereof, or a function thereof, in a database. Any database herein may be a relational database, such as a Structured Query Language (SQL) based. Any output data herein may be stored in any database as a record, that may be searchable using at least one label.

[0281] Any method, scheme, step, or flow chart herein may be preceded or may comprise by training, once or repetitively, any model herein to identify non-verbal information by prosodic multilayered analysis of intonation units. Any training herein may comprise capturing, by a second microphone, a speech that may vocalize text data that comprises multiple words; sounding, by a sounder to the person, the speech captured by the second microphone; obtaining, automatically or manually, the text data; obtaining, by a first input component from the person, in response to the sounding, an identification of multiple Intonation Units (IUS) and one or more labels associated with the identified IUs, wherein each of the IUs comprises one or more words from the text data; and providing the speech captured by the second microphone, the obtained identified IUs, the obtained associated one or more labels, and the obtained text data, to the model for training the model.

[0282] Any providing herein may comprise providing of any obtained multiple labels and any obtained text data as a single structured data stream that may use a structure. Any structure herein, such as of any single structured data stream herein, may comprise any obtained multiple labels and any obtained text data at pre-defined positions or locations in any data stream, such as a sequence of consecutive pairs, and each pair herein may comprise a part of the text data (such as one or more words) followed by an obtained label that may be associated with the preceding part of the text data in the pair. Any generating herein may comprise generating of any inference output data as a single structured data stream that may use, or may be based on, any structure herein, that may comprise a sequence of consecutive pairs. Each pair herein may comprise one or more of the multiple words followed by an obtained label that may be associated with the preceding one or more of the multiple words in the pair. Any generating herein may comprise generating of any inference output data as a single structured data stream that may comprise a sequence of consecutive pairs. Any pair herein may comprise one or more words from the multiple words followed by a label that may be associated with the preceding one or more words in the pair. Any acoustic model herein may use, may comprise, or may be based on, an asynchronous transformer that may be based on, or may use, an encoder-decoder Transformer architecture. Any single structured data stream herein may be generated by predicting tokens as outputs of one or more text decoders, and each token herein may be predicted based on former predicted or used tokens. Any method herein, or any generating herein, may comprise replacing, in each pair in any sequence, a predicted token that may be associated with a predicting of the one or more words from the multiple words, with predefined one or more words, that may comprise, or may be based on, one or more words that are transcribed from the captured speech, such as part of any training of the acoustic model.

[0283] Any method herein may further comprise checking whether the first label is included in a labels list, and responsive to the checking of whether the first label is included in the labels list, may perform a first action or a second action, where the actions may be identical to, similar to, or different from, each other. Any action herein, such as any first action, may be performed in response to determining that the first label may be included in the labels list. Alternatively or in addition, any action herein, such as any second action, may be performed in response to determining that the first label may not be included in the labels list. Any list herein, such as the labels list, may include at least 1, 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, or 1,000 labels. Alternatively or in addition, any list herein, such as the labels list, may include less than 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, or 2,000 labels.

[0284] Any method herein may further comprise checking whether the first category may be included in a categories list, and responsive to the checking whether the first category is included in the categories list, may perform a first action or a second action. Any categories list herein may include 1, 2, 3, 4, or 5 categories. Any action herein, such as any first action, may be performed in response to further determining that the first category may be included in the categories list. Alternatively or in addition, any action herein, such as any second action, may be performed in response to further determining that the first category may not be included in the categories list.

[0285] Any method herein may further comprise checking whether a first word of the multiple words that may associated with the first label may be included in a words list, and responsive to the checking whether the first word may be included in the words list, may perform a first action or a second action. Any action herein, such as any first action, may be performed in response to further determining that the first word may be included in the words list. Alternatively or in addition, any action herein, such as any second action, may be performed in response to further determining that the first word may not be included in the words list. Any words list herein may include at least 1, 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, 2,000, 5,000 or 10,000 words. Further, any words list herein may include less than 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, 2,000, 5,000, 10,000 or 20,000 words.

[0286] Alternatively or in addition, any method, scheme, step, or flow chart herein may comprise aligning the speech captured by the second microphone with the obtained identified IUS for synchronization thereof. Any providing herein to any model herein may comprise providing the aligned information. Any obtaining herein of the text data may comprise manually obtaining. Any training herein may comprise obtaining, by a second input component from the person, the text data. Alternatively or in addition, any obtaining herein of the text data may comprise obtaining, by a second input component from the person, the text data. Any second input device herein may be same as, or may be identical to...

Claims

CLAIMS1. A method for identifying non-verbal information by prosodic multilayered analysis of intonation units, the method comprising: capturing, by a first microphone, a speech that vocalize text data that comprises multiple words; providing the captured speech to a trained weakly- supervised deep learning acoustic model; and generating, by the trained acoustic model, an inference output data, wherein the output data comprise the multiple words, identification of first and second Intonation Units (IUS), a first label in a first category, and a second label in a second category, wherein each of the identified first and second IUs comprises one or more words of the multiple words, wherein the first IU is associated with the first label and the second IU is associated with the second label, and wherein the first and second categories are selected from a group that consists of a ‘Genre’ category, a ‘Prototype’ category, a ‘Discourse Function’ category, an ‘Emotion’ category, an ‘Emphasis’ category, and an ‘Attitude’ category.

2. The method according to any one of the preceding claims, wherein the speech is in the English language.

3. The method according to any one of the preceding claims, wherein the speech is in a nonEnglish language.

4. The method according to any one of the preceding claims, wherein a duration of the capturing of the speech is at least 0.1, 0.2, 0.3, 0.5, 0.7, 1.0, 1.1, 1.2, 1.5, 1.7, 2, 3, 5, 10, 12, 15, 20, 30, 50, or 100 second, or at least 1, 2, 3, 5, 10, 12, 15, 20, 30, 50, 100, 120, 150, 200, 300, 500, or 1,000 minutes.

5. The method according to any one of the preceding claims, wherein a duration of the capturing of the speech is no more than 0.1, 0.2, 0.3, 0.5, 0.7, 1.0, 1.1, 1.2, 1.5, 1.7, 2, 3, 5, 10, 12, 15, 20, 30, 50, or 100 second, or at least 1, 2, 3, 5, 10, 12, 15, 20, 30, 50, 100, 120, 150, 200, 300, 500, or 1,000 minutes.

6. The method according to any one of the preceding claims, wherein the first IUs comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words.

7. The method according to claim 6, wherein the first label is associated with a single word of the first IU.

8. The method according to any one of the preceding claims, wherein each one of the first and second IUS comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words.

9. The method according to any one of the preceding claims, wherein each one of the first and second IUs comprises less than 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words.

10. The method according to any one of the preceding claims, wherein a duration of at least one of the first and second IUs is at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long.

11. The method according to any one of the preceding claims, wherein a duration of each one of the first and second IUs is at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long.

12. The method according to any one of the preceding claims, wherein a duration of at least one of the first and second IUs is less than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long.

13. The method according to any one of the preceding claims, wherein a duration of each one of the first and second IUs is less than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long.

14. The method according to any one of the preceding claims, wherein the text data comprises one or more phrases, one or more sentences, or one or more clauses.

15. The method according to any one of the preceding claims, wherein all steps are performed in a single enclosure.

16. The method according to any one of the preceding claims, wherein one of, or each of, the identified first and second IUs comprises a segment of the captured speech bounded by two boundaries that are characterized by a threshold speech-rate deviation, pitch intensity pattern, pitch decay pattern, speech rate pattern, or speech rate decay pattern, or any combination thereof.

17. The method according to any one of the preceding claims, wherein one of, or each one of, the identified first and second IUs comprises a smallest speech unit that conveys non-verbal information or a label.

18. The method according to any one of the preceding claims, wherein one of the first and second IUs is associated with multiple distinct labels.

19. The method according to claim 18, wherein one of the first and second IUs is associated with at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 multiple distinct labels.

20. The method according to any one of the preceding claims, further comprising adding multiple punctuation marks or symbols to the multiple words in the output data.

21. The method according to claim 20, wherein the adding is based on, or uses, boundaries or locations of the identified first and second IUS, the associated first and second labels, or any combination thereof.

22. The method according to claim 20, wherein the multiple punctuation marks or symbols comprise a comma, a period (dot), a quotation mark, parenthesis marks, a dash mark, an exclamation mark, a colon mark, a semi-colon mark, or any combination thereof.

23. The method according to any one of the preceding claims, wherein the text data is in a first language, the method further comprising translating the multiple words into a second language using, or based on, boundaries or locations of the identified first and second IUs, the associated first and second labels, or any combination thereof.

24. The method according to any one of the preceding claims, further comprising embedding marks, symbols, pictograms, logograms, ideograms, or emojis that are based on, or responsive to, the first or second labels, in the multiple words in the output data.

25. The method according to any one of the preceding claims, further comprising sounding, by a sounder, using a speech synthesis or a Text-To-Speech (TTS) scheme, the multiple words in the output data being annotated, combined, responsive to, or modified, with the first and second labels.

26. The method according to claim 25, wherein the speech synthesis or the Text-To- Speech (TTS) scheme uses, or is based on, a concatenative synthesis or formant synthesis.

27. The method according to any one of the preceding claims, further comprising extracting, using a Speech-To-Text (STT) scheme, the text data from the speech captured by the first microphone.

28. The method according to claim 27, further comprising aligning the captured speech and the multiple words.

29. The method according to claim 27, further comprising providing the multiple words to the trained weakly- supervised deep learning acoustic model, and wherein the output data is generated in response to the provided extracted text data.

30. The method according to claim 27, wherein the STT comprises, uses, or is based on, a Hidden Markov Models (HMMs), a Dynamic Time Warping (DTW), a Neural Network, a Deep feedforward Neural Network (DNN), Denoising Autoencoders, or any combination thereof.

31. The method according to any one of the preceding claims, further comprising manipulating the captured speech, wherein the providing of the captured speech comprises providing the manipulated captured speech.

32. The method according to claim 31, wherein the manipulating comprises filtering of a noise, a music, or a background sound, and enhancing speech-related features.

33. The method according to claim 32, wherein the manipulating comprises amplifying, filtering, converting, range matching, resolution increasing, integrating, deviating, equalizing, compressing, de-compressing, coding, decoding, modulating, demodulating, pattern recognizing, smoothing, or any combination thereof.

34. The method according to claim 31, wherein the manipulating comprises performing a function that is based on, uses, or comprises, a discrete, continuous, monotonic, non-monotonic, elementary, algebraic, linear, polynomial, quadratic, Cubic, Nth-root based, exponential, transcendental, quintic, quartic, logarithmic, hyperbolic, or trigonometric function.

35. The method according to claim 31, wherein the manipulating comprises applying a feature engineering technique.

36. The method according to claim 35, wherein the feature engineering technique comprises, uses, or is based on, Imputation; Categorical Imputation; Numerical Imputation; Discretization; Categorical encoding; Splitting; Outliers removal, Outliers values replacing; capping the maximum and minimum Outliers values; Variable transformation; Scaling; Min-Max Scaling; Standardization / Variance Scaling; Feature Creation; creating interaction features; dimensionality reduction; logarithmic or power variables transformations; feature selection; or any combination thereof.

37. The method according to claim 35, wherein the feature engineering technique comprises, uses, or is based on, creation, transformation, extraction, selection, or any combination thereof, of one or more features or variables.

38. The method according to any one of the preceding claims, further comprising storing the captured speech in a memory.

39. The method according to claim 38, further comprising retrieving the captured speech from the memory, wherein the providing of the captured speech comprises providing the retrieved captured speech.

40. The method according to any one of the preceding claims, further comprising storing the output data in a memory.

41. The method according to claim 40, further comprising storing, in the memory, the captured speech being associated with the output data.

42. The method according to any one of the preceding claims, further preceded by training the model for a duration of at least 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100 or 200 minutes, or at least 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100 or 200 hours.

43. The method according to any one of the preceding claims, further preceded by training the model for a duration of less than 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100, or 200 minutes, or at least 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100 or 200 hours.

44. The method according to any one of the preceding claims, further preceded by training the model by at least 100, 200, 300, 500, 700, 1,000, 1,200, 1,500, 2,000, 2,500, 3,000, 5,000, 7,000, 10,000, 20,000, 30,000, 40,000, 50,000, 70,000, or 100,000 words, labels, or IUS.

45. The method according to any one of the preceding claims, further preceded by training the model by less than 100, 200, 300, 500, 700, 1,000, 1,200, 1,500, 2,000, 2,500, 3,000, 5,000, 7,000, 10,000, 20,000, 30,000, 40,000, 50,000, 70,000, or 100,000 words, labels, or IUs.

46. The method according to any one of the preceding claims, wherein the model comprises, uses, or is based on, a machine learning model that is primarily configured for, or trained for, speech recognition, transcription, or translation, and that uses, comprises, or is based on, an encoder-decoder Transformer architecture.

47. The method according to any one of the preceding claims, wherein the model comprises, uses, or is based on, an open-source software.

48. The method according to any one of the preceding claims, wherein the model comprises, uses, or is based on, Mel-frequency Cepstrum or a log-Mel spectrogram.

49. The method according to any one of the preceding claims, wherein the model comprises, uses, or is based on, a sinusoidal positional encoding or sinusoidal positional intermixing, with a learned positional encoding.

50. The method according to any one of the preceding claims, wherein the model comprises, uses, or is based on, Whisper model or architecture, by OpenAI.

51. The method according to any one of the preceding claims, wherein the output data or a part thereof comprises, uses, or is based on, a format that uses, or is based on, a plain text, an annotated text, a Voice Extensible Markup Language (VoiceXML), a VoiceXML (VXML), an Extensible Markup Language (XML), JavaScript Object Notation (JSON) format, a Pronunciation Lexicon Specification (PLS), a Comma-Separated Values (CSV), a Structured Query Language (SQL), or a Java Speech API Markup Language (JSML).51a. The method according to any one of the preceding claims, further comprising converting the output data or a part thereof to a format that that uses, or is based on, a plain text, an annotated text, a Voice Extensible Markup Language (VoiceXML), a VoiceXML (VXML), an Extensible Markup Language (XML), JavaScript Object Notation (JSON) format, a Pronunciation Lexicon Specification (PLS), a Comma-Separated Values (CSV), a Structured Query Language (SQL), or a Java Speech API Markup Language (JSML).5 lb. The method according to any one of the preceding claims, wherein the output data or a part thereof comprises, uses, or is based on, an annotated text that uses, or is based on, a format that uses, or is based on, Tones and Break Indices (ToBI), INCEpTION, or Emotion Markup Language (EML or EmotionML) format.51c. The method according to any one of the preceding claims, further comprising converting the output data or a part thereof to an annotated text in a format that uses, or is based on, Tones and Break Indices (ToBI) format, INCEpTION, or Emotion Markup Language (EML or EmotionML) format.

52. The method according to any one of the preceding claims, wherein the output data comprise multiple labels that include at least the first and second labels, and wherein each of the multiple labels is in a ‘Genre’ category, a ‘Prototype’ category, a ‘Discourse Function’ category, an ‘Emotion’ category, an ‘Emphasis’ category, or an ‘Attitude’ category.

53. The method according to any one of the preceding claims, wherein the first label is in the Genre category.

54. The method according to claim 53, wherein the first label is associated with part of, or whole of, the captured speech.

55. The method according to claim 53, wherein the first label is associated with multiple words, a sentence, multiple sentences, a clause, or a paragraph, in the text data of the output data.

56. The method according to claim 55, wherein the first label is associated with a sentence in the multiple words, and wherein the first label comprises exclamation, request, command, or suggestion.

57. The method according to claim 55, wherein the first label is associated with a rhetorical mode.

58. The method according to claim 57, wherein the first label comprises narration, description, exposition, argumentation, or any combination thereof.

59. The method according to claim 53, wherein the first label comprises an informal addressing, formal addressing, a narrative, a quotation, a casual conversation, a professional exchange, a debate, a testimonial, an instructional, a tutorial, an interview, an announcement, a reportage, or any combination thereof.

60. The method according to claim 53, wherein the first label comprises a gender that is ‘male’ or ‘female’, wherein the first label comprises a dialect type, an accent type, or an estimated age.

61. The method according to claim 53, wherein the first label comprises a naturality indicator that is ‘human’ or ‘machine’, respectively indicating human or machine generated speech.

62. The method according to any one of the preceding claims, wherein the first label is in the prototype category that identifies a boundary in the captured speech or in the text data of the output data, and wherein the first label is associated with the first IU.

63. The method according to claim 62, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the prototype category.

64. The method according to claim 62, wherein the first label is associated with a tone change, or wherein the first label is associated with a punctuation.

65. The method according to claim 64, wherein the first label is associated with a flattish tone, a falling tone, a rising tone, or a disfluency, in the captured speech.

66. The method according to claim 64, wherein the first label is associated with a comma, a period, a question mark, or a truncated text, in the text data of the output data.

67. The method according to claim 62, wherein the first label comprises a ‘continuation’, a ‘conclusion’, a ‘question’, or a ‘truncated’ in the text data of the output data.

68. The method according to any one of the preceding claims, wherein the first label is in the emphasis category that identifies, emphasizes, intensifies, or otherwise pointed out, in the first IU, or a word in the first IU.

69. The method according to claim 68, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the emphasis category.

70. The method according to claim 68, wherein the first label comprises a ‘contrastive’, a ‘strong’, a ‘weak’, or a ‘de-emphasis’.

71. The method according to any one of the preceding claims, wherein the first label is in the emotion category exhibited in the first IU.

72. The method according to claim 71, wherein the first label involves intentional and non- intentional emotion or feeling exhibited in the first IU.

73. The method according to claim 71, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the emotion category.

74. The method according to claim 71, wherein the first label comprises an emotion that is part of a “Negative and forceful” class.

75. The method according to claim 74, wherein the first label comprises Anger; Annoyance; Contempt; Disgust; or Irritation.

76. The method according to claim 71, wherein the first label comprises an emotion that is part of a “Negative and not in control” class.

77. The method according to claim 76, wherein the first label comprises Anxiety; Embarrassment; Fear; Helplessness; Powerlessness; or Worry.

78. The method according to claim 71, wherein the first label comprises an emotion that is part of a “Negative thoughts” class.

79. The method according to claim 78, wherein the first label comprises Doubt; Envy; Frustration; Guilt; or Shame.

80. The method according to claim 71, wherein the first label comprises an emotion that is part of a “Negative and passive” class.

81. The method according to claim 80, wherein the first label comprises Boredom; Despair; Disappointment; Hurt; or Sadness.

82. The method according to claim 71, wherein the first label comprises an emotion that is part of an “Agitation” class.

83. The method according to claim 82, wherein the first label comprises Stress; Shock; or Tension.

84. The method according to claim 71, wherein the first label comprises an emotion that is part of a “Positive and lively” class.

85. The method according to claim 84, wherein the first label comprises Amusement; Delight; Elation; Excitement; Happiness; Joy; or Pleasure.

86. The method according to claim 71, wherein the first label comprises an emotion that is part of a “Caring” class.

87. The method according to claim 86, wherein the first label comprises Affection; Empathy; Friendliness; or Love.

88. The method according to claim 71, wherein the first label comprises an emotion that is part of a “Positive thoughts” class.

89. The method according to claim 88, wherein the first label comprises Pride; Courage; Hope; Humility; Satisfaction; or Trust.

90. The method according to claim 71, wherein the first label comprises an emotion that is part of a “Quiet positive” class.

91. The method according to claim 90, wherein the first label comprises Calmness; Contentment; Relaxation; Relief; or Serenity.

92. The method according to claim 71, wherein the first label comprises an emotion that is part of a “Reactive” class.

93. The method according to claim 92, wherein the first label comprises Interest; Politeness; or Surprise.

94. The method according to claim 93, wherein the first label comprises Disgusted;Apprehended; Hesitant; Angry; Delighted; Happy; Content; Upset; Nervous; Insecure;Confused; Enthusiastic Frustrated; Relieved; Sad; Hopeful; Jealous; Content; Anxious; Curious; Desperate; Optimistic; Pessimistic; or any combination thereof.

95. The method according to any one of the preceding claims, wherein the first label is in the attitude category that involves an attitude or sentiment about someone or something exhibited in the first IU.

96. The method according to claim 95, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the attitude category.

97. The method according to claim 95, wherein the first label comprises Neutral; Negative; Positive; Hesitation (intentional or non-intentional); Inclusion; Exclusion; Solidarity or alignment; Distance; Irony; Sarcasm; Respect; Disrespect; Position of power or authority; Lack of power or authority; Modal distance; Reservation; Reserve; Reluctance; Self-deprecation; Self-humor; Shame; Indifference; Positive surprise; Negative surprise; Indignation; Protest; Admiration; Criticism; Apologetic; Agreement; Disagreement; Encouragement; Discouragement; Reassurance; Uncertainty; Certainty; Sympathy; Empathy; Anticipation; or any combination thereof.

98. The method according to claim 96, wherein the first label corresponds to a cognitive attitude, an affective attitude, or a conative attitude, or wherein the first label corresponds to a negative or positive attitude.

99. The method according to any one of the preceding claims, wherein the first label is in the discourse function category that involves a relation between two or more IUS or elements in the captured speech.

100. The method according to claim 99, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the discourse function category.

101. The method according to claim 99, wherein the relation is between two or more words, or between two or more sentences.

102. The method according to claim 99, wherein the first label comprises Open or closed list; Apposition; Rhetorical question; Request; Suggestion; Confirmation; Transition to another subject; Bi-partites; if / then; when / then; cause&effect; contradiction [X but Y]; Proclamation; Narration; Title; Main subject; Background; Elaboration; Conclusion; Recap; About to make a point; Making a point; Point made; Repetition; Apposition; Parentheticals; Digression; or any combination thereof.

103. The method according to any one of the preceding claims, wherein the first microphone is housed in, is attached to, or is integrated with, a first device having a first enclosure, and whereinthe generating of the inference output data is performed in a second device having a second enclosure.

104. The method according to claim 103, wherein the first device comprises a client device.

105. The method according to claim 104, further comprising storing, operating, or using, by the client device, a client operating system.

106. The method according to claim 105, wherein the client operating system consists of, comprises, or is based on, one out of Microsoft Windows 7, Microsoft Windows XP, Microsoft Windows 8, Microsoft Windows 8.1, Linux, and Google Chrome OS.

107. The method according to claim 105, wherein the client operating system is a mobile operating system.

108. The method according to claim 107, wherein the mobile operating system comprises Android version 2.2 (Froyo), Android version 2.3 (Gingerbread), Android version 4.0 (Ice Cream Sandwich), Android Version 4.2 (Jelly Bean), Android version 4.4 (KitKat), Apple iOS version 3, Apple iOS version 4, Apple iOS version 5, Apple iOS version 6, Apple iOS version 7, Microsoft Windows® Phone version 7, Microsoft Windows® Phone version 8, Microsoft Windows® Phone version 9, or Blackberry® operating system.

109. The method according to claim 107, wherein the client operating system is a Real-Time Operating System (RTOS).

110. The method according to claim 109, wherein the RTOS comprises FreeRTOS, SafeRTOS, QNX, VxWorks, or Micro-Controller Operating Systems (pC / OS).

111. The method according to claim 104, wherein the first enclosure comprises a hand-held enclosure or a portable enclosure.

112. The method according to claim 104, wherein the client device consists of, comprises, is part of, or is integrated with, a notebook computer, a laptop computer, a media player, a Digital Still Camera (DSC), a Digital video Camera (DVC or digital camcorder), a Personal Digital Assistant (PDA), a cellular telephone, a digital camera, a video recorder, or a smartphone.

113. The method according to claim 104, wherein the client device consists of, comprises, is part of, or is integrated with, a smartphone that comprises, or is based on, an Apple iPhone 6 or a Samsung Galaxy S6.

114. The method according to claim 103, wherein the second device comprises, consists of, or is integrated with, a server device.

115. The method according to claim 114, wherein the server device is a dedicated device that manages network resoucres; is not a client device and is not a consumer device; is continuously online with greater availability and maximum up time to receive requests almost all of timeefficiently processes multiple requests from multiple client devices at the same time; generates various logs associated with the client devices and traffic from or to the client devices; primarily interfaces and responds to requests from client devices; has greater fault tolerance and higher reliability with lower failure rates; provides scalability for increasing resources to serve increasing client demands; or any combination thereof.

116. The method according to claim 115, wherein the server device comprises the first device.

117. The method according to claim 114, wherein the server device is storing, operating, or using, a server operating system.

118. The method according to claim 117, wherein the server operating system consists or, comprises of, or based on, one out of Microsoft Windows Server®, Linux, or UNIX.

119. The method according to claim 117, wherein the server operating system consists of, or comprises, Microsoft Windows Server® 2003 R2, 2008, 2008 R2, 2012, or 2012 R2 variant, Linux™ or GNU / Linux-based Debian, GNU / Linux, Debian GNU / kFreeBSD, Debian GNU / Hurd, Fedora™, Gentoo™, Linspire™, Mandriva, Red Hat® Linux, SuSE, and Ubuntu®, UNIX® variant Solaris™, AIX®, Mac™ OS X, FreeBSD®, OpenBSD, or NetBSD®.

120. The method according to claim 116, wherein the server device is cloud-based implemented as part of a public cloud-based service.

121. The method according to claim 120, wherein the public cloud-based service comprises, is provided by, or is based on, Amazon Web Services® (AWS®), Microsoft® Azure™, or Google® Compute Engine™ (GCP).

122. The method according to any one of the preceding claims, wherein the generating of the inference output data is provided as an Infrastructure as a Service (laaS) or as a Software as a Service (SaaS).

123. The method according to any one of the preceding claims, for use with a first device in a first enclosure and with a second device in a second enclosure that communicate over a network, wherein the microphone is housed in, attached to, or integrated with, the first device, and wherein the model or the output data is generated in the second device.

124. The method according to claim 123, wherein the first and second devices are configured to communicate over the Internet.

125. The method according to claim 124, wherein the second device comprises a server device, the first device comprises a client device, and wherein the first and second devices communicate using a client- server architecture.

126. The method according to claim 123, wherein the first and second devices communicate over a wireless network, the method further comprising:coupling, by a first antenna in the first device, to the wireless network; coupling, by a second antenna in the second device, to the wireless network; sending, by the first device using a first wireless transceiver that is coupled to the first antenna, to the wireless network via the first antenna, a first data; and receiving, by the second device using a second wireless transceiver that is coupled to the second antenna, from the wireless network via the second antenna, the first data.

127. The method according to claim 126, wherein the first data comprises part of, or whole of, the captured speech, any representation thereof, or any function thereof.

128. The method according to claim 126, further comprising: sending, by the second device using the second wireless transceiver that is coupled to the second antenna, a second data; and receiving, by the first device using the first wireless transceiver that is coupled to the first antenna, the second data.

129. The method according to claim 128, wherein the second data comprises part of, or whole of, the output data, any representation thereof, or any function thereof.

130. The method according to claim 126, further comprising: sending, by the second device using the second wireless transceiver that is coupled to the second antenna, a second data; and receiving, by a third device using a third wireless transceiver that is coupled to a third antenna, the second data.

131. The method according to claim 130, wherein the second data comprises part of, or whole of, the output data, any representation thereof, or any function thereof.

132. The method according to claim 126, wherein the wireless network is over a licensed radio frequency band.

133. The method according to claim 126, wherein the wireless network is over an unlicensed radio frequency band.

134. The method according to claim 133, wherein the unlicensed radio frequency band is an Industrial, Scientific and Medical (ISM) radio band.

135. The method according to claim 134, wherein the ISM band comprises, or consists of, a 2.4 GHz band, a 5.8 GHz band, a 61 GHz band, a 122 GHz band, or a 244 GHz band.

136. The method according to claim 126, wherein the wireless network is a Wireless Personal Area Network (WPAN), the antenna comprises a WPAN antenna, and the wireless transceiver comprises a WPAN transceiver.

137. The method according to claim 136, wherein the WPAN is according to, compatible with, or based on, Bluetooth™ or Institute of Electrical and Electronics Engineers (IEEE) IEEE 802.15.1-2005 standards, or wherein the WPAN is a wireless control network that is according to, or based on, Zigbee™, IEEE 802.15.4-2003, or Z-Wave™ standards.

138. The method according to claim 136, wherein the WPAN is according to, compatible with, or based on, Bluetooth Low-Energy (BLE).

139. The method according to claim 126, wherein the wireless network is a Wireless Local Area Network (WLAN), the antenna comprises a WLAN antenna, and the wireless transceiver comprises a WLAN transceiver.

140. The method according to claim 139, wherein the WLAN is according to, compatible with, or based on, IEEE 802.11-2012, IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.1 In, or IEEE 802.1 lac.

141. The method according to claim 126, wherein the wireless network is a Wireless Wide Area Network (WWAN).

142. The method according to claim 141, wherein the WWAN is according to, compatible with, or based on, WiMAX network that is according to, compatible with, or based on, IEEE 802.16- 2009.

143. The method according to claim 141, wherein the wireless network is a cellular telephone network.

144. The method according to claim 143, wherein the wireless network is a cellular telephone network that is a Third Generation (3G) network that uses Universal Mobile Telecommunications System (UMTS), Wideband Code Division Multiple Access (W-CDMA) UMTS, High Speed Packet Access (HSPA), UMTS Time-Division Duplexing (TDD), CDMA2000 IxRTT, Evolution - Data Optimized (EV-DO), or Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE) EDGE-Evolution, or wherein the cellular telephone network is a Fourth Generation (4G) network that uses Evolved High Speed Packet Access (HSPA+), Mobile Worldwide Interoperability for Microwave Access (WiMAX), Long-Term Evolution (LTE), LTE-Advanced, Mobile Broadband Wireless Access (MBWA), or is based on IEEE 802.20-2008.

145. The method according to claim 123, wherein the first device communicates with the second device over a wired network, the method further comprising: coupling, by a connector in the second device, to the wired network; transmitting, by a wired transceiver in the second device that is coupled to the connector, a first data to the wired network via the connector; andreceiving, by the wired transceiver in the second device that is coupled to the connector, the first data from the wired network via the connector.

146. The method according to claim 145, wherein the first data comprises part of, or whole of, the captured speech, any representation thereof, or any function thereof.

147. The method according to claim 145, wherein the wired network is a Personal Area Network (PAN), the connector is a PAN connector, and the wired transceiver is a PAN transceiver.

148. The method according to claim 145, wherein the wired network is a Local Area Network (LAN), the connector is a LAN connector, and the wired transceiver is a LAN transceiver.

149. The method according to claim 148, wherein the LAN is Ethernet based.

150. The method according to claim 149, wherein the LAN is according to, is compatible with, or is based on, IEEE 802.3-2008 standard.

151. The method according to claim 150, wherein the LAN is of according to, is compatible with, or is based on, a standard selected from the group consisting of lOBase-T, lOOBase-T, lOOBase-TX, 100Base-T2, 100Base-T4, lOOOBase-T, lOOOBase-TX, 10GBase-CX4, and lOGBase-T; and the LAN connector is an RJ-45 connector.

152. The method according to claim 150, wherein the LAN is according to, is compatible with, or is based on, a standard selected from the group consisting of lOBase-FX, lOOBase-SX, lOOBase-BX, lOOBase-LXlO, lOOOBase-CX, lOOOBase-SX, lOOOBase-LX, lOOOBase-LXlO, lOOOBase-ZX, lOOOBase-BXlO, lOGBase-SR, lOGBase-LR, lOGBase-LRM, lOGBase-ER, lOGBase-ZR, and 10GBase-LX4, and the LAN connector is a fiber-optic connector.

153. The method according to claim 145, wherein the wired network is a packet-based or switched-based Wide Area Network (WAN), the connector is a WAN connector, and the wired transceiver is a WAN transceiver.

154. The method according to any one of the preceding claims, further comprising storing a part of, or whole of, the output data, a representation thereof, or a function thereof, in a database.

155. The method according to claim 154, further comprising storing a part of, or whole of, the captured speech, a representation thereof, or a function thereof, in a database.

156. The method according to claim 154, wherein the database is a relational database.

157. The method according to claim 156, wherein the relational database is Structured Query Language (SQL) based.

158. The method according to claim 154, wherein the output data is stored in the database as a record that is searchable using a label as a query.

159. The method according to any one of the preceding claims, further preceded by training the model to identify non-verbal information by prosodic multilayered analysis of multiple intonation units.

160. The method according to claim 159, further comprising repeating the training.

161. The method according to claim 159, wherein the training comprises: capturing, by a second microphone, a speech that vocalize additional text data that comprises multiple words; sounding, by a sounder to a person, the speech captured by the second microphone; obtaining, automatically or manually, the text data or the multiple words; obtaining, by a first input component from the person, in response to the sounding, an identification of multiple Intonation Units (IUS) and multiple labels associated with the identified IUs, wherein each of the IUs comprises one or more words from the additional text data or multiple words; and providing the speech captured by the second microphone, the obtained identified IUs, the obtained associated multiple labels, and the obtained text data or multiple words, to the model for training the model.

162. The method according to claim 161, wherein the providing comprises providing of the obtained multiple labels and the obtained text data as a single structured data stream that uses a structure.

163. The method according to claim 162, wherein the structure of the single structured data stream comprises the obtained multiple labels and the obtained text data at pre-defined positions or locations in the data stream.

164. The method according to claim 163, wherein the structure comprises a sequence of consecutive pairs, wherein each pair comprises a part of the text data followed by an obtained label that is associated with a preceding part of the text data in the pair.

165. The method according to claim 164, wherein the part of the text data in each of the pairs comprises one or more words.

166. The method according to claim 162, wherein the generating comprises generating of the inference output data as a single structured data stream that uses, or is based on, the structure.

167. The method according to claim 166, wherein the structure comprises a sequence of consecutive pairs, wherein each pair comprises one or more of the multiple words followed by an obtained label that is associated with the preceding one or more of the multiple words in the pair.

168. The method according to any one of the preceding claims, wherein the generating comprises generating of the inference output data as a single structured data stream that comprises a sequence of consecutive pairs, wherein each pair comprises one or more words from the multiple words followed by a label that is associated with the preceding one or more words in the pair.

169. The method according to claim 168, wherein the model uses, comprises, or is based on, an asynchronous transformer that is based on, or uses, an encoder-decoder Transformer architecture, wherein the single structured data stream is generated by predicting tokens as outputs of one or more text decoders, and wherein each token is predicted based on former predicted or used tokens.

170. The method according to claim 169, further comprising replacing, in each pair in the sequence, a predicted token associated with a predicting of the one or more words from the multiple words, with pre-defined one or more words.

171. The method according to claim 170, wherein each of the pre-defined one or more words comprises, or is based on, one or more words that are transcribed from the captured speech.

172. The method according to claim 170, wherein each of the pre-defined one or more words comprises, or is based on, one or more words that are transcribed from the captured speech as part of a training of the acoustic model.

173. The method according to claim 161, the training further comprising aligning the speech captured by the second microphone with the obtained identified IUS for synchronization thereof, and wherein the providing to the model comprises providing an alignment information.

174. The method according to claim 161, wherein the obtaining of the text data or multiple words comprises manually obtaining, and wherein the training further comprises obtaining, by a second input component from the person, the additional text data or multiple words.

175. The method according to claim 161, wherein the obtaining of the additional text data or multiple words comprises obtaining, by a second input component from the person, the additional text data.

176. The method according to claim 175, wherein the second input component is same as, or identical to, the first input component.

177. The method according to claim 175, wherein the second input component is different from the first input component.

178. The method according to claim 161, wherein the obtaining of the additional text data comprises manually obtaining, and wherein the training further comprising extracting, using aSpeech-To-Text (STT) scheme, the additional text data from the speech captured by the second microphone.

179. The method according to claim 178, wherein the STT comprises, uses, or is based on, a Hidden Markov Models (HMMs), a Dynamic Time Warping (DTW), a Neural Network, a Deep feedforward Neural Network (DNN), Denoising Autoencoders, or any combination thereof.

180. The method according to claim 178, further comprising displaying, by a display to the person, the extracted additional text data, and wherein the obtaining by the first input component from the person is in response to the displaying.

181. The method according to claim 180, wherein the display consists of, or comprises, a display screen for visually presenting information.

182. The method according to claim 181, wherein the display or the display screen consists of, or comprises, a monochrome, grayscale or color display, having an array of light emitters or light reflectors.

183. The method according to claim 181, wherein the display or the display screen consists of, or comprises, a projector selected from the group consisting of an Eidophor projector, Liquid Crystal on Silicon (LCoS or LCDS) projector, LCD projector, MEMS projector, and Digital Light Processing (DLP™) projector.

184. The method according to claim 181, wherein the display or the display screen selected from the group consisting of a Cathode-Ray Tube (CRT), a Field Emission Display (FED), an Electroluminescent Display (ELD), a Vacuum Fluorescent Display (VFD), or an Organic Light- Emitting Diode (OLED) display, a passive-matrix (PMOLED) display, an active-matrix OLEDs (AMOLED) display, a Liquid Crystal Display (LCD) display, a Thin Film Transistor (TFT) display, an LED-backlit LCD display, or an Electronic Paper Display (EPD) display that is based on Gyricon technology, Electro- Wetting Display (EWD), and Electrofluidic display technology.

185. The method according to claim 181, wherein the display or the display screen consists of, or comprises, a segment display based on a seven-segment display, a fourteen- segment display, a sixteen- segment display, or a dot matrix display, and is operative to only display at least one of digits, alphanumeric characters, words, characters, arrows, symbols, ASCII, and non-ASCII characters.

186. The method according to claim 161, wherein the sounder is configured to convert an electrical energy to omnidirectional, unidirectional, or bidirectional pattern of emitted, audible or inaudible, sound waves.

187. The method according to claim 186, wherein the sound waves are audible, and wherein the sounder comprises an electromagnetic loudspeaker, a piezoelectric speaker, an electrostatic loudspeaker (ESL), a ribbon or planar magnetic loudspeaker, or a bending wave loudspeaker.

188. The method according to claim 186, wherein the sounder comprises circumaural, supraaural, earbuds, in-ear headphones.

189. The method according to claim 161, wherein the first input component comprises a keyboard or a pointing device.

190. The method according to claim 189, wherein the training further comprises applying a transfer learning scheme.

191. The method according to claim 190, further comprising providing, to the model, multiple token sets that are unknown to the model.

192. The method according to claim 191, wherein each of the token sets comprises, or is associated with, a combination of a word from the text data and a label associated with the respective word.

193. The method according to any one of the preceding claims, wherein one or more of the steps, or the model, comprises, is provided as, uses, or is interfaced by, a Software Development Kit (SDK).

194. The method according to any one of the preceding claims, wherein one or more of the steps, or the model, comprises, is provided as, uses, or is interfaced by, an Application Programming Interface (API).

195. The method according to any one of the preceding claims, wherein at least one of the steps comprises executing, by a processor, a software or firmware stored in a computer readable medium.

196. The method according to any one of the preceding claims, wherein the first microphone comprises an omnidirectional, unidirectional, or bidirectional microphone, that is based on the sensing an incident sound-based motion of a diaphragm or a ribbon.

197. The method according to any one of the preceding claims, wherein the first microphone comprises a condenser, an electret, a dynamic, a ribbon, a carbon, or a piezoelectric microphone.

198. The method according to any one of the preceding claims, wherein the capturing uses multiple microphones that include the first microphone.

199. The method according to claim 198, wherein the multiple microphones are arranged as a directional microphones array configured to estimate a number, magnitude, frequency, Direction-Of- Arrival (DOA), distance, or speed of a sound impinging the microphones array.

200. The method according to any one of the preceding claims, further comprising converting, by an Analog-to-Digital (A / D) converter coupled to the first microphone, the captured speech to a digital data stream.

201. The method according to claim 200, wherein the digital data stream comprises a compressed on uncompressed digital data stream.

202. The method according to claim 200, wherein the digital data stream format comprises Pulse-Code Modulation (PCM), Linear Pulse-Code Modulation (LPCM), Waveform Audio File Format (WAV), Windows Media Audio (WMA), or MP3.

203. The method according to any one of the preceding claims, for use with a labels list, the method further comprising checking whether the first label is included in the labels list, and responsive to the checking of whether the first label is included in the labels list, performing a first action or a second action.

204. The method according to claim 203, wherein the first action is performed in response to determining that the first label is included in the labels list.

205. The method according to claim 204, wherein the second action is performed in response to determining that the first label is not included in the labels list.

206. The method according to claim 203, wherein the labels list includes at least 1, 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, or 1,000 labels.

207. The method according to claim 203, wherein the labels list includes less than 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, or 2,000 labels.

208. The method according to claim 203, for use with a categories list, the method further comprising checking whether the first category is included in the categories list, and responsive to the checking whether the first category is included in the categories list, performing a first action or a second action.

209. The method according to claim 208, wherein the first action is performed in response to further determining that the first category is included in the categories list.

210. The method according to claim 208, wherein the second action is performed in response to further determining that the first category is not included in the categories list.

211. The method according to claim 208, wherein the categories list includes 1, 2, 3, 4, or 5 categories.

212. The method according to claim 203, for use with a words list, the method further comprising checking whether a first word of the multiple words that is associated with the firstlabel is included in the words list, and responsive to the checking whether the first word is included in the words list, performing a first action or a second action.

213. The method according to claim 212, wherein the first action is performed in response to further determining that the first word is included in the words list.

214. The method according to claim 212, wherein the second action is performed in response to further determining that the first word is not included in the words list.

215. The method according to claim 212, wherein the words list includes at least 1, 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, 2,000, 5,000 or 10,000 words.

216. The method according to claim 212, wherein the words list includes less than 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, 2,000, 5,000, 10,000 or 20,000 words.

217. The method according to claim 203, wherein the first action is identical to, similar to, or different from, the second action.

218. The method according to claim 203, wherein the first or second action comprises initiating a communication over a network with a remote device.

219. The method according to claim 218, further comprising sending an alert message over the network to the remote device.

220. The method according to claim 203, wherein the performing of the first or second action comprises activating, deactivating, or controlling, an actuator that converts electrical energy to effect or produce a physical phenomenon.

221. The method according to claim 220, wherein the actuator is electrically powered from a power source, and converts electrical power from the power source to effect or produce the physical phenomenon.

222. The method according to claim 221, wherein the power source is an Alternating Current (AC) or a Direct Current (DC) power source.

223. The method according to claim 222, wherein the power source is a Direct Current (DC) power source that comprises a primary or a rechargeable battery.

224. The method according to claim 221, wherein the power source is an Alternating Current (AC) that comprises a domestic AC power that is nominally 120V AC / 60Hz or 230V AC / 50Hz.

225. The method according to claim 220, wherein the actuator is electrically powered from a power source via a switch connected between the power source and the actuator, wherein activating of the actuator comprises switching power from the power source to the actuator bythe switch, and wherein deactivating of the actuator comprises disconnecting switching power from the power source to the actuator by the switch.

226. The method according to claim 225, wherein the switch is an electrically-controlled AC power Single-Pole-Double-Throw (SPDT) switch, wherein the switch comprises, is based on, is part of, or consists of, a relay, or wherein the switch is based on, comprises, or consists of, an electrical circuit.

227. The method according to claim 220, wherein the actuator is configured to influence, affect, create, or change the phenomenon in an object that is a gas, an air, a liquid, or a solid.

228. The method according to claim 220, wherein the actuator is operative to provide timedependent characteristic of the phenomenon selected from the group consisting of a time- integrated, an average, an RMS (Root Mean Square) value, a frequency, a period, a duty-cycle, a time-integrated, and a time-derivative.

229. The method according to claim 220, wherein the actuator is operative to provide spacedependent characteristic that is selected from the group consisting of a pattern, a linear density, a surface density, a volume density, a flux density, a current, a direction, a rate of change in a direction, and a flow.

230. The method according to claim 220, wherein the actuator consists of, or comprises, an electric light source that converts electrical energy into light, and emits visible or non-visible light for illumination or indication, and the non-visible light is infrared, ultraviolet, X-rays, or gamma rays, and wherein the electric light source consists of, or comprises, a lamp, an incandescent lamp, a gas discharge lamp, a fluorescent lamp, a Solid-State Lighting (SSL), a Light Emitting Diode (LED), an Organic LED (OLED), a polymer LED (PLED), or a laser diode.

231. The method according to claim 220, wherein the actuator consists of, or comprises, a motion actuator that causes linear or rotary motion, or wherein the motion actuator consists of, or comprises, a pneumatic actuator, hydraulic actuator, or electrical actuator.

232. The method according to claim 231, wherein the motion actuator consists of, or comprises, an electrical motor, that is a brushed motor, a brushless motor, an uncommutated DC motor, a DC stepper motor that is a Permanent Magnet (PM) motor, a Variable reluctance (VR) motor, a hybrid synchronous stepper motor, or wherein the motion actuator consists of, or comprises, an AC motor that is an induction motor, a synchronous motor, an eddy current motor, a singlephase AC induction motor, a two-phase AC servo motor, or a three-phase AC synchronous motor, and the AC motor is a split-phase motor, a capacitor-start motor, a Permanent-Split Capacitor (PSC) motor, an electrostatic motor, a piezoelectric actuator, or is a MEMS-basedmotor, a linear hydraulic actuator, a linear pneumatic actuator, a linear induction electric motor (LIM), a Linear Synchronous electric Motor (LSM), a piezoelectric motor, a Surface Acoustic Wave (SAW) motor, a Squiggle motor, an ultrasonic motor, or a micro- or nanometer combdrive capacitive actuator, a Dielectric or Ionic based Electroactive Polymers (EAPs) actuator, a solenoid, a thermal bimorph, or a piezoelectric unimorph actuator.

233. The method according to claim 220, wherein the actuator consists of, or comprises, a compressor or a pump and is operative to move, force, or compress a liquid, a gas or a slurry, wherein the pump is a direct lift pump, an impulse pump, a displacement pump, a valveless pump, a velocity pump, a centrifugal pump, a vacuum pump, or a gravity pump, wherein the actuator consists of, or comprises, a positive displacement pump that is a rotary lobe pump, a progressive cavity pump, a rotary gear pump, a piston pump, a diaphragm pump, a screw pump, a gear pump, a hydraulic pump, or a vane pump, or a rotary-type positive displacement pump that is an internal gear pump, a screw pump, a shuttle block pump, a flexible vane pump, a sliding vane pump, a rotary vane pump, a circumferential piston pump, a helical twisted roots pump, or a liquid ring vacuum pump, or wherein the actuator consists of, or comprises, a reciprocating-type positive displacement type that is a piston pump, a diaphragm pump, a plunge pump, a diaphragm valve pump, or a radial piston pump.

234. The method according to claim 220, wherein the actuator consists of, or comprises, a display screen for visually presenting information, wherein the display or the display screen consists of, or comprises, a monochrome, grayscale or color display, having an array of light emitters or light reflectors, wherein the display or the display screen consists of, or comprises, a projector selected from the group consisting of an Eidophor projector, Liquid Crystal on Silicon (LCoS or LCOS) projector, LCD projector, MEMS projector, and Digital Light Processing (DLP™) projector, wherein the display or the display screen consists of, or comprises, a video display supporting Standard-Definition (SD) or High-Definition (HD) standards, and is capable of scrolling, static, bold or flashing the presented information, wherein the video display is a 3D video display.

235. The method according to claim 220, wherein the actuator consists of, or comprises, a thermoelectric actuator and is a heater or a cooler, operative for affecting a temperature of a solid, a liquid, or a gas object, and is coupled to the object by conduction, convection, force convention, thermal radiation, or by a transfer of energy by phase changes, and wherein the thermoelectric actuator consists of, or comprises, a cooler based on a heat pump driving a refrigeration cycle using a compressor-based electric motor, an electric heater that is a resistanceheater or a dielectric heater, or an induction heater that is solid-state based or is an active heat pump that uses, or is based on, the Peltier effect.

236. The method according to claim 220, wherein the actuator consists of, or comprises, a chemical or an electrochemical actuator, and is operative for producing, changing, or affecting a matter structure, properties, composition, process, or reactions, or wherein the electrochemical actuator is operative for producing, changing, or affecting, an oxidation / reduction or an electrolysis reaction.

237. The method according to claim 220, wherein the actuator consists of, or comprises, an electromagnetic coil or an electromagnet operative for generating a magnetic or electric field, wherein the actuator consists of, or comprises, an electrical signal generator, wherein the signal generator is operative to output repeating or non-repeating electronic signals, wherein the signal generator is an analog signal generator having an analog voltage or analog current output, and wherein the output of the analog signal generator is a sine wave, a sawtooth wave, a step (pulse), a square wave, or a triangular wave, an Amplitude Modulation (AM), a Frequency Modulation (FM), or a Phase Modulation (PM) signal, or wherein the signal generator is an Arbitrary Waveform Generator (AWG) or a logic signal generator.

238. The method according to claim 220, wherein the actuator consists of, or comprises, a sounder for converting an electrical energy to omnidirectional, unidirectional, or bidirectional pattern of emitted, audible or inaudible, sound waves, wherein the sound waves are audible, and wherein the sounder is an electromagnetic loudspeaker, a piezoelectric speaker, an electrostatic loudspeaker (ESL), a ribbon or planar magnetic loudspeaker, or a bending wave loudspeaker, wherein the sounder is operative to emit a single or multiple tones, wherein the sounder is operative to perform continuous or intermittent operation, wherein the sounder is an electromechanical or a ceramic -based, and is an electric bell, a buzzer (or beeper), a chime, a whistle or a ringer, or wherein the sounder is a loudspeaker, and wherein the actuator is operative to store and play one or more digital audio content files.

239. The method according to any one of the preceding claims, further comprising executing a chatbot, providing the inferred output data to the chatbot to be processed therein, and to provide a chatbot response that is based on at least part of the output data.

240. A non-transitory computer readable medium having computer executable instructions stored thereon, wherein the instructions include the steps of claim 1.

241. A method for training a model to identify non-verbal information by prosodic multilayered analysis of multiple intonation units, the method comprises: capturing, by a first microphone, a speech that vocalize text data that comprises multiple words; sounding, by a sounder to a person, the speech captured by the first microphone; obtaining, automatically or manually, the text data or the multiple words; obtaining, by a first input component from the person, in response to the sounding, an identification of multiple Intonation Units (IUS) and multiple labels associated with the identified IUs, wherein each of the IUs comprises one or more words from the text data or multiple words; and providing the speech captured by the first microphone, the obtained identified IUs, the obtained associated multiple labels, and the obtained text data or multiple words, to a trained weakly- supervised deep learning acoustic model for training the model.

242. The method according to claim 241, further comprising repeating the training.

243. The method according to claims 241 or 242, wherein the providing comprises providing of the obtained multiple labels and the obtained text data as a single structured data stream that uses a structure.

244. The method according to claim 243, wherein the structure of the single structured data stream comprises the obtained multiple labels and the obtained text data at pre-defined positions or locations in the data stream.

245. The method according to claim 244, wherein the structure comprises a sequence of consecutive pairs, wherein each pair comprises a part of the text data followed by an obtained label that is associated with a preceding part of the text data in the pair.

246. The method according to claim 245, wherein the part of the text data in each of the pairs comprises one or more words.

247. The method according to claim 243, further followed by generating, by the trained acoustic model, an inference output data.

248. The method according to claim 247, wherein the generating comprises generating of the inference output data as a single structured data stream that uses, or is based on, the structure.

249. The method according to claim 248, wherein the structure comprises a sequence of consecutive pairs, wherein each pair comprises one or more of the multiple words followed by an obtained label that is associated with preceding one or more of the multiple words in the pair.

250. The method according to any one of claims 241-249, wherein the text data comprises one or more phrases, one or more sentences, or one or more clauses.

251. The method according to any one of claims 241-250, wherein all steps are performed in a single enclosure.

252. The method according to any one of claims 241-251, further comprising aligning the speech captured by the first microphone with the obtained identified IUS for synchronization thereof, and wherein the providing to the model comprises providing an alignment information.

253. The method according to any one of claims 241-252, wherein the obtaining of the text data or multiple words comprises manually obtaining, and wherein the training further comprises obtaining, by a second input component from the person, the text data or multiple words.

254. The method according to any one of claims 241-253, wherein the obtaining of the text data or multiple words comprises obtaining, by a second input component from the person, the text data.

255. The method according to claim 254, wherein the second input component is same as, or identical to, the first input component.

256. The method according to claim 254, wherein the second input component is different from the first input component.

257. The method according to any one of claims 241-256, wherein the obtaining of the text data comprises manually obtaining, and wherein the training further comprising extracting, using a Speech-To-Text (STT) scheme, the text data from the speech captured by the first microphone.

258. The method according to claim 257, wherein the STT comprises, uses, or is based on, a Hidden Markov Models (HMMs), a Dynamic Time Warping (DTW), a Neural Network, a Deep feedforward Neural Network (DNN), Denoising Autoencoders, or any combination thereof.

259. The method according to claim 257, further comprising displaying, by a display to the person, the extracted text data, and wherein the obtaining by the first input component from the person is in response to the displaying.

260. The method according to claim 259, wherein the display consists of, or comprises, a display screen for visually presenting information.

261. The method according to claim 260, wherein the display or the display screen consists of, or comprises, a monochrome, grayscale or color display, having an array of light emitters or light reflectors.

262. The method according to claim 260, wherein the display or the display screen consists of, or comprises, a projector selected from the group consisting of an Eidophor projector, Liquid Crystal on Silicon (LCoS or LCDS) projector, LCD projector, MEMS projector, and Digital Light Processing (DLP™) projector.

263. The method according to claim 260, wherein the display or the display screen selected from the group consisting of a Cathode-Ray Tube (CRT), a Field Emission Display (FED), an Electroluminescent Display (ELD), a Vacuum Fluorescent Display (VFD), or an Organic Light- Emitting Diode (OLED) display, a passive-matrix (PMOLED) display, an active-matrix OLEDs (AMOLED) display, a Liquid Crystal Display (LCD) display, a Thin Film Transistor (TFT) display, an LED-backlit LCD display, or an Electronic Paper Display (EPD) display that is based on Gyricon technology, Electro- Wetting Display (EWD), and Electrofluidic display technology.

264. The method according to claim 260, wherein the display or the display screen consists of, or comprises, a segment display based on a seven-segment display, a fourteen- segment display, a sixteen- segment display, or a dot matrix display, and is operative to only display at least one of digits, alphanumeric characters, words, characters, arrows, symbols, ASCII, and non-ASCII characters.

265. The method according to any one of claims 241-264, wherein the sounder is configured to convert an electrical energy to omnidirectional, unidirectional, or bidirectional pattern of emitted, audible or inaudible, sound waves.

266. The method according to claim 265, wherein the sound waves are audible, and wherein the sounder comprises an electromagnetic loudspeaker, a piezoelectric speaker, an electrostatic loudspeaker (ESL), a ribbon or planar magnetic loudspeaker, or a bending wave loudspeaker.

267. The method according to claim 265, wherein the sounder comprises circumaural, supraaural, earbuds, in-ear headphones.

268. The method according to any one of claims 241-267, wherein the first input component comprises a keyboard or a pointing device.

269. The method according to claim 268, wherein the training further comprises applying a transfer learning scheme.

270. The method according to claim 269, further comprising providing, to the model, multiple token sets that are unknown to the model.

271. The method according to claim 270, wherein each of the token sets comprises, or is associated with, a combination of a word from the text data and a label associated with the respective word.

272. The method according to any one of claims 241-271, wherein one or more of the steps, or the model, comprises, is provided as, uses, or is interfaced by, a Software Development Kit (SDK).

273. The method according to any one of claims 241-272, wherein one or more of the steps, or the model, comprises, is provided as, uses, or is interfaced by, an Application Programming Interface (API).

274. The method according to any one of claims 241-273, wherein at least one of the steps comprises executing, by a processor, a software or firmware stored in a computer readable medium.

275. The method according to any one of claims 241-274, wherein the first microphone comprises an omnidirectional, unidirectional, or bidirectional microphone, that is based on a sensing an incident sound-based motion of a diaphragm or a ribbon.

276. The method according to any one of claims 241-275, wherein the first microphone comprises a condenser, an electret, a dynamic, a ribbon, a carbon, or a piezoelectric microphone.

277. The method according to any one of the preceding claims, wherein the capturing uses multiple microphones that include the first microphone.

278. The method according to claim 277, wherein the multiple microphones are arranged as a directional microphones array configured to estimate a number, magnitude, frequency, Direction-Of- Arrival (DOA), distance, or speed of a sound impinging the microphones array.

279. The method according to any one of the preceding claims, further comprising converting, by an Analog-to-Digital (A / D) converter coupled to the first microphone, the captured speech to a digital data stream.

280. The method according to claim 279, wherein the digital data stream comprises a compressed on uncompressed digital data stream, and wherein the digital data stream format comprises Pulse-Code Modulation (PCM), Linear Pulse-Code Modulation (LPCM), Waveform Audio File Format (WAV), Windows Media Audio (WMA), or MP3.

281. The method according to any one of claims 241-280, further comprising training the model for a duration of at least 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100 or 200 minutes, or at least 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100 or 200 hours.

282. The method according to any one of claims 241-281, further comprising training the model for a duration of less than 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100, or 200 minutes, or at least 1, 2, 3, 5, 7, 10, 12, 15, 20, 25, 30, 50, 70, 100 or 200 hours.

283. The method according to any one of claims 241-282, further comprising training the model by at least 100, 200, 300, 500, 700, 1,000, 1,200, 1,500, 2,000, 2,500, 3,000, 5,000, 7,000, 10,000, 20,000, 30,000, 40,000, 50,000, 70,000, or 100,000 words, labels, or IUS.

284. The method according to any one of claims 241-283, further comprising training the model by less than 100, 200, 300, 500, 700, 1,000, 1,200, 1,500, 2,000, 2,500, 3,000, 5,000, 7,000, 10,000, 20,000, 30,000, 40,000, 50,000, 70,000, or 100,000 words, labels, or IUS.

285. The method according to any one of claims 241-284, wherein the model comprises, uses, or is based on, a machine learning model that is primarily configured for, or trained for, speech recognition, transcription, or translation, and that uses, comprises, or is based on, an encoderdecoder Transformer architecture.

286. The method according to any one of claims 241-285, wherein the model comprises, uses, or is based on, an open-source software.

287. The method according to any one of claims 241-286, wherein the model comprises, uses, or is based on, Mel-frequency Cepstrum or a log-Mel spectrogram.

288. The method according to any one of claims 241-287, wherein the model comprises, uses, or is based on, a sinusoidal positional encoding or sinusoidal positional intermixing, with a learned positional encoding.

289. The method according to any one of claims 241-288, wherein the model comprises, uses, or is based on, Whisper model or architecture, by OpenAI.

290. The method according to any one of claims 241-289, further for using the trained model for inferring non-verbal information by prosodic multilayered analysis of intonation units, the method comprising: capturing, by a second microphone, an additional speech that vocalize an additional text data that comprises additional multiple words; providing the captured additional speech to the trained weakly- supervised deep learning acoustic model; and generating, by the trained acoustic model, an inference output data, wherein the output data comprise the additional multiple words, identification of first and second Intonation Units (IUs), a first label in a first category, and a second label in a second category, wherein each of the identified first and second IUs comprises one or more words of the multiple words, wherein the first IU is associated with the first label and the second IU is associated with the second label, and wherein the first and second categories are selected from a group that consists of a ‘Genre’ category, a ‘Prototype’ category, a ‘Discourse Function’ category, an ‘Emotion’ category, an ‘Emphasis’ category, and an ‘Attitude’ category.

291. The method according to claim 290, wherein the additional speech is in the English language.

292. The method according to claim 290, wherein the additional speech is in a non-English language.

293. The method according to claim 290, wherein a duration of the capturing of the additional speech is at least 0.1, 0.2, 0.3, 0.5, 0.7, 1.0, 1.1, 1.2, 1.5, 1.7, 2, 3, 5, 10, 12, 15, 20, 30, 50, or 100 second, or at least 1, 2, 3, 5, 10, 12, 15, 20, 30, 50, 100, 120, 150, 200, 300, 500, or 1,000 minutes.

294. The method according to claim 290, wherein a duration of the capturing of the additional speech is no more than 0.1, 0.2, 0.3, 0.5, 0.7, 1.0, 1.1, 1.2, 1.5, 1.7, 2, 3, 5, 10, 12, 15, 20, 30, 50, or 100 second, or at least 1, 2, 3, 5, 10, 12, 15, 20, 30, 50, 100, 120, 150, 200, 300, 500, or 1,000 minutes.

295. The method according to claim 290, wherein the first IUS comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words.

296. The method according to claim 295, wherein the first label is associated with a single word of the first IU.

297. The method according to claim 290, wherein each one of the first and second IUs comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words.

298. The method according to claim 290, wherein each one of the first and second IUs comprises less than 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, or 15 words.

299. The method according to claim 290, wherein a duration of at least one of the first and second IUs is at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long.

300. The method according to claim 290, wherein a duration of each one of the first and second IUs is at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long.

301. The method according to claim 290, wherein a duration of at least one of the first and second IUs is less than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long.

302. The method according to claim 290, wherein a duration of each one of the first and second IUs is less than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.5, 1.7, 2.0, 2.2, or 2.5 seconds long.

303. The method according to claim 290, wherein one of, or each of, the identified first and second IUs comprises a segment of the captured additional speech bounded by two boundariesthat are characterized by a threshold speech-rate deviation, pitch intensity pattern, pitch decay pattern, speech rate pattern, or speech rate decay pattern, or any combination thereof.

304. The method according to claim 290, wherein one of, or each one of, the identified first and second IUS comprises a smallest speech unit that conveys non-verbal information or a label.

305. The method according to claim 290, wherein one of the first and second IUs is associated with multiple distinct labels.

306. The method according to claim 305, wherein one of the first and second IUs is associated with at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 multiple distinct labels.

307. The method according to claim 290, further comprising adding multiple punctuation marks or symbols to the multiple words in the output data.

308. The method according to claim 307, wherein the adding is based on, or uses, boundaries or locations of the identified first and second IUs, the associated first and second labels, or any combination thereof.

309. The method according to claim 307, wherein the multiple punctuation marks or symbols comprise a comma, a period (dot), a quotation mark, parenthesis marks, a dash mark, an exclamation mark, a colon mark, a semi-colon mark, or any combination thereof.

310. The method according to claim 290, wherein the additional text data is in a first language, the method further comprising translating the multiple words into a second language using, or based on, boundaries or locations of the identified first and second IUs, the associated first and second labels, or any combination thereof.

311. The method according to claim 290, further comprising embedding marks, symbols, pictograms, logograms, ideograms, or emojis that are based on, or responsive to, the first or second labels, in the multiple words in the output data.

312. The method according to claim 290, further comprising sounding, by a sounder, using a speech synthesis or a Text-To-Speech (TTS) scheme, the multiple words in the output data being annotated, combined, responsive to, or modified, with the first and second labels.

313. The method according to claim 312, wherein the speech synthesis or the Text-To- Speech (TTS) scheme uses, or is based on, a concatenative synthesis or formant synthesis.

314. The method according to claim 290, further comprising extracting, using a Speech-To-Text (STT) scheme, the additional text data from the speech captured by the second microphone.

315. The method according to claim 314, further comprising aligning the captured additional speech and the multiple words.

316. The method according to claim 314, further comprising providing the multiple words to the trained weakly- supervised deep learning acoustic model, and wherein the output data is generated in response to the provided extracted additional text data.

317. The method according to claim 314, wherein the STT comprises, uses, or is based on, a Hidden Markov Models (HMMs), a Dynamic Time Warping (DTW), a Neural Network, a Deep feedforward Neural Network (DNN), Denoising Autoencoders, or any combination thereof.

318. The method according to claim 290, further comprising manipulating the captured additional speech, wherein the providing of the captured additional speech comprises providing the manipulated captured additional speech.

319. The method according to claim 318, wherein the manipulating comprises filtering of a noise, a music, or a background sound, and enhancing speech-related features.

320. The method according to claim 319, wherein the manipulating comprises amplifying, filtering, converting, range matching, resolution increasing, integrating, deviating, equalizing, compressing, de-compressing, coding, decoding, modulating, demodulating, pattern recognizing, smoothing, or any combination thereof.

321. The method according to claim 318, wherein the manipulating comprises performing a function that is based on, uses, or comprises, a discrete, continuous, monotonic, non-monotonic, elementary, algebraic, linear, polynomial, quadratic, Cubic, Nth-root based, exponential, transcendental, quintic, quartic, logarithmic, hyperbolic, or trigonometric function.

322. The method according to claim 318, wherein the manipulating comprises applying a feature engineering technique.

323. The method according to claim 322, wherein the feature engineering technique comprises, uses, or is based on, Imputation; Categorical Imputation; Numerical Imputation; Discretization; Categorical encoding; Splitting; Outliers removal, Outliers values replacing; capping the maximum and minimum Outliers values; Variable transformation; Scaling; Min-Max Scaling; Standardization / Variance Scaling; Feature Creation; creating interaction features; dimensionality reduction; logarithmic or power variables transformations; feature selection; or any combination thereof.

324. The method according to claim 322, wherein the feature engineering technique comprises, uses, or is based on, creation, transformation, extraction, selection, or any combination thereof, of one or more features or variables.

325. The method according to claim 290, further comprising storing the captured additional speech in a memory.

326. The method according to claim 325, further comprising retrieving the captured additional speech from the memory, wherein the providing of the captured speech comprises providing the retrieved captured additional speech.

327. The method according to claim 290, further comprising storing the output data in a memory.

328. The method according to claim 327, further comprising storing, in the memory, the captured additional speech being associated with the output data.

329. The method according to claim 290, wherein the output data or a part thereof comprises, uses, or is based on, a format that uses, or is based on, a plain text, an annotated text, a Voice Extensible Markup Language (VoiceXML), a VoiceXML (VXML), an Extensible Markup Language (XML), JavaScript Object Notation (JSON) format, a Pronunciation Lexicon Specification (PLS), a Comma-Separated Values (CSV), a Structured Query Language (SQL), or a Java Speech API Markup Language (JSML).

330. The method according to claim 290, further comprising converting the output data or a part thereof to a format that that uses, or is based on, a plain text, an annotated text, a Voice Extensible Markup Language (VoiceXML), a VoiceXML (VXML), an Extensible Markup Language (XML), JavaScript Object Notation (JSON) format, a Pronunciation Lexicon Specification (PLS), a Comma-Separated Values (CSV), a Structured Query Language (SQL), or a Java Speech API Markup Language (JSML).

331. The method according to claim 290, wherein the output data or a part thereof comprises, uses, or is based on, an annotated text that uses, or is based on, a format that uses, or is based on, Tones and Break Indices (ToBI), INCEpTION, or Emotion Markup Language (EML or EmotionML) format.

332. The method according to claim 290, further comprising converting the output data or a part thereof to an annotated text in a format that uses, or is based on, Tones and Break Indices (ToBI) format, INCEpTION, or Emotion Markup Language (EML or EmotionML) format.

333. The method according to claim 290, wherein the output data comprise multiple labels that include at least the first and second labels, and wherein each of the multiple labels is in a ‘Genre’ category, a ‘Prototype’ category, a ‘Discourse Function’ category, an ‘Emotion’ category, an ‘Emphasis’ category, or an ‘Attitude’ category.

334. The method according to claim 290, wherein the first label is in the Genre category.

335. The method according to claim 334, wherein the first label is associated with part of, or whole of, the captured additional speech.

336. The method according to claim 334, wherein the first label is associated with multiple words, a sentence, multiple sentences, a clause, or a paragraph, in the additional text data of the output data.

337. The method according to claim 336, wherein the first label is associated with a sentence in the multiple words, and wherein the first label comprises exclamation, request, command, or suggestion.

338. The method according to claim 336, wherein the first label is associated with a rhetorical mode.

339. The method according to claim 338, wherein the first label comprises narration, description, exposition, argumentation, or any combination thereof.

340. The method according to claim 334, wherein the first label comprises an informal addressing, formal addressing, a narrative, a quotation, a casual conversation, a professional exchange, a debate, a testimonial, an instructional, a tutorial, an interview, an announcement, a reportage, or any combination thereof.

341. The method according to claim 334, wherein the first label comprises a gender that is ‘male’ or ‘female’, wherein the first label comprises a dialect type, an accent type, or an estimated age.

342. The method according to claim 334, wherein the first label comprises a naturality indicator that is ‘human’ or ‘machine’, respectively indicating human or machine generated speech.

343. The method according to claim 290, wherein the first label is in the prototype category that identifies a boundary in the captured additional speech or in the additional text data of the output data, and wherein the first label is associated with the first IU.

344. The method according to claim 343, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the prototype category.

345. The method according to claim 343, wherein the first label is associated with a tone change, or wherein the first label is associated with a punctuation.

346. The method according to claim 345, wherein the first label is associated with a flattish tone, a falling tone, a rising tone, or a disfluency, in the captured additional speech.

347. The method according to claim 345, wherein the first label is associated with a comma, a period, a question mark, or a truncated text, in the additional text data of the output data.

348. The method according to claim 343, wherein the first label comprises a ‘continuation’, a ‘conclusion’, a ‘question’, or a ‘truncated’ in the additional text data of the output data.

349. The method according to claim 290, wherein the first label is in the emphasis category that identifies, emphasizes, intensifies, or otherwise pointed out, in the first IU, or a word in the first IU.

350. The method according to claim 349, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the emphasis category.

351. The method according to claim 349, wherein the first label comprises a ‘contrastive’, a ‘strong’, a ‘weak’, or a ‘de-emphasis’.

352. The method according to claim 290, wherein the first label is in the emotion category exhibited in the first IU.

353. The method according to claim 352, wherein the first label involves intentional and non- intentional emotion or feeling exhibited in the first IU.

354. The method according to claim 352, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the emotion category.

355. The method according to claim 352, wherein the first label comprises an emotion that is part of a “Negative and forceful” class.

356. The method according to claim 355, wherein the first label comprises Anger; Annoyance; Contempt; Disgust; or Irritation.

357. The method according to claim 352, wherein the first label comprises an emotion that is part of a “Negative and not in control” class.

358. The method according to claim 357, wherein the first label comprises Anxiety; Embarrassment; Fear; Helplessness; Powerlessness; or Worry.

359. The method according to claim 352, wherein the first label comprises an emotion that is part of a “Negative thoughts” class.

360. The method according to claim 359, wherein the first label comprises Doubt; Envy; Frustration; Guilt; or Shame.

361. The method according to claim 352, wherein the first label comprises an emotion that is part of a “Negative and passive” class.

362. The method according to claim 361, wherein the first label comprises Boredom; Despair; Disappointment; Hurt; or Sadness.

363. The method according to claim 352, wherein the first label comprises an emotion that is part of an “Agitation” class.

364. The method according to claim 363, wherein the first label comprises Stress; Shock; or Tension.

365. The method according to claim 352, wherein the first label comprises an emotion that is part of a “Positive and lively” class.

366. The method according to claim 365, wherein the first label comprises Amusement; Delight; Elation; Excitement; Happiness; Joy; or Pleasure.

367. The method according to claim 352, wherein the first label comprises an emotion that is part of a “Caring” class.

368. The method according to claim 367, wherein the first label comprises Affection; Empathy; Friendliness; or Love.

369. The method according to claim 352, wherein the first label comprises an emotion that is part of a “Positive thoughts” class.

370. The method according to claim 369, wherein the first label comprises Pride; Courage; Hope; Humility; Satisfaction; or Trust.

371. The method according to claim 352, wherein the first label comprises an emotion that is part of a “Quiet positive” class.

372. The method according to claim 371, wherein the first label comprises Calmness; Contentment; Relaxation; Relief; or Serenity.

373. The method according to claim 352, wherein the first label comprises an emotion that is part of a “Reactive” class.

374. The method according to claim 373, wherein the first label comprises Interest; Politeness; or Surprise.

375. The method according to claim 374, wherein the first label comprises Disgusted; Apprehended; Hesitant; Angry; Delighted; Happy; Content; Upset; Nervous; Insecure; Confused; Enthusiastic Frustrated; Relieved; Sad; Hopeful; Jealous; Content; Anxious; Curious; Desperate; Optimistic; Pessimistic; or any combination thereof.

376. The method according to claim 290, wherein the first label is in the attitude category that involves an attitude or sentiment about someone or something exhibited in the first IU.3T1. The method according to claim 376, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the attitude category.

378. The method according to claim 376, wherein the first label comprises Neutral; Negative; Positive; Hesitation (intentional or non-intentional); Inclusion; Exclusion; Solidarity or alignment; Distance; Irony; Sarcasm; Respect; Disrespect; Position of power or authority; Lack of power or authority; Modal distance; Reservation; Reserve; Reluctance; Self-deprecation; Self-humor; Shame; Indifference; Positive surprise; Negative surprise; Indignation; Protest; Admiration; Criticism; Apologetic; Agreement; Disagreement; Encouragement;Discouragement; Reassurance; Uncertainty; Certainty; Sympathy; Empathy; Anticipation; or any combination thereof.

379. The method according to claim 376, wherein the first label corresponds to a cognitive attitude, an affective attitude, or a conative attitude, or wherein the first label corresponds to a negative or positive attitude.

380. The method according to claim 290, wherein the first label is in the discourse function category that involves a relation between two or more IUS or elements in the captured additional speech.

381. The method according to claim 380, wherein the output data comprises at least 2, 3, 4, 5, 7, 10, 12, 15, 20, 30, 50, or 100 labels in the discourse function category.

382. The method according to claim 380, wherein the relation is between two or more words, or between two or more sentences.

383. The method according to claim 380, wherein the first label comprises Open or closed list; Apposition; Rhetorical question; Request; Suggestion; Confirmation; Transition to another subject; Bi-partites; if / then; when / then; cause&effect; contradiction [X but Y]; Proclamation; Narration; Title; Main subject; Background; Elaboration; Conclusion; Recap; About to make a point; Making a point; Point made; Repetition; Apposition; Parentheticals; Digression; or any combination thereof.

384. The method according to claim 290, wherein the second microphone is housed in, is attached to, or is integrated with, a first device having a first enclosure, and wherein the generating of the inference output data is performed in a second device having a second enclosure.

385. The method according to claim 384, wherein the first device comprises a client device.

386. The method according to claim 385, further comprising storing, operating, or using, by the client device, a client operating system.

387. The method according to claim 386, wherein the client operating system consists of, comprises, or is based on, one out of Microsoft Windows 7, Microsoft Windows XP, Microsoft Windows 8, Microsoft Windows 8.1, Linux, and Google Chrome OS.

388. The method according to claim 386, wherein the client operating system is a mobile operating system.

389. The method according to claim 388, wherein the mobile operating system comprises Android version 2.2 (Froyo), Android version 2.3 (Gingerbread), Android version 4.0 (Ice Cream Sandwich), Android Version 4.2 (Jelly Bean), Android version 4.4 (KitKat), Apple iOS version 3, Apple iOS version 4, Apple iOS version 5, Apple iOS version 6, Apple iOS version 7,Microsoft Windows® Phone version 7, Microsoft Windows® Phone version 8, Microsoft Windows® Phone version 9, or Blackberry® operating system.

390. The method according to claim 388, wherein the client operating system is a Real-Time Operating System (RTOS).

391. The method according to claim 390, wherein the RTOS comprises FreeRTOS, SafeRTOS, QNX, VxWorks, or Micro-Controller Operating Systems (pC / OS).

392. The method according to claim 384, wherein the first enclosure comprises a hand-held enclosure or a portable enclosure.

393. The method according to claim 384, wherein the first device consists of, comprises, is part of, or is integrated with, a notebook computer, a laptop computer, a media player, a Digital Still Camera (DSC), a Digital video Camera (DVC or digital camcorder), a Personal Digital Assistant (PDA), a cellular telephone, a digital camera, a video recorder, or a smartphone.

394. The method according to claim 384, wherein the client device consists of, comprises, is part of, or is integrated with, a smartphone that comprises, or is based on, an Apple iPhone 6 or a Samsung Galaxy S6.

395. The method according to claim 384, wherein the second device comprises, consists of, or is integrated with, a server device.

396. The method according to claim 395, wherein the server device is a dedicated device that manages network resoucres; is not a client device and is not a consumer device; is continuously online with greater availability and maximum up time to receive requests almost all of time efficiently processes multiple requests from multiple client devices at the same time; generates various logs associated with the client devices and traffic from or to the client devices; primarily interfaces and responds to requests from client devices; has greater fault tolerance and higher reliability with lower failure rates; provides scalability for increasing resources to serve increasing client demands; or any combination thereof.

397. The method according to claim 396, wherein the server device comprises the first device.

398. The method according to claim 395, wherein the server device is storing, operating, or using, a server operating system.

399. The method according to claim 398, wherein the server operating system consists or, comprises of, or based on, one out of Microsoft Windows Server®, Linux, or UNIX.

400. The method according to claim 398, wherein the server operating system consists of, or comprises, Microsoft Windows Server® 2003 R2, 2008, 2008 R2, 2012, or 2012 R2 variant, Linux™ or GNU / Linux-based Debian, GNU / Linux, Debian GNU / kFreeBSD, DebianGNU / Hurd, Fedora™, Gentoo™, Linspire™, Mandriva, Red Hat® Linux, SuSE, and Ubuntu®, UNIX® variant Solaris™, AIX®, Mac™ OS X, FreeBSD®, OpenBSD, or NetBSD®.

401. The method according to claim 395, wherein the server device is cloud-based implemented as part of a public cloud-based service, and wherein the public cloud-based service comprises, is provided by, or is based on, Amazon Web Services® (AWS®), Microsoft® Azure™, or Google® Compute Engine™ (GCP).

402. The method according to claim 290, wherein the generating of the inference output data is provided as an Infrastructure as a Service (laaS) or as a Software as a Service (SaaS).

403. The method according to claim 290, for use with a first device in a first enclosure and with a second device in a second enclosure that communicate over a network, wherein the microphone is housed in, attached to, or integrated with, the first device, and wherein the model or the output data is generated in the second device.

404. The method according to claim 403, wherein the first and second devices are configured to communicate over the Internet.

405. The method according to claim 404, wherein the second device comprises a server device, the first device comprises a client device, and wherein the first and second devices communicate using a client- server architecture.

406. The method according to claim 403, wherein the first and second devices communicate over a wireless network, the method further comprising: coupling, by a first antenna in the first device, to the wireless network; coupling, by a second antenna in the second device, to the wireless network; sending, by the first device using a first wireless transceiver that is coupled to the first antenna, to the wireless network via the first antenna, a first data; and receiving, by the second device using a second wireless transceiver that is coupled to the second antenna, from the wireless network via the second antenna, the first data.

407. The method according to claim 406, wherein the first data comprises part of, or whole of, the captured additional speech, any representation thereof, or any function thereof.

408. The method according to claim 406, further comprising: sending, by the second device using the second wireless transceiver that is coupled to the second antenna, a second data; and receiving, by the first device using the first wireless transceiver that is coupled to the first antenna, the second data.

409. The method according to claim 408, wherein the second data comprises part of, or whole of, the output data, any representation thereof, or any function thereof.

410. The method according to claim 406, further comprising: sending, by the second device using the second wireless transceiver that is coupled to the second antenna, a second data; and receiving, by a third device using a third wireless transceiver that is coupled to a third antenna, the second data.

411. The method according to claim 410, wherein the second data comprises part of, or whole of, the output data, any representation thereof, or any function thereof.

412. The method according to claim 406, wherein the wireless network is over a licensed radio frequency band.

413. The method according to claim 406, wherein the wireless network is over an unlicensed radio frequency band.

414. The method according to claim 413, wherein the unlicensed radio frequency band is an Industrial, Scientific and Medical (ISM) radio band.

415. The method according to claim 414, wherein the ISM band comprises, or consists of, a 2.4 GHz band, a 5.8 GHz band, a 61 GHz band, a 122 GHz band, or a 244 GHz band.

416. The method according to claim 406, wherein the wireless network is a Wireless Personal Area Network (WPAN), the antenna comprises a WPAN antenna, and the wireless transceiver comprises a WPAN transceiver.

417. The method according to claim 416, wherein the WPAN is according to, compatible with, or based on, Bluetooth™ or Institute of Electrical and Electronics Engineers (IEEE) IEEE 802.15.1-2005 standards, or wherein the WPAN is a wireless control network that is according to, or based on, Zigbee™, IEEE 802.15.4-2003, or Z-Wave™ standards.

418. The method according to claim 416, wherein the WPAN is according to, compatible with, or based on, Bluetooth Low-Energy (BLE).

419. The method according to claim 406, wherein the wireless network is a Wireless Local Area Network (WLAN), the antenna comprises a WLAN antenna, and the wireless transceiver comprises a WLAN transceiver.

420. The method according to claim 419, wherein the WLAN is according to, compatible with, or based on, IEEE 802.11-2012, IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.1 In, or IEEE 802.1 lac.

421. The method according to claim 406, wherein the wireless network is a Wireless Wide Area Network (WWAN).

422. The method according to claim 421, wherein the WWAN is according to, compatible with, or based on, WiMAX network that is according to, compatible with, or based on, IEEE 802.16- 2009.

423. The method according to claim 421, wherein the wireless network is a cellular telephone network.

424. The method according to claim 406, wherein the wireless network is a cellular telephone network that is a Third Generation (3G) network that uses Universal Mobile Telecommunications System (UMTS), Wideband Code Division Multiple Access (W-CDMA) UMTS, High Speed Packet Access (HSPA), UMTS Time-Division Duplexing (TDD), CDMA2000 IxRTT, Evolution - Data Optimized (EV-DO), or Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE) EDGE-Evolution, or wherein the cellular telephone network is a Fourth Generation (4G) network that uses Evolved High Speed Packet Access (HSPA+), Mobile Worldwide Interoperability for Microwave Access (WiMAX), Long-Term Evolution (LTE), LTE-Advanced, Mobile Broadband Wireless Access (MBWA), or is based on IEEE 802.20-2008.

425. The method according to claim 403, wherein the first device communicates with the second device over a wired network, the method further comprising: coupling, by a connector in the second device, to the wired network; transmitting, by a wired transceiver in the second device that is coupled to the connector, a first data to the wired network via the connector; and receiving, by the wired transceiver in the second device that is coupled to the connector, the first data from the wired network via the connector.

426. The method according to claim 425, wherein the first data comprises part of, or whole of, the captured additional speech, any representation thereof, or any function thereof.

427. The method according to claim 425, wherein the wired network is a Personal Area Network (PAN), the connector is a PAN connector, and the wired transceiver is a PAN transceiver.

428. The method according to claim 425, wherein the wired network is a Local Area Network (LAN), the connector is a LAN connector, and the wired transceiver is a LAN transceiver.

429. The method according to claim 428, wherein the LAN is Ethernet based.

430. The method according to claim 429, wherein the LAN is according to, is compatible with, or is based on, IEEE 802.3-2008 standard.

431. The method according to claim 430, wherein the LAN is of according to, is compatible with, or is based on, a standard selected from the group consisting of lOBase-T, lOOBase-T,lOOBase-TX, 100Base-T2, 100Base-T4, lOOOBase-T, lOOOBase-TX, 10GBase-CX4, and lOGBase-T; and the LAN connector is an RJ-45 connector.

432. The method according to claim 430, wherein the LAN is according to, is compatible with, or is based on, a standard selected from the group consisting of lOBase-FX, lOOBase-SX, lOOBase-BX, lOOBase-LXlO, lOOOBase-CX, lOOOBase-SX, lOOOBase-LX, lOOOBase-LXlO, lOOOBase-ZX, lOOOBase-BXlO, lOGBase-SR, lOGBase-LR, lOGBase-LRM, lOGBase-ER, lOGBase-ZR, and 10GBase-LX4, and the LAN connector is a fiber-optic connector.

433. The method according to claim 425, wherein the wired network is a packet-based or switched-based Wide Area Network (WAN), the connector is a WAN connector, and the wired transceiver is a WAN transceiver.

434. The method according to claim 290, further comprising storing a part of, or whole of, the output data, a representation thereof, or a function thereof, in a database.

435. The method according to claim 434, further comprising storing a part of, or whole of, the captured additional speech, a representation thereof, or a function thereof, in a database.

436. The method according to claim 434, wherein the database is a relational database.

437. The method according to claim 436, wherein the relational database is Structured Query Language (SQL) based.

438. The method according to claim 434, wherein the output data is stored in the database as a record that is searchable using a label as a query.

439. The method according to claim 290, wherein the generating comprises generating of the inference output data as a single structured data stream that comprises a sequence of consecutive pairs, wherein each pair comprises one or more words from the additional multiple words followed by a label that is associated with the preceding one or more words in the pair.

440. The method according to claim 439, wherein the model uses, comprises, or is based on, an asynchronous transformer that is based on, or uses, an encoder-decoder Transformer architecture, wherein the single structured data stream is generated by predicting tokens as outputs of one or more text decoders, and wherein each token is predicted based on former predicted or used tokens.

441. The method according to claim 440, further comprising replacing, in each pair in the sequence, a predicted token associated with a predicting of the one or more words from the additional multiple words, with pre-defined one or more words.

442. The method according to claim 441, wherein each of the pre-defined one or more words comprises, or is based on, one or more words that are transcribed from the captured additional speech, or wherein each of the pre-defined one or more words comprises, or is based on, one ormore words that are transcribed from the captured additional speech as part of a training of the acoustic model.

443. The method according to claim 290, for use with a labels list, the method further comprising checking whether the first label is included in the labels list, and responsive to the checking of whether the first label is included in the labels list, performing a first action or a second action.

444. The method according to claim 443, wherein the first action is performed in response to determining that the first label is included in the labels list.

445. The method according to claim 444, wherein the second action is performed in response to determining that the first label is not included in the labels list.

446. The method according to claim 443, wherein the labels list includes at least 1, 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, or 1,000 labels.

447. The method according to claim 443, wherein the labels list includes less than 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, or 2,000 labels.

448. The method according to claim 443, for use with a categories list, the method further comprising checking whether the first category is included in the categories list, and responsive to the checking whether the first category is included in the categories list, performing a first action or a second action.

449. The method according to claim 448, wherein the first action is performed in response to further determining that the first category is included in the categories list.

450. The method according to claim 448, wherein the second action is performed in response to further determining that the first category is not included in the categories list.

451. The method according to claim 448, wherein the categories list includes 1, 2, 3, 4, or 5 categories.

452. The method according to claim 443, for use with a words list, the method further comprising checking whether a first word of the additional multiple words that is associated with the first label is included in the words list, and responsive to the checking whether the first word is included in the words list, performing a first action or a second action.

453. The method according to claim 452, wherein the first action is performed in response to further determining that the first word is included in the words list.

454. The method according to claim 452, wherein the second action is performed in response to further determining that the first word is not included in the words list.

455. The method according to claim 452, wherein the words list includes at least 1, 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, 2,000, 5,000 or 10,000 words.

456. The method according to claim 452, wherein the words list includes less than 2, 3, 5, 8, 10, 12, 15, 18, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, 2,000, 5,000, 10,000 or 20,000 words.

457. The method according to claim 443, wherein the first action is identical to, similar to, or different from, the second action.

458. The method according to claim 443, wherein the first or second action comprises initiating a communication over a network with a remote device.

459. The method according to claim 458, further comprising sending an alert message over the network to the remote device.

460. The method according to claim 443, wherein the performing of the first or second action comprises activating, deactivating, or controlling, an actuator that converts electrical energy to effect or produce a physical phenomenon.

461. The method according to claim 460, wherein the actuator is electrically powered from a power source, and converts electrical power from the power source to effect or produce the physical phenomenon.

462. The method according to claim 461, wherein the power source is an Alternating Current (AC) or a Direct Current (DC) power source.

463. The method according to claim 462, wherein the power source is a Direct Current (DC) power source that comprises a primary or a rechargeable battery.

464. The method according to claim 461, wherein the power source is an Alternating Current (AC) that comprises a domestic AC power that is nominally 120V AC / 60Hz or 230V AC / 50Hz.

465. The method according to claim 460, wherein the actuator is electrically powered from a power source via a switch connected between the power source and the actuator, wherein activating of the actuator comprises switching power from the power source to the actuator by the switch, and wherein deactivating of the actuator comprises disconnecting switching power from the power source to the actuator by the switch.

466. The method according to claim 465, wherein the switch is an electrically-controlled AC power Single-Pole-Double-Throw (SPDT) switch, wherein the switch comprises, is based on, is part of, or consists of, a relay, or wherein the switch is based on, comprises, or consists of, an electrical circuit.

467. The method according to claim 460, wherein the actuator is configured to influence, affect, create, or change the phenomenon in an object that is a gas, an air, a liquid, or a solid.

468. The method according to claim 460, wherein the actuator is operative to provide timedependent characteristic of the phenomenon selected from the group consisting of a time- integrated, an average, an RMS (Root Mean Square) value, a frequency, a period, a duty-cycle, a time-integrated, and a time-derivative.

469. The method according to claim 460, wherein the actuator is operative to provide spacedependent characteristic that is selected from the group consisting of a pattern, a linear density, a surface density, a volume density, a flux density, a current, a direction, a rate of change in a direction, and a flow.

470. The method according to claim 460, wherein the actuator consists of, or comprises, an electric light source that converts electrical energy into light, and emits visible or non-visible light for illumination or indication, and the non-visible light is infrared, ultraviolet, X-rays, or gamma rays, and wherein the electric light source consists of, or comprises, a lamp, an incandescent lamp, a gas discharge lamp, a fluorescent lamp, a Solid-State Lighting (SSL), a Light Emitting Diode (LED), an Organic LED (OLED), a polymer LED (PLED), or a laser diode.

471. The method according to claim 460, wherein the actuator consists of, or comprises, a motion actuator that causes linear or rotary motion, or wherein the motion actuator consists of, or comprises, a pneumatic actuator, hydraulic actuator, or electrical actuator.

472. The method according to claim 471, wherein the motion actuator consists of, or comprises, an electrical motor, that is a brushed motor, a brushless motor, an uncommutated DC motor, a DC stepper motor that is a Permanent Magnet (PM) motor, a Variable reluctance (VR) motor, a hybrid synchronous stepper motor, or wherein the motion actuator consists of, or comprises, an AC motor that is an induction motor, a synchronous motor, an eddy current motor, a singlephase AC induction motor, a two-phase AC servo motor, or a three-phase AC synchronous motor, and the AC motor is a split-phase motor, a capacitor-start motor, a Permanent-Split Capacitor (PSC) motor, an electrostatic motor, a piezoelectric actuator, or is a MEMS-based motor, a linear hydraulic actuator, a linear pneumatic actuator, a linear induction electric motor (LIM), a Linear Synchronous electric Motor (LSM), a piezoelectric motor, a Surface Acoustic Wave (SAW) motor, a Squiggle motor, an ultrasonic motor, or a micro- or nanometer combdrive capacitive actuator, a Dielectric or Ionic based Electroactive Polymers (EAPs) actuator, a solenoid, a thermal bimorph, or a piezoelectric unimorph actuator.

473. The method according to claim 460, wherein the actuator consists of, or comprises, a compressor or a pump and is operative to move, force, or compress a liquid, a gas or a slurry, wherein the pump is a direct lift pump, an impulse pump, a displacement pump, a valveless pump, a velocity pump, a centrifugal pump, a vacuum pump, or a gravity pump, wherein the actuator consists of, or comprises, a positive displacement pump that is a rotary lobe pump, a progressive cavity pump, a rotary gear pump, a piston pump, a diaphragm pump, a screw pump, a gear pump, a hydraulic pump, or a vane pump, or a rotary-type positive displacement pump that is an internal gear pump, a screw pump, a shuttle block pump, a flexible vane pump, a sliding vane pump, a rotary vane pump, a circumferential piston pump, a helical twisted roots pump, or a liquid ring vacuum pump, or wherein the actuator consists of, or comprises, a reciprocating-type positive displacement type that is a piston pump, a diaphragm pump, a plunge pump, a diaphragm valve pump, or a radial piston pump.

474. The method according to claim 460, wherein the actuator consists of, or comprises, a display screen for visually presenting information, wherein the display or the display screen consists of, or comprises, a monochrome, grayscale or color display, having an array of light emitters or light reflectors, wherein the display or the display screen consists of, or comprises, a projector selected from the group consisting of an Eidophor projector, Liquid Crystal on Silicon (LCoS or LCOS) projector, LCD projector, MEMS projector, and Digital Light Processing (DLP™) projector, wherein the display or the display screen consists of, or comprises, a video display supporting Standard-Definition (SD) or High-Definition (HD) standards, and is capable of scrolling, static, bold or flashing the presented information, wherein the video display is a 3D video display.

475. The method according to claim 460, wherein the actuator consists of, or comprises, a thermoelectric actuator and is a heater or a cooler, operative for affecting a temperature of a solid, a liquid, or a gas object, and is coupled to the object by conduction, convection, force convention, thermal radiation, or by a transfer of energy by phase changes, and wherein the thermoelectric actuator consists of, or comprises, a cooler based on a heat pump driving a refrigeration cycle using a compressor-based electric motor, an electric heater that is a resistance heater or a dielectric heater, or an induction heater that is solid-state based or is an active heat pump that uses, or is based on, the Peltier effect.

476. The method according to claim 460, wherein the actuator consists of, or comprises, a chemical or an electrochemical actuator, and is operative for producing, changing, or affecting a matter structure, properties, composition, process, or reactions, or wherein the electrochemicalactuator is operative for producing, changing, or affecting, an oxidation / reduction or an electrolysis reaction.

477. The method according to claim 460, wherein the actuator consists of, or comprises, an electromagnetic coil or an electromagnet operative for generating a magnetic or electric field, wherein the actuator consists of, or comprises, an electrical signal generator, wherein the signal generator is operative to output repeating or non-repeating electronic signals, wherein the signal generator is an analog signal generator having an analog voltage or analog current output, and wherein the output of the analog signal generator is a sine wave, a sawtooth wave, a step (pulse), a square wave, or a triangular wave, an Amplitude Modulation (AM), a Frequency Modulation (FM), or a Phase Modulation (PM) signal, or wherein the signal generator is an Arbitrary Waveform Generator (AWG) or a logic signal generator.

478. The method according to claim 460, wherein the actuator consists of, or comprises, a sounder for converting an electrical energy to omnidirectional, unidirectional, or bidirectional pattern of emitted, audible or inaudible, sound waves, wherein the sound waves are audible, and wherein the sounder is an electromagnetic loudspeaker, a piezoelectric speaker, an electrostatic loudspeaker (ESL), a ribbon or planar magnetic loudspeaker, or a bending wave loudspeaker, wherein the sounder is operative to emit a single or multiple tones, wherein the sounder is operative to perform continuous or intermittent operation, wherein the sounder is an electromechanical or a ceramic -based, and is an electric bell, a buzzer (or beeper), a chime, a whistle or a ringer, or wherein the sounder is a loudspeaker, and wherein the actuator is operative to store and play one or more digital audio content files.

479. The method according to claim 290, further comprising executing a chatbot, providing the inferred output data to the chatbot to be processed therein, and to provide a chatbot response that is based on at least part of the output data.

480. A non-transitory computer readable medium having computer executable instructions stored thereon, wherein the instructions include the steps of claim 241.

Citation Information

Patent Citations

  • Multilingual prosody generation

    US20160071512A1

  • System and method for extracting and using prosody features

    US20170103748A1

  • Extracting content from speech prosody

    US20200380960A1

  • Keyword determinations from conversational data

    US20200388288A1

Cited By

  • Glottis state monitoring method and system based on weak supervised learning

    CN116843945A

  • A glottal state monitoring method and system based on weakly supervised learning

    CN116843945B

  • Video note generation method and device based on AI

    CN120705353A

  • Multitask speech emotion recognition method based on parallel processing hybrid expert network

    CN120808822A

  • Synthesized data driven sublingual annotation method and device, equipment and storage medium

    CN121122249A